Stacked self- multi-head attention and fully connected layers. With optional layer normalization applied to the final output. See 'Attention Is All You Need' https://arxiv.org/abs/1706.03762 for details. The use of this stack for batch major transformer is deprecated. Use GPipeBatchMajo
source not stored for this graph (policy: none)
no outgoing calls
searching dependent graphs…