MCPcopy Create free account
hub / github.com/tensorflow/lingvo / GPipeTransformerStack

Class GPipeTransformerStack

lingvo/core/layers_with_gpipe.py:576–888  ·  view source on GitHub ↗

Stacked self- multi-head attention and fully connected layers. With optional layer normalization applied to the final output. See 'Attention Is All You Need' https://arxiv.org/abs/1706.03762 for details. The use of this stack for batch major transformer is deprecated. Use GPipeBatchMajo

Source from the content-addressed store, hash-verified

source not stored for this graph (policy: none)

Calls

no outgoing calls

Used in the wild real call sites across dependent graphs

searching dependent graphs…