Attention Mechanism
Attention Score
Causal (Masked) Attention
Cross-Attention
Key Vector
Multi-Head Attention
Padding Mask
Query Vector
Scaled Dot-Product Attention
Self-Attention Mechanism
Value Vector
Decoder Block
Encoder Block
Position-wise Feed-Forward Network
Positional Embeddings
Pre-LN vs Post-LN
Transformer
Autoregressive
Sequence-to-Sequence
-
+8 more
-
-
-
-
-
A
Neural Machine Translation by Jointly Learning to Align and Translate (Bahdanau, Cho, Bengio) Pap… 1 term
-
-
-
-
-
-
H
Text generation — autoregressive generation with LLMs (Hugging Face Transformers docs) Docs 1 term
-
-
-
-
-
A
Understanding the Difficulty of Training Transformers (Liu et al., 2020) — Post-LN instability and Admin Pap… 1 term
No terms or links match your search.