Attention
The mechanism that powers Transformers. Scaled dot-product, softmax, masking, and causal attention.
Learning path
Work through each subtopic in order. Click any file below to open its lesson. Track your progress with the checkbox at the bottom.
Subtopics
Each subtopic includes a concise explanation, code examples, and references.
Attention
Token-to-token connection graph.
Self-Attention
Attention heatmap.
Query
Q-vector inspector.
Key
Key-vector inspector.
Value
Value-flow visualization.
Q/K/V Matrices
Three synchronized matrix viewers.
Attention Scores
Score matrix heatmap.
Scaled Dot-Product Attention
Formula pipeline.
Softmax
Logit-to-probability animation.
Attention Weights
Token heatmap.
Causal Attention
Triangular attention matrix.
Attention Mask
Interactive mask grid.