Transformer
The full architecture: attention + FFN + residuals + norm. From encoder-decoder to GPT-style decoder-only.
Learning path
Work through each subtopic in order. Click any file below to open its lesson. Track your progress with the checkbox at the bottom.
Subtopics
Each subtopic includes a concise explanation, code examples, and references.
Transformer Architecture
Fully animated Transformer block.
Transformer Block
Expandable architecture diagram.
Feed-Forward Network
Neuron expansion/compression animation.
Residual Connection
Skip-connection animation.
Layer Normalization
Distribution visualization.
Dropout
Neuron masking animation.
Encoder
Encoder stack visualization.
Decoder
Decoder stack visualization.
Decoder-Only Transformer
GPT-style architecture.
GPT Architecture
Complete GPT pipeline.