Alignment
Making models helpful and safe. RLHF, reward models, PPO, and DPO.
Learning path
Work through each subtopic in order. Click any file below to open its lesson. Track your progress with the checkbox at the bottom.
Subtopics
Each subtopic includes a concise explanation, code examples, and references.
Human Preferences
Pairwise response comparison.
Preference Data
Preference dataset table.
Reward Model
Preference โ reward visualization.
RLHF
Complete RLHF pipeline.
PPO
Policy-update visualization.
DPO
Chosen vs rejected probability comparison.