Evaluation
Measuring LLM quality. Perplexity, benchmarks, human eval, hallucination detection.
Learning path
Work through each subtopic in order. Click any file below to open its lesson. Track your progress with the checkbox at the bottom.
Subtopics
Each subtopic includes a concise explanation, code examples, and references.
Model Evaluation
Evaluation dashboard.
Training Loss
Live loss curve.
Validation Loss
Train vs validation graph.
Perplexity
Token prediction confidence visualization.
Accuracy
Correct/incorrect result grid.
Benchmarking
Model comparison table/chart.
Human Evaluation
Pairwise comparison UI.
Hallucination Evaluation
Claim verification interface.
Latency Evaluation
Request latency distribution.
Throughput Evaluation
Tokens-per-second dashboard.