LLM Systems
Production infrastructure: serving, load balancing, caching, monitoring, cost optimization.
Learning path
Work through each subtopic in order. Click any file below to open its lesson. Track your progress with the checkbox at the bottom.
Subtopics
Each subtopic includes a concise explanation, code examples, and references.
Model Serving
Production-serving architecture.
API Server
Request/response flow.
Request Handling
Request queue.
Load Balancing
Multiple inference workers.
GPU Allocation
GPU cluster allocation panel.
Caching
Cache-hit/miss visualization.
Monitoring
Production dashboard.
Logging
Request trace viewer.
Cost Optimization
Cost calculator.