MDSection 31 of 38 ยท 4 subtopics
Inference Optimization
Faster inference: Flash Attention, speculative decoding, memory and compute trade-offs.
Learning path
Work through each subtopic in order. Click any file below to open its lesson. Track your progress with the checkbox at the bottom.
Subtopics
Each subtopic includes a concise explanation, code examples, and references.