Multimodal LLMs
Beyond text: images, audio, video. Vision encoders, cross-modal alignment.
Learning path
Work through each subtopic in order. Click any file below to open its lesson. Track your progress with the checkbox at the bottom.
Subtopics
Each subtopic includes a concise explanation, code examples, and references.
Multimodal Models
Text/image/audio input workspace.
Vision-Language Models
Image + text alignment view.
Image Embeddings
Image-to-vector pipeline.
Vision Encoder
Image patch visualization.
Audio Models
Waveform โ representation visualization.
Video Understanding
Frame timeline.
Cross-Modal Representation
Shared embedding space.