SLIDE 001
≈1 minuteLECTURE 01 · FALL 2026/27
How machines learn
to make moving worlds
LECTURE 01 · PRIMARY SOURCES
Read the systems, not the hype.
Three 2026 sources anchor the lecture’s current model landscape. Read each for what it discloses, what it measures, and what remains outside the evidence.
READ · APR 202601
Evolution of Video Generative Foundations
A technical survey of architectures, training objectives, data, evaluation, and the shift toward multimodal long narratives.
arXiv ↗TECHNICAL REPORT · APR 202602
Seedance 2.0
Evaluation across text-to-video, image-to-video, reference-to-video, audio, and multimodal task following.
ByteDance / arXiv ↗OPEN-SOURCE REPORT · 202603
MiniMax H3
A disclosed audio-video system with Context-IR, a 33B dense omni-modal transformer, latent visual and audio paths, and 2K regeneration.
MiniMax ↗TEACH OR STUDY
Download the complete technical edition.
Narrated lecture40-minute synchronized MP4Download video ↓
ElevenLabs audioComplete slide-timed narrationDownload MP3 ↓
PowerPointEditable 102-slide deckDownload PPTX ↓
Instructor transcript36-page slide-by-slide scriptDownload DOCX ↓
Transcript PDFPrint-ready teaching editionOpen PDF ↗
Plain textSearchable classroom scriptDownload TXT ↓
