AMCC 5160 · Fall 2026/27

AI-Driven Animation
& Video Generation

An art-and-technology course for creating, directing, understanding, and critically rethinking generative media—from images to interactive worlds.

Harry Yang · Tuesdays · 18:30 · MIL13002

Course at a glance

When

Tuesdays · 18:30–21:20

Where

MIL13002

Credits

3

Class size

50

Course description

This course treats generative AI as both a technical system and an artistic medium. Students learn how contemporary image, animation, video, audio, and world models work; use them to direct purposeful creative experiments; situate the work in visual culture; and evaluate authorship, provenance, consent, bias, and audience experience.

01

Understand systems

Build mental models that outlast individual products.

02

Direct experience

Connect technical choices to intention and audience.

03

Evaluate evidence

Make limitations and responsible use part of the work.

One medium · two inseparable lenses

See the system.
Direct the experience.

Lecture + discussion + lab

Art lens

Intention becomes form.

  • Visual language, composition, color, time, sound, and rhythm
  • Art direction, performance, narrative, interaction, and audience
  • Reference, context, authorship, critique, and exhibition

Technology lens

Capability becomes control.

  • Data, representations, transformers, diffusion, and flow
  • Conditioning, editing, adaptation, temporal modeling, and agents
  • Evaluation, reproducibility, provenance, safety, and limits
By the end, students can

explain a pipeline · direct a coherent work · build and document an experiment · critique with evidence · account for AI use · present a refined project

Assessment

A clear path from sharing to showcase.

Letter grades · 100% total

Weeks 2–11

Weekly presentation

10%

20% each

Assignments 1–3

60%

Nov 17 or 24

Final presentation

30%

Assignment deadlines Details will be announced in class and on Canvas.

A1 · 20%Sep 3023:59 HKT
A2 · 20%Oct 3123:59 HKT
A3 · 20%Nov 3023:59 HKT
Weekly presentation10%

Share current
or past work.

Around 10 minutes. Bring one artistic, design, research, or technical project and connect it to the course.

  • Explain its context, creative intention, and technical approach.
  • Share one lesson, limitation, or question.
  • Sign up for one of five slots in Weeks 2–11.
Final presentation30%

Refine one
assignment.

Choose one of Assignments 1–3 and present it as your strongest refined showcase. Ten minutes plus critique and Q&A.

  • State the problem, audience, or artistic intention.
  • Show evidence, one representative failure, and your own contribution.
  • Explain AI-tool use, iteration, and what should be tried next.

Tentative weekly schedule

Thirteen weeks.
One evolving medium.

Open a week to see its focus.

01Sep 1 · This weekFrom generative media to multimodal simulationSeeing + systems
Focus

How images become time, systems, and worlds. Technical mental models meet visual language, authorship, critique, provenance, and a first controlled experiment.

02Sep 8Representation, visual grammar, and unified multimodal modelsForm + representation
Focus

Pixels, latents, and visual tokens alongside composition, framing, color, and reference. How representation shapes what artists can edit, preserve, and vary.

03Sep 15Transformers, diffusion, flow—and the aesthetics of variationModels + chance
Focus

Autoregressive prediction, denoising, diffusion transformers, flow matching, guidance, seeds, and sampling—paired with rule-making, chance, seriality, and selection.

04Sep 22Image generation, editing, and art directionImage + direction
Focus

Subject and style control, reference grounding, masks, iterative editing, visual prompting, and agent-assisted workflows. Build a coherent visual system—not a bag of adjectives.

05Sep 29Adaptation, creative workflows, and authorshipCraft + workflow
Focus

LoRA and adapters, distillation, quantization, open versus hosted models, node and code pipelines, dataset curation, reproducibility, appropriation, and artistic voice.

06Oct 6Video foundations: time, motion, and cinematographyTime + motion
Focus

Video latents, temporal attention, motion representation, camera control, shot scale, blocking, rhythm, montage, and current open and hosted model families.

07Oct 13Controllable motion, choreography, and video-as-promptControl + choreography
Focus

Identity, pose, trajectory, depth, optical flow, and reference-video semantics—applied to performance, choreography, camera movement, continuity, and failure diagnosis.

08Oct 20Native audio-video, performance, and real-time charactersSound + performance
Focus

Synchronized sound, dialogue, lip-sync, expressive avatars, low-latency generation, voice, soundscape, performance direction, and cross-modal evaluation.

09Oct 27Critique, evaluation, provenance, and responsible releaseCritique + responsibility
Focus

Human preference and technical metrics meet studio critique. Prompt alignment, consistency, physics, synchronization, model cards, consent, copyright, and disclosure.

10Nov 3Interactive world models and explorable artWorlds + interaction
Focus

Action-conditioned prediction, explorable environments, persistent scenes, real-time simulation, installation, games, and physical-AI data. Case studies include Genie 3 and Cosmos 3.

11Nov 10Video reasoning, creative agents, and speculative practiceReasoning + speculation
Focus

Generation as a reasoning tool, multimodal agents, limits of treating video generators as world models, speculative futures, exhibition strategy, and presentation coaching.

12Nov 17Final presentations IShowcase
Focus

Students present one selected assignment as a refined showcase, followed by peer critique and discussion.

13Nov 24Final presentations IIShowcase
Focus

The second showcase session, peer critique, course synthesis, and future directions.

Current frontier · Fall 2026

The syllabus follows capabilities—not product hype.

01

Unified multimodality

Understanding, generation, and editing increasingly share one context and one workflow.

02

Native audio-video

Systems such as Veo 3.1 generate synchronized dialogue, ambience, effects, and moving images together.

03

Video as reasoning

Recent work tests whether generative video models can also segment, edit, infer affordances, and simulate tool use.

04

Interactive worlds

Genie 3 and Cosmos 3 push from passive clips toward action-conditioned, explorable, and physical-AI environments.

References & further reading Open list +

Teaching team

People, not prompts.