# AMCC5160 Lecture 02 — English / 中文

90-minute narrated video · 123 slides


## 001 — Generative Editing (00:00:00)

**00:00:00** Imagine you are the art director.

想象一下，你是艺术总监。

**00:00:02** Your exhibition opens tomorrow.

你的展览明天开幕。

**00:00:04** You love this glass pear, but you want to see it in ivory ceramic.

你很喜欢这只玻璃梨，但想看看它变成象牙白陶瓷的样子。

**00:00:08** You give the editor a tiny request: change the material, keep everything else.

你给编辑器一个小小的要求：改变材质，其他一切保持不变。

**00:00:13** Then you notice the reflection.

然后，你注意到了倒影。

**00:00:15** Should it still look like glass?

它还应该像玻璃一样吗？

**00:00:18** Suddenly, a three-word edit contains a whole argument about the world.

突然，一个简单的编辑指令，包含了关于世界如何运作的一整套判断。

**00:00:22** That is our starting point.

这就是我们的起点。

**00:00:24** We will follow this fictional artwork through image editing, exhibition design, and a moving shot.

我们将跟随这件虚构作品，经历图像编辑、展览设计，再进入一段运动镜头。

**00:00:31** The exhibition is called AFTER RAIN.

展览名叫《雨后》，英文是 AFTER RAIN。

**00:00:34** The pear is our recurring character.

这只梨会反复登场。

**00:00:37** By the end, you should be able to explain a model's mechanism, defend a visual choice, and catch a failure that a beautiful preview can hide.

到课程结束，你应该能够解释模型机制、为视觉选择提供理由，并发现精美预览可能掩盖的失败。

**00:00:46** First, look carefully at what you think must survive.

首先，仔细看看你认为必须保留的东西。


## 002 — One object, three creative decisions (00:00:50)

**00:00:50** We are going to give the same object three different jobs.

我们要让同一个物体承担三种不同的任务。

**00:00:53** First it is a technical puzzle: can its material change while its identity survives?

首先，它是一个技术难题：能否改变材质，同时保留它的身份特征？

**00:00:59** Then it becomes an artwork: can its placement and lighting make us feel that it is fragile?

接着，它成为艺术作品：能否通过摆放和灯光，让我们感到它的脆弱？

**00:01:06** Finally it becomes a character in time: can it disappear behind something and return as the same object?

最后，它成为时间中的角色：能否在被遮挡后重新出现，仍然是同一个物体？

**00:01:12** Those jobs demand different judgments.

这些任务需要不同的判断。

**00:01:15** A technically clean image may say very little.

技术上干净的图像，可能表达得很少。

**00:01:18** An expressive poster may contain an intentional impossibility.

富有表现力的海报，可能故意包含不可能的景象。

**00:01:23** A wonderful still may belong to a broken video.

一张精彩的静帧，可能来自一段存在问题的视频。

**00:01:26** As its job changes, you may find yourself approving a change that you would have rejected five minutes earlier.

随着任务改变，你可能会接受一个五分钟前还会拒绝的变化。


## 003 — The reflection is part of the edit (00:01:33)

**00:01:33** For a few seconds, ignore the pear.

先暂时别看梨。

**00:01:36** Look at the water beneath it.

看看它下面的水。

**00:01:38** Now look back at the body.

现在再看回梨的主体。

**00:01:40** If the body becomes opaque ceramic, what happens to the light that used to pass through the glass?

如果主体变成不透明的陶瓷，原先穿过玻璃的光会怎样？

**00:01:46** The answer cannot live entirely inside the object's outline.

答案不可能完全局限在物体轮廓之内。

**00:01:49** This comparison is a prepared classroom illustration.

这组对比是事先准备的课堂示意图。

**00:01:53** Use it to locate the consequences of the request: transmission, highlights, and the reflection below.

用它来找出指令引发的影响：透光、高光，以及下方的倒影。

**00:02:00** We still want the blue stem and recognizable silhouette.

我们仍然希望保留蓝色的柄和可辨认的轮廓。

**00:02:04** But preserving every surrounding pixel could preserve the wrong optics. a successful local edit may require a carefully justified change somewhere else.

但保留周围的每一个像素，可能会保留错误的光学效果。成功的局部编辑，有时需要在别处做出有充分理由的改变。


## 004 — Necessary change or accidental drift? (00:02:14)

**00:02:14** Let us make the opening comparison more disciplined.

让我们更严谨地审视开场的对比。

**00:02:17** The disappearance of transmitted amber light is consistent with changing transparent glass into opaque ceramic.

从透明玻璃变为不透明陶瓷后，透过的琥珀色光消失，是合理的。

**00:02:25** A changed blue stem, by contrast, would need a separate justification because the instruction did not request that transformation.

相比之下，如果蓝色的柄也变了，就需要另外解释，因为指令并没有要求这种变化。

**00:02:33** The reflection is more interesting.

倒影则更有意思。

**00:02:35** A material change can require its appearance to change, while its placement should still agree with the object and the water.

材质改变可能要求倒影外观随之变化，但它的位置仍应与物体和水面一致。

**00:02:44** We cannot classify every difference using a simple rule that change is bad.

我们不能用“变化就是坏事”这一简单规则来判断每个差异。

**00:02:49** We need a model of the intended scene.

我们需要理解目标场景应当如何运作。

**00:02:51** That is why the edit contract includes allowed consequences as well as invariants.

因此，编辑约定既要包含不变量，也要包含允许发生的连带变化。


## 005 — Tonight’s route (00:02:57)

**00:02:57** Here is the route through that puzzle.

下面是我们破解这个难题的路线。

**00:03:00** We begin inside an image editor: its representations, sampling process, and controls.

我们先进入图像编辑器内部，研究它的表示方式、采样过程和控制手段。

**00:03:09** Then we take the art director's seat and decide what those controls should accomplish.

然后坐到艺术总监的位置，决定这些控制手段应该实现什么。

**00:03:14** In the final technical section, the image starts moving and the preservation problem becomes harder.

在最后一个技术部分，图像开始运动，保留原有信息的问题也变得更难。

**00:03:21** The spoken script is planned for about ninety minutes.

讲解稿按约九十分钟规划。

**00:03:24** The marked discussions, the break, and student presentations need additional class time.

标出的讨论、课间休息和学生展示，需要额外的课堂时间。

**00:03:30** You will see the same pear return in different situations.

你会看到同一只梨出现在不同情境中。

**00:03:34** Each return asks us to revise our judgment, rather than learn a completely unrelated example.

每次重逢，都要求我们修正判断，而不是学习一个毫不相关的新例子。

**00:03:41** At the end, three readings let us question who sets the goal and who decides whether it was achieved.

最后三篇阅读，将让我们追问：谁来设定目标，又由谁判断目标是否达成？


## 006 — Three decisions you can defend (00:03:47)

**00:03:47** Here are three decisions I want you to be able to defend by the end.

到课程结束，我希望你能为这三类决定提供理由。

**00:03:51** When an edit fails, which control would you change, and why?

编辑失败时，你会调整哪一种控制手段，为什么？

**00:03:56** When two posters both look good, which belongs in this exhibition?

两张海报都很好看时，哪一张更适合这个展览？

**00:04:01** When a video looks convincing, which moment would you inspect before accepting it?

视频看起来可信时，你会在接受它之前检查哪个时刻？

**00:04:07** You do need a link between mechanism and consequence.

你确实需要把机制与结果联系起来。

**00:04:10** A mask tells us where; a reference can show us what; an artistic brief tells us why the result matters.

蒙版告诉我们改哪里；参考图可以展示改成什么；艺术创作说明告诉我们结果为何重要。

**00:04:17** Today we will practice making those links aloud.

今天，我们会练习把这些联系说清楚。

**00:04:21** The aim is to leave with reasons you can use when the next tool has a different interface.

目标是让你带走一套理由，即使下一款工具换了界面，也能继续使用。


## 007 — How image editors work (00:04:26)

**00:04:26** The pear gives us a concrete technical problem: change its material while retaining the information that makes it recognizable.

这只梨给我们一个具体的技术问题：改变材质，同时保留让人认出它的信息。

**00:04:34** We can describe this as conditional generation under preservation constraints.

我们可以把它描述为：在保留约束下进行条件生成。

**00:04:39** The inputs specify the change; the contract specifies what should survive.

输入指定要改变什么；编辑约定指定要保留什么。

**00:04:44** Throughout this section, connect each mechanism to one part of that problem.

在这一部分，请把每种机制对应到这个问题的某一环节。

**00:04:49** We begin by making our expectations explicit.

我们先把预期说清楚。


## 008 — Before Part 1: image editing (00:04:53)

**00:04:53** Before we open the machinery, decide what success should look like.

在打开技术黑箱之前，先确定什么才算成功。

**00:04:57** When glass becomes ceramic, what should stay the same?

玻璃变成陶瓷时，哪些东西应该保持不变？

**00:05:00** Which parts of the scene must change with the material?

场景的哪些部分必须随材质一起改变？

**00:05:03** What could a model misunderstand in that request?

模型可能会如何误解这个要求？

**00:05:07** Discuss these three questions and point to visible evidence for your choices.

请讨论这三个问题，并指出支持你们选择的可见证据。

**00:05:11** If your partner wants to preserve a detail that you want to change, identify the artistic intention behind each answer.

如果同伴想保留的细节恰好是你想改变的，请找出两种答案背后的艺术意图。

**00:05:19** Pause the video here.

请在这里暂停视频。

**00:05:21** We will return to these expectations when we write the worked ceramic-edit contract.

等到我们为陶瓷编辑写出完整约定时，再回头检视这些预期。


## 009 — The edit contract (00:05:30)

**00:05:30** The gallery has become a forest.

画廊变成了森林。

**00:05:32** Let us write the edit contract.

让我们写下编辑约定。

**00:05:34** The intended change is the environment.

我们想改变的是环境。

**00:05:37** The pear's identity, its pedestal, and the framing should remain recognizable.

梨的身份特征、底座和构图，都应该仍然可辨认。

**00:05:43** The ambient light and reflections are allowed to adapt to the forest.

环境光和反射则允许适应森林。

**00:05:48** Why include that last category?

为什么要加入最后这一类？

**00:05:50** Because preserving every visible relationship would contradict the new setting.

因为保留所有可见关系，会与新环境矛盾。

**00:05:55** Imagine retaining a bright gallery-window reflection inside a dark forest.

想象一下，在昏暗森林里还保留着明亮画廊窗户的反光。

**00:05:59** The model might preserve the source faithfully and still produce an implausible scene.

模型可能忠实保留了原图，却仍然生成了不合理的场景。

**00:06:05** Before generating, separate the requested change, the invariants, and the consequences that should adapt.

生成之前，要区分所要求的变化、不变量，以及应当随之调整的连带效果。

**00:06:12** Afterwards, inspect those categories separately.

生成之后，再分别检查这些类别。


## 010 — An invariant can be semantic or exact (00:06:16)

**00:06:16** Suppose a curator says, keep the pear exactly the same.

假设策展人说：让这只梨完全保持原样。

**00:06:20** There are at least two meanings hiding in that sentence.

这句话至少隐藏着两种含义。

**00:06:24** One means that a visitor should recognize the same sculpture in a new room.

一种是：观众在新房间里，仍然能认出这是同一件雕塑。

**00:06:29** Its highlights may change with the room.

它的高光可以随着房间而变化。

**00:06:31** The other means that selected pixel values must remain identical.

另一种是：选定像素的数值必须完全一致。

**00:06:36** Those are different requirements.

这是两种不同的要求。

**00:06:38** Generative conditioning can encourage recognizable identity.

生成式条件控制可以帮助保留可辨认的身份特征。

**00:06:41** Copying source pixels or compositing can enforce exact preservation in a specified region.

复制原图像素或进行合成，可以强制指定区域完全不变。

**00:06:47** Neither requirement automatically makes the whole image coherent: a copied region may now have the wrong lighting.

但这两种要求都不会自动保证整幅图像协调：复制过来的区域，光照可能已经不合适。

**00:06:55** Before choosing a tool, settle the meaning of same.

选择工具之前，先确定“相同”到底是什么意思。


## 011 — Editing as conditional generation (00:06:59)

**00:06:59** Read the expression as a distribution of possible edited results, given the source, instruction, and optional references.

把这个表达式理解为：给定原图、指令和可选参考信息后，可能产生的编辑结果的分布。

**00:07:07** The symbol theta represents the model's learned parameters.

符号 theta 表示模型学习到的参数。

**00:07:11** The vertical bar means 'given.' We are not asking for a unique answer in every case.

竖线表示“给定”。我们并不是在所有情况下都要求唯一答案。

**00:07:17** Two different forest scenes might both satisfy the same brief.

两个不同的森林场景，都可能满足同一份创作说明。

**00:07:21** However, a source image is a condition rather than a promise that every unmentioned pixel will be copied.

但原图只是条件，并不意味着每一个未提及的像素都会被复制。

**00:07:27** To judge success, we need a more precise contract than 'it looks plausible.'

要判断是否成功，我们需要比“看起来合理”更精确的约定。


## 012 — A condition is not a hard constraint (00:07:33)

**00:07:33** The probability expression says that the system produces a candidate given conditions.

概率表达式表示：系统在给定条件下生成一个候选结果。

**00:07:38** It does not contain a certificate that the candidate passes our checks.

它并不附带证明，保证这个结果通过我们的检查。

**00:07:43** The acceptance step is a separate decision that we impose on the result.

是否接受，是我们对结果做出的另一个独立判断。

**00:07:47** Imagine two forest outputs.

想象两个森林版本。

**00:07:50** Both look plausible, but one changes the pear's stem and one retains it.

两个都看起来合理，但一个改变了梨柄，另一个保留了它。

**00:07:54** The model can assign probability to both; our contract can reject one.

模型可以给两者都分配概率，而我们的约定可以拒绝其中一个。

**00:08:00** In practice, acceptance may involve human inspection, image comparisons, or explicit production constraints.

实际上，验收可能依赖人工检查、图像对比，或明确的制作约束。


## 013 — A smaller workspace: image latents (00:08:07)

**00:08:07** The RGB image has 1024 by 1024 pixels and three channels: just over three million scalar values.

RGB 图像有 1024 乘 1024 个像素，每个像素有三个通道，总共略多于三百万个标量值。

**00:08:15** With spatial reduction by sixteen and sixty-four latent channels, the representation becomes 64 by 64 by 64: 262,144 values.

空间尺寸缩小十六倍、潜在通道数为六十四时，表示变成 64 乘 64 乘 64，即 262,144 个数值。

**00:08:29** Notice the trap.

注意这里的陷阱。

**00:08:31** Sixteen times smaller along each spatial axis does not mean sixteen times fewer values overall, because the number of channels changes too.

每个空间轴缩小十六倍，并不意味着总数值量减少十六倍，因为通道数也变了。

**00:08:39** Here the scalar count falls by a factor of twelve.

在这个例子中，标量数量减少了十二倍。

**00:08:43** That is specific to this configuration.

这个比例只适用于这一具体配置。

**00:08:45** The model works in a learned representation, and its decoder must recover visible detail.

模型在学习得到的表示中工作，解码器必须从中恢复可见细节。

**00:08:52** Our next question is what that compression already loses before we request any edit.

接下来要问的是：在我们提出任何编辑要求之前，压缩已经丢失了什么？


## 014 — Is the detail lost before the edit? (00:08:57)

**00:08:57** Suppose the small exhibition caption is already blurry after an edit.

假设编辑后，展览的小字号说明文字变模糊了。

**00:09:02** You might spend an hour rewriting the prompt to say, preserve every letter.

你可能花一个小时重写提示词，强调要保留每一个字母。

**00:09:08** There is a quicker diagnostic: encode the original image and decode it again without requesting a change.

有一个更快的诊断方法：不要求任何改变，先把原图编码，再解码回来。

**00:09:15** If the letters are already damaged in that round trip, part of the problem lies in representation and reconstruction.

如果这次往返已经损坏了文字，那么问题的一部分就在表示和重建环节。

**00:09:23** A more emphatic instruction cannot restore a detail the editing path never retained faithfully.

再强硬的指令，也无法恢复编辑流程从未忠实保留的细节。

**00:09:29** Look at fine text, thin edges, and small textures first.

首先检查小字、细线边缘和微小纹理。

**00:09:33** If reconstruction is sound but the edit damages them, investigate the editing stage.

如果重建没有问题，而编辑损坏了它们，就调查编辑阶段。

**00:09:39** It also explains why exact exhibition typography deserves its own editable layer.

这也说明，要求精确的展览文字排版，应当保留在独立的可编辑图层中。


## 015 — Why token count matters (00:09:45)

**00:09:45** Now consider the number of tokens a transformer processes.

现在考虑 Transformer 需要处理多少个 token。

**00:09:49** In our toy example, one token per position on a 64 by 64 grid gives 4,096 tokens.

在这个简化例子里，64 乘 64 的网格上，每个位置对应一个 token，共有 4,096 个。

**00:09:58** Grouping each two-by-two region gives 1,024 tokens, a reduction by four.

把每个二乘二区域组合起来，就得到 1,024 个 token，数量减少四倍。

**00:10:06** Dense self-attention compares token pairs.

稠密自注意力会比较 token 对。

**00:10:09** Squaring the two counts gives 16,777,216 and 1,048,576 pair scores.

将两个数量分别平方，得到 16,777,216 和 1,048,576 个配对分数。

**00:10:21** That is a reduction by sixteen.

数量减少了十六倍。

**00:10:24** This is why representation choices matter so much for cost.

这就是为什么表示方式会如此显著地影响成本。

**00:10:28** It is not a claim that the whole system runs sixteen times faster.

但这不等于整个系统的运行速度会提高十六倍。

**00:10:32** Text tokens, reference tokens, other layers, and implementation details also matter.

文本 token、参考图 token、其他网络层和实现细节，也会影响结果。


## 016 — Double the resolution: count the cost (00:10:38)

**00:10:38** Before reading the second number, make a prediction.

看第二个数字之前，先预测一下。

**00:10:42** We double the image width and double its height, keeping the patch scheme fixed.

我们把图像宽度和高度都加倍，同时保持分块方式不变。

**00:10:47** How many image tokens do we get?

图像 token 会变成多少？

**00:10:49** Four times as many.

原来的四倍。

**00:10:51** In dense self-attention, each token can compare with every other token.

在稠密自注意力中，每个 token 都可以与其他所有 token 比较。

**00:10:56** Four times four gives sixteen times as many pair scores.

四乘四，配对分数的数量就变成十六倍。

**00:10:59** This calculation describes the image-token attention component, not a promise that the entire system becomes sixteen times slower.

这个计算只描述图像 token 的注意力部分，并不保证整个系统会慢十六倍。

**00:11:08** Other operations and implementations matter.

其他操作和实现方式同样重要。

**00:11:12** The useful habit is to ask what grew: pixels, tokens, comparisons, or measured runtime.

一个有用的习惯是追问：增加的究竟是像素、token、比较次数，还是实测运行时间？

**00:11:19** They are related quantities, but they are not interchangeable.

这些量相互关联，但不能互换。


## 017 — Inside a current image editor (00:11:23)

**00:11:23** Let us walk through this architecture from the input side.

让我们从输入端走一遍这个架构。

**00:11:27** The source image contributes semantic information through a vision-language component and visual information through a VAE.

原图通过视觉语言组件提供语义信息，并通过 VAE 提供视觉信息。

**00:11:35** The target stream begins with an evolving noisy representation.

目标分支从一个不断演化的带噪表示开始。

**00:11:39** The transformer processes that stream together with the conditions.

Transformer 将这个分支与条件信息一起处理。

**00:11:44** During training, we possess an example of the desired edited target, so we can measure what the model should learn.

训练时，我们拥有所需编辑结果的样本，因此可以衡量模型应该学习什么。

**00:11:51** During inference, we have the source and request but not the desired target.

推理时，我们有原图和要求，却没有所需的目标图像。

**00:11:55** The model must generate it.

模型必须把它生成出来。

**00:11:57** That difference is essential.

这个区别至关重要。

**00:12:00** The architecture does not quietly receive the answer when you ask it to edit your photograph.

当你让模型编辑照片时，架构并没有偷偷收到答案。


## 018 — The missing target at inference (00:12:06)

**00:12:06** There is a piece of information on the left that we do not possess on the right: the desired edited image.

左边有一项信息，是右边没有的：期望得到的编辑图像。

**00:12:12** During supervised training, a source, an instruction, and a target show the model what a successful change looks like.

在监督训练中，原图、指令和目标图一起，向模型展示成功的变化是什么样子。

**00:12:20** During use, we provide the request because that target does not yet exist.

使用时，我们之所以提出要求，正是因为目标图像还不存在。

**00:12:26** This distinction explains a surprisingly common confusion.

这个区别解释了一个相当常见的混淆。

**00:12:29** Showing a reference to an editor is not automatically teaching new weights.

给编辑器看一张参考图，并不自动意味着学习新的权重。

**00:12:34** It may simply be conditioning this particular generation.

它可能只是在为这一次生成提供条件。

**00:12:38** Adaptation, which we will reach with LoRA, changes trainable parameters.

适配则会改变可训练参数，讲到 LoRA 时我们会展开。

**00:12:43** It is available when we construct the learning signal, and absent when we ask the trained model to make our ceramic pear.

构造学习信号时，目标图像是已知的；要求训练好的模型生成陶瓷梨时，它则是未知的。


## 019 — Two representations of the source (00:12:50)

**00:12:50** Think of these two representations using our pear.

用这只梨来理解这两种表示。

**00:12:54** Semantic features help with questions such as: which object is the pear, what does ceramic mean, and what does 'it' refer to in the instruction?

语义特征帮助回答：哪个物体是梨，陶瓷是什么意思，以及指令中的“它”指什么？

**00:13:05** Visual latents provide appearance information, including shapes and textures that help keep the particular source recognizable.

视觉潜变量提供外观信息，包括形状和纹理，帮助保留原图的独特特征。

**00:13:13** This is a conceptual distinction, not a claim that the branches have perfectly isolated responsibilities.

这是一种概念上的区分，并不是说两个分支的职责完全隔离。

**00:13:20** Learned representations can overlap in what they encode.

学习到的表示所编码的信息可能重叠。

**00:13:24** If an edit fails, ask whether the model misunderstood the request or lost important visual information.

编辑失败时，要问：模型是误解了要求，还是丢失了重要的视觉信息？


## 020 — OpenSubject: identity across new scenes (00:13:31)

**00:13:31** How do you teach a system the difference between a pear and this particular pear?

如何教会系统区分“一只梨”和“这只特定的梨”？

**00:13:36** That question connects to OpenSubject, research I coauthored with Yexin Liu and our collaborators.

这个问题与 OpenSubject 有关，这是我与Yexin Liu及其他合作者共同完成的研究。

**00:13:43** A video offers something valuable: repeated observations of a subject as the view changes.

视频提供了一种宝贵资源：随着视角变化，对同一主体进行反复观察。

**00:13:49** The shared identity is a learning opportunity.

共同的身份特征提供了学习机会。

**00:13:53** Follow the main path in the figure.

请沿着图中的主路径看。

**00:13:55** We curate clips, verify subjects across frames, and select diverse pairs.

我们筛选视频片段，跨帧验证主体，并挑选具有多样性的图像对。

**00:14:01** Inpainting or outpainting helps synthesize reference inputs, followed by verification.

通过图像修补或扩图来合成参考输入，再进行验证。

**00:14:08** The corpus contains 2.5 million samples; the contribution is training data and a benchmark.

数据集包含二百五十万个样本；这项工作的贡献是训练数据和评测基准。

**00:14:15** For our exhibition, imagine the same sculpture photographed in several rooms.

对于我们的展览，可以想象在不同房间拍摄同一件雕塑。

**00:14:20** We want freedom to change the setting without losing its distinctive features.

我们希望自由改变环境，同时不丢失它的独特特征。

**00:14:24** That is a classroom application of the identity problem, rather than a claim that our fictional pear was tested in the paper.

这是身份保持问题的课堂应用，并不意味着论文测试过这只虚构的梨。


## 021 — Flow matching: the training target (00:14:32)

**00:14:32** The editor needs a way to learn how to move from noise toward an image.

编辑器需要学习如何从噪声走向图像。

**00:14:36** In this simple flow-matching setup, we construct an intermediate latent by mixing noise, epsilon, with a known target latent, z one.

在这个简单的流匹配设定中，我们将噪声 epsilon 与已知目标潜变量 z₁ 混合，构造中间潜变量。

**00:14:47** The mixing variable t tells us where we are along that training path.

混合变量 t 告诉我们处于这条训练路径的什么位置。

**00:14:52** For this straight interpolation, the target velocity is the target latent minus the noise.

对于这种直线插值，目标速度等于目标潜变量减去噪声。

**00:14:57** The model sees an intermediate state and learns to predict that direction, given the source and instruction.

模型看到中间状态，并在原图和指令的条件下，学习预测这个方向。

**00:15:04** There is also an important clock distinction: t here is the generative process's time.

还要区分两种时钟：这里的 t 是生成过程中的时间。

**00:15:09** Later, our video will have another time axis—the seconds that pass in the scene.

稍后的视频还有另一条时间轴：场景中实际流逝的秒数。


## 022 — The interpolation endpoints (00:15:15)

**00:15:15** Before trusting an equation, check its easiest cases.

相信一个方程之前，先检查最简单的情况。

**00:15:19** At t equals zero, the coefficient on the target becomes zero, so we recover the noise.

当 t 等于零时，目标的系数为零，所以我们得到噪声。

**00:15:26** At t equals one, the coefficient on noise vanishes, so we recover the target latent.

当 t 等于一时，噪声的系数消失，所以我们得到目标潜变量。

**00:15:34** Along this simple straight path, the difference between target and noise gives the direction we train the model to predict.

沿着这条简单直线路径，目标减去噪声，就是我们训练模型预测的方向。

**00:15:41** You do not need to imagine a recognizable half-finished image at every intermediate point.

不必想象每个中间点都对应一张能辨认的半成品图像。

**00:15:47** The calculation takes place in a learned representation, and this path is a teaching construction rather than the only possible design.

计算发生在学习得到的表示中；这条路径用于教学，并不是唯一可能的设计。


## 023 — Predict the next sampling step (00:15:57)

**00:15:57** Here is the smallest possible sampling calculation.

这是一个最小的采样计算例子。

**00:16:01** One coordinate currently equals 0.20.

当前某个坐标值是 0.20。

**00:16:04** Its predicted velocity is 0.60, and the next step has size 0.10.

预测速度为 0.60，下一步的步长为 0.10。

**00:16:12** Before looking at the result, say what operation should happen.

看结果之前，先说说应该做什么运算。

**00:16:17** We add a small amount of motion to the current state: 0.20 plus 0.10 times 0.60, giving 0.26.

在当前状态上加一小步：0.20 加上 0.10 乘 0.60，得到 0.26。

**00:16:27** In a real latent, many coordinates update together.

在真实的潜变量中，很多坐标会同时更新。

**00:16:31** These numbers are invented to show an Euler step, and practical samplers can use more elaborate solvers.

这些数字是为了演示欧拉步而编造的，实际采样器可以采用更复杂的求解器。

**00:16:37** But the sequence is now less mysterious: predict a direction, take a step, repeat, then decode.

但这个过程现在不再那么神秘：预测方向，迈出一步，重复，然后解码。


## 024 — A better solver cannot clarify the brief (00:16:45)

**00:16:45** Imagine following a beautifully accurate set of directions to the wrong gallery.

想象你沿着极其精确的路线指引，却走向了错误的画廊。

**00:16:50** Taking smaller steps will not repair the destination.

把步子迈小一点，并不能修正目的地。

**00:16:54** The same distinction helps with generative sampling.

同样的区分也有助于理解生成采样。

**00:16:57** A better numerical solver can follow a learned field more accurately.

更好的数值求解器，可以更准确地沿着学到的向量场前进。

**00:17:02** That alone does not settle an ambiguous editing request.

但仅凭这一点，无法解决含糊的编辑要求。

**00:17:06** If the pear's boundary is unstable across sampling settings, numerical or model behavior may matter.

如果不同采样设置下梨的边界不稳定，数值计算或模型行为可能有影响。

**00:17:13** If we never specified whether the glass reflection should become ceramic, we also have a brief problem.

如果我们从未说明玻璃倒影是否应变成陶瓷倒影，那么创作说明本身也有问题。

**00:17:20** Ask whether the system struggled to realize a clear instruction, or whether we have not agreed on what success means.

要问：系统难以实现一个清楚的指令，还是我们尚未就成功的含义达成一致？

**00:17:27** Those situations call for different next actions.

这两种情况需要不同的下一步行动。


## 025 — Where edit supervision comes from (00:17:31)

**00:17:31** Where does the ability to follow an edit instruction come from?

遵循编辑指令的能力从何而来？

**00:17:35** A training example can contain a source image, an instruction, and the corresponding edited target.

一个训练样本可以包含原图、指令，以及对应的编辑目标图像。

**00:17:42** InstructPix2Pix is an early example that used synthetic editing data to teach this relationship.

InstructPix2Pix 是一个较早的例子，它用合成编辑数据来学习这种关系。

**00:17:48** Suppose a training pair says 'make the object ceramic,' but its target also moves the camera and replaces the background.

假设某个训练对的指令是“把物体变成陶瓷”，但目标图还移动了相机并替换了背景。

**00:17:56** The learning signal no longer cleanly identifies the requested transformation.

这时，学习信号就不再能清晰地指向所要求的变换。

**00:18:01** The model may learn unwanted associations.

模型可能学到不希望出现的关联。

**00:18:05** This gives us an important data-quality question: did the target accomplish the requested change while preserving what should remain?

因此，一个重要的数据质量问题是：目标图是否完成了要求的变化，同时保留了应该保留的内容？

**00:18:14** Attractive targets alone are not enough to teach dependable editing.

仅有好看的目标图，还不足以教会可靠的编辑。


## 026 — What a mismatched training pair teaches (00:18:18)

**00:18:18** A training pair is a lesson, and lessons can accidentally teach the wrong thing.

一个训练对就是一堂课，而课程可能无意中教错东西。

**00:18:23** Imagine a source glass pear and a target ceramic pear that also has a different background.

想象原图是一只玻璃梨，目标图是一只陶瓷梨，但背景也变了。

**00:18:29** The written request says only to change the material.

文字要求却只说改变材质。

**00:18:32** Which visible changes should the model associate with that request?

模型应当把哪些可见变化与这个要求联系起来？

**00:18:36** If this mismatch recurs in the training data, the model can learn an unwanted association between a material change and a scene change.

如果训练数据中反复出现这种不匹配，模型可能把材质变化与场景变化错误地联系起来。

**00:18:44** Compare the instruction with the actual difference between source and target.

请将指令与原图、目标图之间的实际差异进行比较。

**00:18:49** Ask what supervision is rewarding, including changes nobody meant to label.

要问监督信号在奖励什么，包括那些没人打算标注的变化。

**00:18:54** The preservation contract begins in the examples used to teach the editor.

保留约定，应当从用来教会编辑器的样本开始。


## 027 — Classifier-free guidance (CFG) (00:18:59)

**00:18:59** Sometimes a conditioned prediction moves in the right direction but too weakly.

有时，条件预测的方向是对的，但力度不足。

**00:19:04** Classifier-free guidance uses the difference between a conditioned prediction and a baseline prediction to steer the result.

无分类器引导利用条件预测与基线预测之间的差异，来引导结果。

**00:19:12** Read the formula as baseline, plus a scaled change in direction.

可以把公式读成：基线，加上经过缩放的方向变化。

**00:19:16** When s equals one, the baseline terms cancel and we recover the conditioned prediction.

当 s 等于一时，基线项抵消，我们得到条件预测。

**00:19:22** Above one, we extrapolate beyond it.

大于一时，我们会外推到条件预测之外。

**00:19:25** That can strengthen the requested change, but it can strengthen errors as well.

这可以强化所要求的变化，但也可能强化错误。

**00:19:30** In an editing system, the baseline may still retain image information; baseline does not always mean no information at all.

在编辑系统中，基线可能仍保留图像信息；基线并不总意味着完全没有信息。

**00:19:39** Think of guidance as a specific operation on predictions, rather than a universal quality dial.

把引导理解为对预测值的一种具体运算，而不是万能的质量旋钮。


## 028 — Guidance goes beyond the prediction (00:19:46)

**00:19:46** The baseline predicts 0.2 and the conditioned version predicts 0.5.

基线预测是 0.2，条件预测是 0.5。

**00:19:52** With guidance scale two, do we get a value somewhere between them?

引导强度设为二，结果会落在两者之间吗？

**00:19:57** No.

不会。

**00:19:58** We take 0.2 plus twice the difference of 0.3, which gives 0.8.

我们用 0.2 加上两倍的差值 0.3，得到 0.8。

**00:20:05** We have moved beyond the conditioned prediction.

我们已经越过了条件预测。

**00:20:08** If the useful direction includes a slight mistake, we can amplify both.

如果有效方向中包含一点错误，那么两者都可能被放大。

**00:20:13** For the pear, the new material might become clearer while the stem or boundary becomes less faithful.

对于这只梨，新材质可能更明显，但柄或边界可能变得不那么忠实。

**00:20:19** Compare the requested change and preservation separately.

请分别比较所要求的变化和保留效果。

**00:20:24** One slider can move those judgments in opposite directions; a single overall impression can hide the tradeoff.

同一个滑块可能让两项评价朝相反方向变化；单一的总体印象会掩盖这种取舍。


## 029 — Different controls carry different information (00:20:31)

**00:20:31** Different controls communicate different kinds of information.

不同控制手段传达不同类型的信息。

**00:20:35** Text describes a requested change.

文本描述所要求的变化。

**00:20:38** A mask specifies a region.

蒙版指定区域。

**00:20:41** A spatial condition can describe pose, depth, or edges.

空间条件可以描述姿态、深度或边缘。

**00:20:46** An image reference can supply appearance information that would be difficult to express precisely in words.

参考图可以提供难以用语言精确表达的外观信息。

**00:20:52** These are not interchangeable knobs, and every product does not expose all of them.

这些控制并不能互相替代，也不是每款产品都提供全部选项。

**00:20:57** ControlNet and IP-Adapter are examples of distinct architectural approaches.

ControlNet 和 IP-Adapter 就是两种不同架构思路的例子。

**00:21:01** If exact untouched pixels are essential, a generation mask alone may be insufficient; explicit copying or compositing can enforce that requirement.

如果必须精确保留未编辑像素，仅有生成蒙版可能不够；明确复制像素或合成，可以强制实现这一要求。

**00:21:12** Choose a control by asking what information is missing, then check whether the resulting image actually respected it.

选择控制手段时，先问缺少什么信息，再检查结果是否真正遵循了它。

**00:21:18** A mask indicates a permitted region; it does not by itself solve seams or the reflection of a changed material.

蒙版指出允许编辑的区域，但它本身不能解决接缝，也不能处理材质改变后的倒影。

**00:21:25** The next slide names an architecture that learns to use a spatial condition.

下一页介绍一种学习使用空间条件的架构。


## 030 — ControlNet: add spatial conditioning (00:21:30)

**00:21:30** Suppose your sentence is understood, but the pear keeps changing shape.

假设模型理解了你的话，但梨的形状总在变化。

**00:21:34** You can describe its outline with more adjectives, or supply spatial evidence.

你可以用更多形容词描述轮廓，也可以直接提供空间证据。

**00:21:39** ControlNet gives an edge map, depth map, or pose a learned route into a compatible diffusion model.

ControlNet 为边缘图、深度图或姿态信息，提供了一条进入兼容扩散模型的可学习路径。

**00:21:46** Trace the two paths in this original architecture.

请追踪原始架构中的两条路径。

**00:21:49** The pretrained path stays frozen.

预训练主路径保持冻结。

**00:21:51** A trainable copy processes the added condition, and zero-initialized one-by-one convolutions connect its features to the main path.

一个可训练副本处理新增条件，零初始化的一乘一卷积把它的特征连接到主路径。

**00:22:00** Initially those connections contribute zero; training learns their contribution.

初始时，这些连接的贡献为零；训练会学出它们应有的贡献。

**00:22:06** The copied branch itself starts from pretrained weights, not all zeros.

副本分支本身从预训练权重开始，而不是全部从零开始。

**00:22:11** Now return to our request.

现在回到我们的要求。

**00:22:13** Edges can help specify the outline.

边缘可以帮助指定轮廓。

**00:22:16** They cannot, by themselves, tell us whether the reflected ceramic looks convincing or the blue stem retains its identity.

但仅凭边缘，无法判断陶瓷倒影是否可信，也无法保证蓝色梨柄保留身份特征。


## 031 — Multiple references need distinct identities (00:22:24)

**00:22:24** With several references, the system must know which reference contributes which information.

当有多张参考图时，系统必须知道每张图分别提供什么信息。

**00:22:30** Imagine asking for the sculpture from image one and the atmosphere from image two.

想象你要求使用图一的雕塑，以及图二的氛围。

**00:22:37** If the references become confused, you may get the wrong object with the right lighting.

如果参考图的角色混淆，可能会得到光照正确、物体却错误的结果。

**00:22:43** The illustrated research approach uses separators and image-index information to distinguish inputs.

图示研究方法使用分隔符和图像索引信息来区分输入。

**00:22:48** In an art brief, name the role of each image rather than presenting a pile of vaguely related inspiration.

在艺术创作说明中，要说明每张图的角色，而不是堆放一批关系模糊的灵感图。

**00:22:55** Then inspect for leakage: did a composition reference accidentally replace the subject?

然后检查信息是否串用：构图参考是否意外替换了主体？

**00:23:02** This figure describes one proposed mechanism, not a universal design used by every editor.

这张图描述的是一种提出的机制，并不是所有编辑器通用的设计。


## 032 — A reference-role failure is visible in the output (00:23:08)

**00:23:08** Look at these published multi-image examples with reference roles in mind.

请带着“参考图角色”的概念，观察这些已发表的多图示例。

**00:23:13** Before judging the output, identify what each input was supposed to contribute.

评价输出之前，先确定每个输入原本应该贡献什么。

**00:23:18** Then trace the subject and the requested transformation into the result.

然后在结果中追踪主体和所要求的变换。

**00:23:23** A reference-role failure can look superficially attractive.

参考图角色混淆的结果，表面上也可能很好看。

**00:23:27** The system may borrow the wrong object's appearance or import a background that was never intended.

系统可能借用了错误物体的外观，或引入了原本不需要的背景。

**00:23:33** We are examining qualitative examples from a particular paper, not conducting a broad comparison between products.

我们正在查看某篇论文的定性示例，并不是对产品做全面比较。

**00:23:40** The figure helps us practice an inspection method: follow the intended contribution of each input and look for unintended transfers between them.

这张图帮助我们练习一种检查方法：追踪每个输入的预期贡献，寻找意外的信息转移。


## 033 — Attention as weighted information gathering (00:23:49)

**00:23:49** Attention lets a token gather information from other tokens with different weights.

注意力让一个 token 以不同权重，从其他 token 收集信息。

**00:23:55** Our simplified scalar example gives weight 0.8 to a value of 0.9, and weight 0.2 to a value of 0.1.

在这个简化的标量例子中，数值 0.9 的权重为 0.8，数值 0.1 的权重为 0.2。

**00:24:05** The weighted result is 0.74.

加权结果是 0.74。

**00:24:08** The real mechanism operates on learned vectors, so this is arithmetic intuition rather than a literal description of artistic decision-making.

真实机制处理的是学习到的向量，因此这里只是在建立算术直觉，并非字面描述艺术决策。

**00:24:17** The important structure is selective combination: every available source need not contribute equally.

关键结构是选择性组合：不是每个可用来源都必须贡献同样多的信息。

**00:24:23** In a multi-reference edit, the system also needs to distinguish where information came from.

在多参考图编辑中，系统还需要区分信息来自哪里。

**00:24:29** Otherwise, gathering information successfully can still produce the wrong mixture of subject identity and visual style.

否则，即使成功收集了信息，也可能错误地混合主体身份与视觉风格。


## 034 — Softmax weights sum to one (00:24:36)

**00:24:36** The weights in this simplified attention example sum to one.

在这个简化注意力例子中，权重之和为一。

**00:24:41** That makes the result a weighted combination of the values.

因此，结果是各个数值的加权组合。

**00:24:45** A larger weight makes the associated value contribute more to this particular calculation.

权重越大，对应数值在这次计算中的贡献就越大。

**00:24:51** Now be cautious about jumping from the calculation to an explanation of the final image.

但要谨慎，不能直接从这个计算跳到对最终图像的解释。

**00:24:57** Real networks contain many layers, heads, and transformations.

真实网络包含许多层、注意力头和变换。

**00:25:01** A high weight at one point may not tell us why a visible feature ultimately appeared.

某处权重很高，未必能解释某个可见特征最终为何出现。

**00:25:06** For practical reference editing, the useful question remains whether the intended information survives in the output, regardless of how compelling an attention visualization looks.

在实际参考图编辑中，关键仍是预期信息是否出现在输出中，无论注意力可视化看起来多么有说服力。


## 035 — LoRA: low-rank adaptation (00:25:18)

**00:25:18** What if a useful concept must recur across many requests?

如果一个有用的概念需要在许多请求中反复出现，怎么办？

**00:25:22** A reference can condition a generation, while LoRA adapts selected learned weights using a low-rank update.

参考图可以为一次生成提供条件；LoRA 则通过低秩更新，调整选定的模型权重。

**00:25:29** The update is a product of two smaller matrices.

更新量是两个较小矩阵的乘积。

**00:25:33** Instead of learning every entry of a large update matrix independently, we express that update as a product of two smaller matrices.

我们不独立学习大型更新矩阵的每个元素，而是用两个较小矩阵的乘积来表示更新。

**00:25:41** For a 4096 by 4096 weight matrix and rank sixteen, the full matrix has 16,777,216 entries.

对于一个 4096 乘 4096、秩设为十六的例子，完整矩阵有 16,777,216 个元素。

**00:25:52** The two factors together have 131,072, a factor of 128 fewer for this update.

两个因子合计只有 131,072 个元素，这次更新的参数量减少了一百二十八倍。

**00:26:01** That is not a claim that the entire model shrinks by 128.

这并不意味着整个模型缩小了一百二十八倍。

**00:26:05** The base weights still exist.

基础权重仍然存在。

**00:26:08** Low rank makes adaptation economical, but does not certify that the learned subject survives a new pose.

低秩让适配更经济，但不能证明学到的主体在新姿态下仍能保持一致。

**00:26:14** That question requires examples the adaptation did not already see.

这个问题需要用适配时未见过的样例来检验。


## 036 — Did it learn the subject or the setting? (00:26:19)

**00:26:19** An adapted model reproduces your favorite training portrait perfectly.

适配后的模型完美复现了你最喜欢的训练肖像。

**00:26:24** Is that enough to use the character in a new film?

这就足以让这个角色出演新电影了吗？

**00:26:27** Consider what else it may have learned: the familiar camera angle, the background, even the lighting that always accompanied the subject.

想一想它还可能学到了什么：熟悉的机位、背景，甚至总与主体一起出现的光照。

**00:26:35** A held-out check deliberately changes those circumstances.

留出测试会有意改变这些条件。

**00:26:39** Ask for a new view or a different setting and inspect the distinctive features.

要求一个新视角或不同环境，并检查独特特征。

**00:26:45** We want a reusable concept, rather than a narrow ability to reproduce familiar combinations.

我们需要的是可复用的概念，而不只是复现熟悉组合的狭窄能力。

**00:26:51** For the pear, keep the blue stem recognizable while moving the exhibition outdoors.

对于这只梨，要在把展览移到户外时，仍保留可辨认的蓝色柄。

**00:26:57** If identity collapses there, more praise for the training examples will not solve the problem.

如果身份特征在那里崩溃，再称赞训练样例也解决不了问题。


## 037 — What does the reward favor? (00:27:03)

**00:27:03** A reward tells a learning process what kinds of outputs to favor.

奖励告诉学习过程应当偏好哪些输出。

**00:27:07** The figure shows a recent approach that separates task-specific reward training and then distills what is learned.

图中是一种近期方法：分别训练任务特定的奖励，再蒸馏学到的能力。

**00:27:14** Look at the categories: editing quality involves more than a single judgment of attractiveness.

看这些类别：编辑质量不只是单一的美观判断。

**00:27:21** Now imagine an artwork whose purpose is to feel awkward or disturbing.

现在想象一件艺术作品，其目的就是让人感到别扭或不安。

**00:27:25** A generic preference for polished images could work against that purpose.

对精致图像的普遍偏好，可能反而违背这个目的。

**00:27:29** Human preference signals are useful, but they do not define artistic merit for every project.

人类偏好信号很有用，但不能为所有项目定义艺术价值。

**00:27:35** This is the bridge to the next section: technical optimization can help produce candidates, while an artist still needs to decide which properties serve the work.

这正是通向下一部分的桥梁：技术优化可以帮助生成候选作品，而艺术家仍要决定哪些属性服务于作品。


## 038 — One reward cannot stand in for every intention (00:27:46)

**00:27:46** These edit examples invite a question about evaluation.

这些编辑示例引出了一个评价问题。

**00:27:50** Which properties would you reward separately?

你会分别奖励哪些属性？

**00:27:53** You might ask whether the instruction was carried out, whether identity survived, and whether the image is visually convincing.

可以问：指令是否完成、身份是否保留，以及图像在视觉上是否可信。

**00:28:01** Those judgments can disagree.

这些判断可能彼此冲突。

**00:28:04** A highly polished output might erase an awkward feature that is central to the artwork.

高度精致的输出，可能抹掉作品中至关重要的别扭特征。

**00:28:09** An unusual composition might serve the brief while attracting a lower generic preference score.

不寻常的构图可能符合创作要求，却得到较低的通用偏好分数。

**00:28:15** We should understand what a reward encourages before treating a high score as artistic approval.

在把高分视为艺术认可之前，我们应当理解奖励究竟鼓励什么。

**00:28:20** The published examples illustrate the research setting; our classroom task is to articulate the intention against which we would judge a particular result.

已发表示例展示的是研究情境；课堂任务则是明确创作意图，并据此判断具体结果。


## 039 — Attractiveness and fidelity are separate tests (00:28:30)

**00:28:30** Which failure would you notice first?

你会先注意到哪种失败？

**00:28:33** On one side, the output is beautiful, but it has quietly replaced our particular sculpture with a generic decorative pear.

一边的结果很美，却悄悄把我们的特定雕塑换成了普通的装饰梨。

**00:28:40** On the other, the right sculpture is present, but its reflection still behaves like the old material.

另一边保留了正确雕塑，但倒影仍像原来的材质。

**00:28:46** The first may win a quick aesthetic vote.

第一种可能赢得快速的审美投票。

**00:28:49** The second may pass a checklist that only asks whether the requested object changed.

第二种可能通过只检查“目标物体是否改变”的清单。

**00:28:53** Neither satisfies the full brief.

两者都没有满足完整要求。

**00:28:56** Has the editor lost identity, broken the scene's optical logic, or missed the work's intention?

编辑器是丢失了身份、破坏了场景光学逻辑，还是偏离了作品意图？

**00:29:02** Naming the mismatch gives us a next action.

说清不匹配之处，才能决定下一步。

**00:29:05** Saying only that an output looks a little wrong leaves the diagnosis unfinished.

只说输出“看起来有点不对”，诊断还没有完成。


## 040 — A useful diagnosis changes the next action (00:29:11)

**00:29:11** A diagnosis becomes useful when it changes the next action.

只有能改变下一步行动的诊断，才有用。

**00:29:15** If the wrong object changes, first suspect an ambiguous reference or region specification.

如果改错了物体，首先怀疑参考对象或区域描述有歧义。

**00:29:22** If the correct object becomes another instance, appearance grounding may be the issue.

如果改对了物体，却变成另一个个体，问题可能在外观信息的约束。

**00:29:28** If the material looks disconnected from its reflection, the dependent change may be missing.

如果材质与倒影脱节，可能缺少必要的连带变化。

**00:29:34** These are candidate explanations, not automatic conclusions.

这些只是候选解释，不是自动成立的结论。

**00:29:38** Choose a small revision that tests one of them.

选择一个小调整，检验其中一种解释。

**00:29:40** If it does not help, reconsider the hypothesis.

如果没有帮助，就重新考虑假设。

**00:29:44** This is more informative than changing the prompt, seed, reference, and model all at once, because then a successful result would leave us unsure which intervention mattered.

这比同时改提示词、随机种子、参考图和模型更有信息量；否则，即使成功，也不知道哪个改动起了作用。


## 041 — Discuss: the ceramic edit (00:29:54)

**00:29:54** We now have several ways to communicate an edit.

现在，我们有了几种表达编辑要求的方法。

**00:29:58** Would you preserve the reflection or let it change, and why?

你会保留倒影，还是允许它变化？为什么？

**00:30:02** When would a mask help more than a longer prompt?

什么时候蒙版比更长的提示词更有帮助？

**00:30:06** What would edges or depth control still leave uncertain?

边缘或深度控制仍会留下哪些不确定性？

**00:30:10** Discuss these three questions using the pear already on screen.

请用屏幕上的这只梨，讨论这三个问题。

**00:30:14** For each control you favor, name the failure it addresses and something it cannot settle.

对于你支持的每种控制手段，指出它解决什么失败，以及它不能决定什么。

**00:30:19** Pause the video here.

请在这里暂停视频。

**00:30:21** The next slide offers one possible contract; compare it with your reasoning rather than treating it as the only artistic answer.

下一页给出一种可能的约定；请与你的推理比较，不要把它当成唯一的艺术答案。


## 042 — A worked ceramic-edit contract (00:30:33)

**00:30:33** Here is one defensible answer to the ceramic puzzle.

对于陶瓷难题，这是一种有理有据的答案。

**00:30:37** Preserve the blue stem and recognizable silhouette.

保留蓝色柄和可辨认的轮廓。

**00:30:40** Allow transmission, highlights, and the reflection to change with the new material.

允许透光、高光和倒影随着新材质改变。

**00:30:47** Then inspect the boundary and water, because that is where the request reaches beyond the object.

然后检查边界和水面，因为要求的影响会在那里超出物体本身。

**00:30:53** Notice the wording: allow a justified consequence, rather than permit arbitrary drift.

注意措辞：允许有理由的连带变化，而不是允许任意漂移。

**00:30:58** We have not given the model permission to redesign the gallery.

我们没有授权模型重新设计画廊。

**00:31:02** We have named the changes necessary to make this material transformation coherent.

我们明确指出了，让这次材质变换协调一致所必需的变化。

**00:31:07** A mask could help localize work, while compositing might protect a region that truly must remain exact.

蒙版可以帮助限定工作区域；合成则可以保护真正必须完全不变的区域。

**00:31:15** The contract tells us how to use those tools.

编辑约定告诉我们如何使用这些工具。


## 043 — The brief, the model, and the review (00:31:18)

**00:31:18** We can now explain our opening puzzle at three levels.

现在，我们可以在三个层面解释开场难题。

**00:31:22** The brief decides what should change and what should survive.

创作说明决定什么应改变、什么应保留。

**00:31:26** The model uses representations, conditioning, and sampling to propose a result.

模型通过表示、条件控制和采样，提出候选结果。

**00:31:31** Our review checks whether that proposal actually satisfies the brief.

我们的审查检验这个候选结果是否真正满足要求。

**00:31:36** Keep those levels separate when you diagnose a failure.

诊断失败时，要区分这三个层面。

**00:31:39** Stronger guidance will not choose the exhibition's purpose.

更强的引导不会替你选择展览的目的。

**00:31:43** A more eloquent artistic statement will not enforce identical pixels.

更动人的艺术陈述不会强制像素完全一致。

**00:31:48** A beautiful preview will not prove preservation.

精美预览不会证明保留成功。

**00:31:51** The useful skill is connecting the right intervention to the observed problem.

有用的能力，是把正确的干预对应到观察到的问题。

**00:31:56** If your partner can tell when to use it and what it leaves uncertain, you have understood more than its name.

如果同伴能说清何时使用它、它还留下什么不确定性，你就不只是记住了它的名字。


## 044 — After Part 1: explain the control (00:32:03)

**00:32:03** We can now explain our controls rather than just name them.

现在，我们能够解释控制手段，而不只是说出名称。

**00:32:06** How does ControlNet differ from a mask or a text prompt?

ControlNet 与蒙版或文本提示词有什么区别？

**00:32:10** Why can stronger guidance make an edit worse?

为什么更强的引导可能让编辑变差？

**00:32:14** What evidence would show that an edit preserved identity?

什么证据能够表明编辑保留了身份特征？

**00:32:18** Discuss these three questions.

请讨论这三个问题。

**00:32:20** Choose a concrete requirement—a silhouette, an untouched region, or a distinctive stem—and connect your explanation to it.

选一个具体要求——轮廓、未编辑区域，或独特的梨柄——并把解释与它联系起来。

**00:32:27** Then consider whether the control guarantees that requirement or only helps express it.

然后考虑：这种控制是保证实现要求，还是仅仅帮助表达要求？

**00:32:32** Pause the video here before the break.

休息之前，请在这里暂停视频。


## 045 — Break (00:32:39)

**00:32:39** We will take ten minutes and resume when the class is ready.

我们休息十分钟，等大家准备好再继续。

**00:32:43** When you come back, think about a series of images you would recognize as belonging to one artwork.

回来时，请想一组你能认出属于同一件艺术作品的图像。

**00:32:48** What makes them belong together: the subject, the palette, the treatment of space, or something else?

是什么让它们属于一体：主体、色彩、空间处理，还是别的因素？

**00:32:56** Leave that question open for now.

暂时先保留这个问题。

**00:32:59** We will use it to move from individual editing operations to a coherent visual language.

我们将用它，从单次编辑操作转向连贯的视觉语言。


## 046 — One exhibition, different visual worlds (00:33:08)

**00:33:08** Look again at the artwork we have been using as a technical test.

再看看这件一直被我们用来做技术测试的作品。

**00:33:12** After the break, it has a different job.

休息之后，它有了不同的任务。

**00:33:15** We are no longer asking only whether the edit obeys a request.

我们不再只问编辑是否遵循了要求。

**00:33:19** We are asking what a visitor might feel, and which visual decisions create that feeling.

我们要问：观众可能感受到什么，哪些视觉决定创造了这种感受？

**00:33:26** Two images can have different surfaces and still belong to one exhibition.

两张图像的表面可以不同，却仍属于同一个展览。

**00:33:30** Two images can share a palette and feel unrelated.

两张图像可以使用相同色彩，却感觉毫不相关。

**00:33:33** The difference is worth arguing about.

这种差别值得讨论。

**00:33:36** In the next section, I will show prepared alternatives for AFTER RAIN.

下一部分，我会展示为《雨后》准备的不同方案。

**00:33:41** Choose a direction in your mind, and be ready to explain a visible reason.

请在心里选择一个方向，并准备说出可见的理由。

**00:33:46** Your neighbor may choose the other one.

你的邻座可能会选另一个。


## 047 — Art direction (00:33:49)

**00:33:49** Now take the art director’s seat.

现在，请坐到艺术总监的位置。

**00:33:51** The model can produce many plausible outputs; we decide which differences matter for AFTER RAIN.

模型能生成许多合理的结果；我们来决定哪些差异对《雨后》重要。

**00:33:57** As we compare the prepared images, name the visible relationship that carries the idea.

比较这些准备好的图像时，请指出承载创意的可见关系。

**00:34:03** That could be scale, light, or the treatment of space.

它可能是尺度、光线，或空间处理。

**00:34:09** Your explanation will be more useful than simply calling an image impressive.

这样的解释，比简单地说图像很惊艳更有用。


## 048 — Before Part 2: artistic intention (00:34:13)

**00:34:13** Before seeing the poster alternatives, consider what you want a visitor to experience.

看海报方案之前，先想想你希望观众经历什么。

**00:34:18** What makes an image feel fragile rather than merely attractive?

什么让图像显得脆弱，而不仅仅是好看？

**00:34:22** Can very different images belong to the same artwork, and why?

差别很大的图像，能否属于同一件作品？为什么？

**00:34:27** Which artistic decisions would you keep for yourself?

哪些艺术决定，你会留给自己？

**00:34:31** Discuss these three questions.

请讨论这三个问题。

**00:34:33** You might disagree about the effect of empty space or the meaning of a material.

你们可能对留白的效果或材质的含义有不同看法。

**00:34:37** Locate the visual evidence behind that disagreement.

请找出分歧背后的视觉证据。

**00:34:41** Pause the video here.

请在这里暂停视频。

**00:34:43** Keep your starting position in mind when we compare the prepared directions.

比较预先准备的方向时，请记住你的初始立场。


## 049 — The exhibition brief (00:34:51)

**00:34:51** Here is our commission: AFTER RAIN, a fictional exhibition about fragile objects in changing environments.

这是我们的委托：《雨后》，一个关于变化环境中脆弱物体的虚构展览。

**00:35:00** Imagine a visitor seeing its poster across a corridor before encountering a five-second moving image inside.

想象观众先在走廊另一端看到海报，再进入展厅看到一段五秒的动态图像。

**00:35:07** What should that visitor expect to feel?

你希望这位观众期待怎样的感受？

**00:35:10** The prepared examples keep the amber pear and blue stem recognizable.

准备好的示例都保留了可辨认的琥珀色梨和蓝色柄。

**00:35:14** We will examine photographic and collage directions, then consider how a moving reveal changes the experience.

我们将研究摄影和拼贴两个方向，再考虑运动中的揭示如何改变体验。

**00:35:22** Nobody needs to generate an image during this lecture.

这节课不需要任何人现场生成图像。

**00:35:25** Fragile is a useful beginning, but it is not yet an art direction.

“脆弱”是一个有用的起点，但还不算艺术方向。

**00:35:30** We need to translate that word into something visible enough to compare and precise enough to revise.

我们需要把这个词转化为足够可见、能够比较，又足够精确、能够修改的东西。


## 050 — Making fragility visible (00:35:36)

**00:35:36** Try replacing fragile with expensive in the brief.

试着把创作说明中的“脆弱”换成“昂贵”。

**00:35:40** You might still choose glass, but would you use the same composition?

你可能仍然选择玻璃，但还会用同样的构图吗？

**00:35:44** A large centered object and assertive lighting might suggest a luxury product.

居中放大的物体配合强势光照，可能让人联想到奢侈品。

**00:35:49** A small object surrounded by quiet space might instead seem exposed.

安静空间中一个小小的物体，则可能显得无所庇护。

**00:35:54** Scale can make the pear feel vulnerable.

尺度可以让梨显得易受伤害。

**00:35:57** Restrained light can make us look closely.

克制的光线可以引导我们细看。

**00:35:59** A reflection can suggest a world less stable than the object itself.

倒影可以暗示一个比物体本身更不稳定的世界。

**00:36:04** These are artistic hypotheses to test against an audience's reading, not a formula that makes every image fragile.

这些是需要通过观众解读来检验的艺术假设，不是让所有图像显得脆弱的公式。

**00:36:12** Point to the feature doing the work.

请指出真正起作用的特征。

**00:36:15** If nobody can locate it in the image, the intention may still be living only in the prompt.

如果没人能在图像中找到它，意图可能仍然只存在于提示词里。


## 051 — Hokusai: scale, rhythm, and tension (00:36:21)

**00:36:21** In Hokusai's Great Wave, look first for Mount Fuji.

看葛饰北斋的《神奈川冲浪里》，先找富士山。

**00:36:24** It is small and distant.

它很小，也很远。

**00:36:26** Now follow the curves of the wave and the boats beneath it.

现在沿着巨浪的曲线，以及下方的小船看。

**00:36:30** Scale and rhythm create a tension we can discuss without turning the artwork into a style label.

尺度和节奏创造了张力；我们可以讨论这种张力，而不是把作品变成一个风格标签。

**00:36:36** For AFTER RAIN, we might borrow a relationship: a small, vulnerable form facing a much larger environment.

对《雨后》来说，我们可以借鉴一种关系：一个小而脆弱的形体，面对远大于自身的环境。

**00:36:43** We do not need to reproduce the wave or ask for a generic imitation of the artist.

不需要复制巨浪，也不需要要求泛泛地模仿这位艺术家。

**00:36:48** The reference becomes useful when we can say which visual decision we are studying.

当我们能够说明正在研究哪个视觉决定时，参考作品才真正有用。

**00:36:53** Then we must test its translation.

然后，还必须检验这种转化。

**00:36:55** Our quiet flooded gallery has a different subject and emotional register.

我们安静而积水的画廊，有着不同的主体和情感基调。


## 052 — Reference as a relationship, not a label (00:37:00)

**00:37:00** Instead of using a reference as a label, describe a relationship you can observe.

不要把参考作品当作标签，而要描述你能观察到的关系。

**00:37:06** You might want a small stable form beneath a large dynamic curve, or a rhythm that moves the eye through the composition.

你可能想要一个巨大动态曲线下的小型稳定形体，或引导视线穿过构图的节奏。

**00:37:14** Then translate that relationship into the new subject and brief.

然后，把这种关系转化到新的主体和创作要求中。

**00:37:18** The result should be judged as its own artwork, not as a contest to resemble the reference.

应把结果作为独立作品评价，而不是比赛谁更像参考图。

**00:37:23** This approach makes references more useful to collaborators because they can understand what you are borrowing.

这种方法让参考资料对合作者更有用，因为他们能理解你借鉴的是什么。

**00:37:29** It also makes iteration more focused: if the intended tension is missing, you can revise scale or rhythm rather than vaguely asking for more influence from the source.

它也让迭代更聚焦：如果缺少预期张力，可以调整尺度或节奏，而不是含糊地要求“更受原作影响”。


## 053 — Where does your eye arrive? (00:37:41)

**00:37:41** Before admiring the surface detail, decide where your eye arrives.

欣赏表面细节之前，先判断你的视线最先落在哪里。

**00:37:45** Is it the pear, the reflection, or the space waiting above them?

是梨、倒影，还是它们上方留出的空间？

**00:37:50** A poster has to organize that first encounter.

海报必须组织这第一次相遇。

**00:37:54** If every region is equally busy, the audience has to invent a hierarchy we have not provided.

如果每个区域都同样繁忙，观众就得自行建立我们没有提供的层次。

**00:38:00** Here, the open area can make the object feel small and give the title somewhere to live.

这里的开放区域能让物体显得小，也给标题留下位置。

**00:38:05** We can ruin the sense of quiet by filling every available gap with decorative detail, even if each addition is attractive on its own.

如果用装饰细节填满每一处空隙，就可能破坏安静感，即使每个新增细节单独看都很漂亮。

**00:38:13** When revising this image, I would first ask whether the composition expresses the intended scale and silence.

修改这张图时，我会先问：构图是否表达了预期的尺度感与寂静？

**00:38:20** More texture comes later, if the work needs it.

如果作品需要，之后再增加纹理。


## 054 — Negative space has a job (00:38:24)

**00:38:24** Negative space does two jobs in this poster.

留白在这张海报中承担两项任务。

**00:38:27** It changes how large or isolated the object feels, and it creates a place for typography.

它改变物体显得多大、多孤立，也为文字排版创造位置。

**00:38:34** These are connected design decisions rather than separate finishing steps.

这是相互关联的设计决定，而不是彼此独立的收尾步骤。

**00:38:39** If you fill the pale area with dramatic detail, the image may become more visually active but leave no calm route for reading the title.

如果在浅色区域填满戏剧性细节，画面可能更活跃，却没有平静的路径让人阅读标题。

**00:38:47** If you reserve too much space, the subject might lose the presence you wanted.

如果留白过多，主体又可能失去你想要的存在感。

**00:38:51** Evaluate the space in relation to the finished use.

请结合最终用途来评价空间。

**00:38:55** A generative background is part of a larger composition when it must carry text, branding, or other exact elements.

当背景需要承载文字、品牌或其他精确元素时，生成背景只是更大构图的一部分。


## 055 — A visual language has rules (00:39:03)

**00:39:03** The collage direction changes the rules of the world.

拼贴方向改变了这个世界的规则。

**00:39:06** Torn edges replace smooth contours.

撕裂边缘取代了平滑轮廓。

**00:39:09** Flat layers replace optical depth.

平面层次取代了光学纵深。

**00:39:12** Amber and indigo keep a connection to our recurring subject.

琥珀色和靛蓝色，让它仍与反复出现的主体保持联系。

**00:39:17** Look at how those decisions affect the pear's apparent weight and vulnerability.

看看这些决定如何影响梨看起来的重量与脆弱感。

**00:39:22** If the body looks like paper but the water behaves like a photograph, do we accept that collision?

如果主体像纸，水却像照片，我们是否接受这种碰撞？

**00:39:28** We might, if it is deliberate and supports the work.

如果它是有意的，并服务于作品，我们可能接受。

**00:39:31** We might reject it as an unresolved mixture.

我们也可能认为它是尚未解决的混杂，因而拒绝。

**00:39:35** The answer depends on the intended visual language.

答案取决于预期的视觉语言。

**00:39:38** The next comparison asks you to locate precisely where that language holds together or breaks.

下一组比较，请你精确指出：这种语言在哪里成立，又在哪里断裂。


## 056 — When a reflection breaks the collage (00:39:44)

**00:39:44** A photographic reflection inside a paper collage can be a mistake, or the most interesting decision in the image.

纸质拼贴中的摄影式倒影，可能是错误，也可能是整张图最有意思的决定。

**00:39:51** We need to know whether the mismatch is an intentional disruption and whether it produces the intended effect.

我们需要知道：这种不匹配是否是有意的打破，以及是否产生了预期效果。

**00:39:59** Suppose the exhibition explores unstable memories.

假设展览探索的是不稳定的记忆。

**00:40:02** An impossibly photographic reflection could make sense.

一个不可能的摄影式倒影，就可能说得通。

**00:40:06** Suppose the brief calls for a coherent world assembled from torn paper.

假设要求是构建一个由撕纸组成的统一世界。

**00:40:10** The same reflection might weaken it.

同样的倒影可能削弱它。

**00:40:13** This does not mean every accident deserves an explanation after the fact.

这并不意味着每个意外都值得事后找理由。

**00:40:18** Make the intention specific, examine what the viewer can actually see, and compare alternatives.

请明确意图，检查观众实际能看见什么，并比较不同方案。

**00:40:25** A strong critique can distinguish a productive contradiction from an excuse for an unresolved result.

有力的评论能区分富有成效的矛盾，与为未完成结果寻找的借口。


## 057 — Reference roles (00:40:32)

**00:40:32** A reference becomes easier to use when we assign it a role.

为参考图分配角色后，它就更容易使用。

**00:40:36** One image might define the subject, another the composition, and another the material treatment.

一张图可以定义主体，另一张定义构图，再一张定义材质处理。

**00:40:43** Name the property you want from each.

请说清你想从每张图中得到什么属性。

**00:40:45** The model may not isolate those roles perfectly.

模型未必能完美隔离这些角色。

**00:40:49** That is why you should look for unwanted transfers, such as importing a reference's background when you only wanted its palette.

因此，要检查意外的信息转移，例如只想借用配色，却把参考图的背景也引入了。

**00:40:56** Try removing one reference and observing what disappears.

试着移除一张参考图，观察什么随之消失。

**00:41:00** Start with the smallest useful set.

从最小但有用的参考集合开始。

**00:41:03** If five references contradict one another, adding a sixth may make the problem harder to diagnose rather than solve it.

如果五张参考图互相矛盾，再加第六张可能让问题更难诊断，而不是解决问题。


## 058 — What did the reference contribute? (00:41:11)

**00:41:11** We have added a reference because we think it helps.

我们添加参考图，是因为认为它有帮助。

**00:41:14** How would we discover whether it actually does?

怎样才能知道它是否真的有帮助？

**00:41:16** Remove it and compare.

移除它，再比较。

**00:41:19** First state its intended role: perhaps the pear's silhouette, perhaps the composition's sense of scale.

先说明它的预期作用：可能是梨的轮廓，也可能是构图的尺度感。

**00:41:26** Then look for that quality in outputs with and without the reference.

然后，在有无这张参考图的输出中寻找这种品质。

**00:41:30** Also inspect what came along uninvited.

也要检查哪些信息不请自来。

**00:41:33** A useful subject reference can carry a background or lighting scheme we did not want.

有用的主体参考，也可能携带我们不想要的背景或光照。

**00:41:38** This is an ablation: change one component to understand its contribution.

这就是消融实验：改变一个组件，以理解它的贡献。

**00:41:44** Random generation makes one lucky comparison weak evidence, so several samples can help.

随机生成会让一次幸运的对比缺乏说服力，因此多个样本会有帮助。

**00:41:51** If we cannot explain what a reference contributes, we may be making the request more complicated without making the direction clearer.

如果说不清一张参考图的贡献，我们可能只是把要求变复杂，却没有让方向更清楚。


## 059 — An art-direction prompt (00:41:59)

**00:41:59** Listen to how this prompt distributes responsibility.

听听这个提示词如何分配职责。

**00:42:03** The pear reference supplies silhouette and the blue stem.

梨的参考图提供轮廓和蓝色柄。

**00:42:06** Placement low on the right organizes the composition.

将主体放在右下方，用来组织构图。

**00:42:10** A pale upper-left area reserves space for the title.

左上方的浅色区域，为标题留出空间。

**00:42:14** Soft daylight and restrained reflections support the mood.

柔和日光与克制的倒影，支撑整体情绪。

**00:42:17** Every phrase points toward something we could inspect in an output.

每个短语，都指向输出中可以检查的东西。

**00:42:22** Compare that with asking for a stunning masterpiece with beautiful lighting.

与“用漂亮光照创作一幅惊艳杰作”这样的要求比较一下。

**00:42:26** This prompt is still a proposal, not a binding contract enforced by the model.

这个提示词仍然只是提议，不是模型强制执行的约束协议。

**00:42:31** If an output fails, we can now name the failed relationship: the title area is crowded, the light is too assertive, or the reference's background has leaked into the scene.

输出失败时，我们现在可以指出哪种关系失败了：标题区太拥挤、光线太强势，或参考图背景渗入了场景。


## 060 — Group Editing: one decision, many views (00:42:43)

**00:42:43** A single poster can look coherent by itself.

单张海报本身可能看起来很统一。

**00:42:47** A series exposes a harder problem: does the same material decision survive across viewpoints?

系列作品暴露出更难的问题：同样的材质决定，能否在不同视角下保持？

**00:42:53** Our Group Editing collaboration studies related images that should be edited consistently.

我们合作开展的 Group Editing 研究，关注应当保持一致编辑的关联图像。

**00:42:58** Treating each image independently can produce slightly different costumes or materials.

独立处理每张图，可能生成略有不同的服装或材质。

**00:43:03** The method arranges related images as pseudo-video frames to use a video model's consistency prior.

该方法把相关图像排列成伪视频帧，利用视频模型的一致性先验。

**00:43:10** VGGT provides geometric correspondences.

VGGT 提供几何对应关系。

**00:43:14** Geometry-enhanced rotary positional embeddings connect geometry features with image latents, while Identity-RoPE supports identity preservation.

几何增强的旋转位置编码，将几何特征与图像潜变量联系起来；Identity-RoPE 则支持身份保持。

**00:43:24** Follow the penguin examples through the figure before reading every label.

阅读每个标签之前，先沿着图中的企鹅示例看。

**00:43:28** Reliable correspondence is part of the technical problem: when views cannot be matched well, repeating the same instruction alone does not establish a coherent series.

可靠对应关系是技术问题的一部分：视角无法准确匹配时，仅重复相同指令，并不能保证系列统一。


## 061 — Iteration as a controlled comparison (00:43:40)

**00:43:40** Iteration becomes informative when we know what changed between attempts.

当我们知道两次尝试之间改了什么，迭代才有信息量。

**00:43:44** Suppose we compare two lighting treatments.

假设我们比较两种灯光处理。

**00:43:47** Keep the subject and composition instructions stable, and record the references and model version.

保持主体和构图指令不变，并记录参考图和模型版本。

**00:43:54** If a seed control exists, holding it fixed can help the first comparison.

如果可以控制随机种子，固定种子有助于初步比较。

**00:43:59** A shared seed still does not guarantee identical composition after a prompt change.

但即使种子相同，改变提示词后，也不能保证构图完全一致。

**00:44:04** If results vary widely, examine several outputs for each condition.

如果结果差异很大，就检查每种条件下的多个输出。

**00:44:09** Otherwise, one lucky sample may decide the whole direction.

否则，一个幸运样本就可能决定整个方向。

**00:44:13** The table is a proposed experiment, not measured results.

这张表是拟议实验，不是实测结果。

**00:44:17** Its purpose is to make each generation answer a question about the artwork.

它的目的，是让每次生成都回答一个关于作品的问题。


## 062 — A lucky image or a reliable direction? (00:44:21)

**00:44:21** Imagine you compare two prompts once, and the second gives a wonderful poster.

想象你只比较了两条提示词各一次，第二条生成了精彩海报。

**00:44:28** Did its wording cause the improvement, or did it receive a favorable random sample?

改进是措辞造成的，还是它碰巧抽到了有利的随机样本？

**00:44:33** From one output, it can be difficult to tell.

只看一个输出，往往很难判断。

**00:44:36** Several outputs per condition reveal whether the direction is stable or merely fortunate.

每种条件生成多个结果，才能看出这个方向是稳定有效，还是仅仅幸运。

**00:44:42** For art direction, we can ask two different questions.

在艺术指导中，我们可以问两个不同的问题。

**00:44:45** Would I exhibit this particular image?

我愿意展出这张具体图像吗？

**00:44:48** Could I reliably develop a series using this direction?

我能用这个方向，可靠地发展出一个系列吗？

**00:44:52** One exceptional output may answer the first while leaving the second unresolved.

一张出色的输出可能回答了第一个问题，却没有解决第二个。

**00:44:57** Repeating generations is helpful when it reveals a pattern relevant to the brief; it becomes a distraction when we keep browsing alternatives to avoid deciding what we value.

重复生成若能揭示与创作要求相关的规律，就有帮助；若只是为了逃避价值判断而不断浏览方案，就成了干扰。


## 063 — Typography is part of the image (00:45:08)

**00:45:08** Now the image becomes a poster.

现在，图像变成了海报。

**00:45:11** The title is an editable typographic layer, which lets us choose its wording, line breaks, and placement exactly.

标题是可编辑的文字图层，让我们精确选择措辞、换行和位置。

**00:45:20** That matters when a work has to carry an exhibition name rather than merely resemble a poster in a generated preview.

作品需要准确呈现展览名称时，这很重要，而不是只在生成预览中“看起来像海报”。

**00:45:28** Look at how the title occupies the area we deliberately left open.

看看标题如何占据我们有意留出的区域。

**00:45:32** The image and type were designed to cooperate.

图像与文字是协同设计的。

**00:45:35** We can adjust hierarchy without asking the model to regenerate the sculpture and risk changing it.

我们可以调整层次，而不必让模型重新生成雕塑、冒着改变它的风险。

**00:45:41** This is a useful division of labor: generation develops the visual material, and direct layout controls exact communication.

这是一种有用的分工：生成负责发展视觉素材，直接排版负责精确传达信息。

**00:45:49** The audience sees one composition.

观众看到的是一个整体构图。

**00:45:52** It does not need to know which parts came from which tool.

他们不需要知道哪个部分来自哪种工具。


## 064 — Typography needs a reading order (00:45:56)

**00:45:56** Typography establishes a sequence of attention.

文字排版建立了注意力的顺序。

**00:45:59** Ask what the viewer should read first, what comes next, and where their eye returns to the artwork.

要问：观众应先读什么、接着读什么，以及视线在哪里回到作品？

**00:46:05** In our poster, reserved space allows the title to be clear without covering the sculpture.

在我们的海报中，预留空间让标题清楚可读，又不会遮住雕塑。

**00:46:10** This is also a production decision.

这也是一个制作决定。

**00:46:13** Exact text can remain editable, while the image carries the material and atmosphere.

精确文字可以保持可编辑，而图像承载材质和氛围。

**00:46:18** If the title feels too dominant, change size, placement, or contrast and inspect the whole composition again.

如果标题太抢眼，就调整大小、位置或对比度，再检查整个构图。

**00:46:25** Do not judge the text in isolation from the image.

不要脱离图像，单独评价文字。

**00:46:29** The poster is the relationship between them, including the empty space that lets each element do its job.

海报是两者之间的关系，也包括让各元素发挥作用的空白。


## 065 — Two directions for the same brief (00:46:35)

**00:46:35** Take a few seconds with both directions before I describe them.

在我描述之前，先花几秒看看两个方向。

**00:46:40** Which would you put outside AFTER RAIN?

你会把哪一个放在《雨后》展厅外？

**00:46:42** Choose privately first, so your answer is not just a response to mine.

先自己选择，避免答案只是对我的回应。

**00:46:47** Then identify the visible feature that made you choose.

然后指出促使你选择的可见特征。

**00:46:51** The photographic direction can invite attention to material and atmosphere.

摄影方向可以引导观众关注材质与氛围。

**00:46:56** The collage direction can make fragility feel constructed through paper edges and layers.

拼贴方向可以通过纸边和层次，让脆弱感显得被建构出来。

**00:47:01** Both can serve the brief, but they make different promises to a visitor.

两者都可能符合要求，但向观众许下的承诺不同。

**00:47:06** A useful defense goes beyond realism versus abstraction.

有用的辩护，不应止于写实与抽象的区别。

**00:47:10** Tell us what the audience is likely to notice or feel, and how the composition produces that reading.

请告诉我们，观众可能注意或感受到什么，以及构图如何产生这种解读。

**00:47:17** We will use the disagreement to decide what to revise, rather than vote for a universally better image.

我们将用分歧来决定修改方向，而不是投票选出普遍更好的图像。


## 066 — Why would you exhibit this one? (00:47:24)

**00:47:24** Selection is an artistic decision.

选择本身就是艺术决定。

**00:47:27** Once generation gives us many plausible alternatives, choosing one determines what the work becomes.

当生成提供许多合理方案时，选中哪一个，就决定作品最终成为什么。

**00:47:34** The reason should connect to the brief: perhaps the small object feels more exposed, or the torn edge makes fragility physically legible.

理由应与创作要求相连：可能是小物体更显得无所庇护，或撕裂边缘让脆弱感变得可触可见。

**00:47:43** The rejected direction is useful evidence too.

被放弃的方向也是有用证据。

**00:47:47** Explain what it does well and why you are choosing something else for this exhibition.

请解释它哪里做得好，以及为什么为这个展览选择了别的方案。

**00:47:52** An unexpected model output can also change the direction, if you choose to develop it deliberately.

如果你决定有意识地发展一个意外输出，它也可以改变创作方向。

**00:47:58** The key question is whether you can now articulate the intention and carry it through subsequent decisions.

关键是：你现在能否说清意图，并在后续决定中贯彻它？

**00:48:05** Surprise can begin a work; it does not finish the artist's judgment.

惊喜可以开启作品，却不能替艺术家完成判断。


## 067 — A critique with observable evidence (00:48:11)

**00:48:11** A useful critique names a visible feature, explains its effect, and proposes a revision.

有用的评论会指出可见特征、解释效果，并提出修改建议。

**00:48:16** For example: the crisp reflection makes the collage feel photographic, so I would simplify its shape to match the flatter layers.

例如：清晰倒影让拼贴显得像摄影，所以我会简化它的形状，以配合更平面的层次。

**00:48:24** Compare that with 'I do not like the reflection.' The first comment gives the artist a relationship to inspect and a possible next move.

与“我不喜欢倒影”相比，前一种评论给艺术家提供了可检查的关系，以及可能的下一步。

**00:48:33** You can disagree about the intended effect, but make the disagreement specific.

你们可以对预期效果有不同看法，但要把分歧说具体。

**00:48:39** This is a classroom critique framework rather than an official grading rubric.

这是课堂评论框架，不是正式评分标准。

**00:48:44** Use it to connect evidence in the image to the idea the work is trying to communicate.

用它把图像中的证据，与作品试图传达的观念联系起来。


## 068 — LightCtrl: lighting as an artistic choice (00:48:49)

**00:48:49** A curator asks, can the same sculpture feel more vulnerable without changing its shape?

策展人问：能否不改变形状，却让同一件雕塑显得更脆弱？

**00:48:55** Lighting is one way to answer.

灯光是一种回答方式。

**00:48:58** Our LightCtrl collaboration studies controllable relighting from a single image, including light direction, intensity, and color temperature.

我们合作开展的 LightCtrl 研究，探索从单张图像进行可控重光照，包括光线方向、强度和色温。

**00:49:07** Follow the chair through the figure.

请沿着图中的椅子示例看。

**00:49:09** The latent proxy encoder extracts compact physical cues.

潜在代理编码器提取紧凑的物理线索。

**00:49:13** A lighting-aware mask guides the denoiser toward regions affected by the change, and preference optimization in the proxy branch supports physical consistency.

光照感知蒙版引导去噪器关注受变化影响的区域，代理分支中的偏好优化则支持物理一致性。

**00:49:24** The method connects an interpretable lighting request with image generation.

这种方法把可解释的光照要求与图像生成连接起来。

**00:49:29** For our pear, look for consequences rather than the word dramatic: where do highlights move, what happens to shadow, and how does the material read?

对于梨，不要只看“戏剧性”这个词，而要看后果：高光移向哪里、阴影如何变化、材质呈现如何？

**00:49:40** Inferring geometry and material from one photograph is ambiguous, so the result still needs inspection.

从单张照片推断几何和材质存在歧义，因此仍需检查结果。

**00:49:47** The exhibition example is our application of the idea, not an additional result reported by the paper.

展览示例是我们对这一思路的应用，不是论文另外报告的实验结果。


## 069 — Discuss: two artistic directions (00:49:54)

**00:49:54** Choose between the two prepared directions for AFTER RAIN.

请在为《雨后》准备的两个方向中做出选择。

**00:49:58** Which poster communicates fragility more clearly, and why?

哪张海报更清楚地传达脆弱感？为什么？

**00:50:02** What does the rejected direction reveal about your choice?

被放弃的方向，揭示了你选择中的什么？

**00:50:06** Which single revision would most change the audience’s reading?

哪一个单独的修改，会最大程度改变观众的解读？

**00:50:10** Discuss these three questions with a specific visual feature in view.

请看着一个具体视觉特征，讨论这三个问题。

**00:50:15** You can prefer quiet space or the instability of torn paper, but explain how that choice serves the exhibition.

你可以偏爱安静空间，也可以偏爱撕纸的不稳定感，但要解释这种选择如何服务于展览。

**00:50:22** Pause the video here.

请在这里暂停视频。

**00:50:23** Listen for a persuasive reason to choose the direction you initially rejected.

试着听取一个有说服力的理由，支持你最初拒绝的方向。


## 070 — A process record makes critique more precise (00:50:32)

**00:50:32** Consider what a process record lets us discuss.

想一想，过程记录能让我们讨论什么。

**00:50:35** An artist can record the intention and invariants, the prompt and reference roles, and one observation about a rejected result.

艺术家可以记录意图与不变量、提示词与参考图角色，以及对一个被拒绝结果的观察。

**00:50:44** That record explains what an attempt was testing.

这些记录说明了每次尝试在检验什么。

**00:50:48** Without that context, a folder of attractive images can be difficult to interpret.

缺少背景信息，一文件夹好看的图像也可能很难解读。

**00:50:53** We may not know which input changed or why a version was rejected.

我们可能不知道改了哪个输入，或为什么拒绝某个版本。

**00:50:57** A few precise notes make a comparison more informative.

几条精确笔记，就能让比较更有信息量。

**00:51:01** In a critique, this lets us ask about the artist’s decisions and the evidence behind them rather than guessing only from the final image.

在评论中，这让我们能够追问艺术家的决定及其证据，而不只是从最终图像猜测。

**00:51:09** It also helps distinguish an intentional departure from an accidental change.

它也有助于区分有意偏离与意外变化。


## 071 — A gallery conversation (00:51:14)

**00:51:14** Imagine we are standing between the two finished posters in a gallery review.

想象我们站在画廊评审现场，面前是两张完成的海报。

**00:51:19** I would begin with a visible decision that carries the idea, then locate an unintended change, then propose one revision worth discussing.

我会先指出一个承载观念的可见决定，再找出一个意外变化，最后提出一个值得讨论的修改。

**00:51:29** The order matters: the critique begins by understanding the work before prescribing a repair.

顺序很重要：评论应先理解作品，再给出修正方案。

**00:51:35** We can hear two contrasting readings of the same poster.

对同一张海报，我们可能听到两种相反的解读。

**00:51:38** One person may find the empty space quiet; another may find it emotionally distant.

一个人觉得留白安静，另一个人却觉得它情感疏离。

**00:51:44** Ask which details support each reading.

请问，哪些细节支持各自的解读？

**00:51:46** There is no image-making task here.

这里没有图像制作任务。

**00:51:49** We are practicing how to make feedback useful to the next decision.

我们是在练习如何让反馈有助于下一次决定。

**00:51:54** A good proposal names what would change and what effect we expect, so that a later version could confirm or challenge the reasoning.

好建议会说明改什么、预期产生什么效果，让后续版本能够证实或挑战这套推理。

**00:52:02** Discuss these three questions, then pause the video for the gallery conversation.

请讨论这三个问题，然后暂停视频，进行画廊对话。


## 072 — After Part 2: defend a decision (00:52:11)

**00:52:11** Before the artwork begins to move, defend a decision about the images.

在作品开始运动之前，请为一个图像决定提供理由。

**00:52:17** How can a technically correct image fail as an artwork?

技术上正确的图像，为什么可能在艺术上失败？

**00:52:20** Which visual rule should remain across a series?

一个系列中，哪条视觉规则应该保持？

**00:52:24** What evidence makes a critique useful for the next revision?

什么证据能让评论对下一轮修改有用？

**00:52:28** Discuss these three questions through one of the prepared posters.

请结合其中一张准备好的海报，讨论这三个问题。

**00:52:31** Make words such as coherent or expressive concrete: locate the feature, describe its effect, and explain what changing it would do.

把“统一”“有表现力”等词说具体：指出特征、描述效果，并解释改变它会怎样。

**00:52:41** Pause the video here.

请在这里暂停视频。

**00:52:43** We will carry those artistic rules into the video section.

我们会把这些艺术规则带入视频部分。


## 073 — Generative video editing (00:52:50)

**00:52:50** Our image now has to survive time.

现在，图像必须经受时间的考验。

**00:52:53** The subject can move, disappear, and return.

主体会运动、消失，然后返回。

**00:52:57** We will examine selected frames from published research, follow the technical representations, and use a prepared AFTER RAIN scenario to decide what a convincing edit requires.

我们将查看已发表研究的精选帧、理解技术表示，并用准备好的《雨后》情境判断可信编辑需要什么。

**00:53:09** Keep the preservation contract, but add an event that could break it.

保留编辑约定，再加入一个可能让它失效的事件。


## 074 — Before Topic 3: images in motion (00:53:14)

**00:53:14** A moving image creates new ways to break our contract.

动态图像会以新的方式破坏约定。

**00:53:17** What new failures become possible when an image moves?

图像运动起来后，会出现哪些新的失败？

**00:53:21** What should happen when the pear disappears behind a column?

梨消失在柱子后面时，应该发生什么？

**00:53:25** Can four convincing frames prove that a video works?

四张可信的帧，能证明整段视频成立吗？

**00:53:28** Discuss these three questions.

请讨论这三个问题。

**00:53:31** Separate a correct change in visibility from an unwanted change in identity.

区分正确的可见性变化，与不希望发生的身份变化。

**00:53:36** Name a moment you would need to see between the selected frames before accepting the shot.

接受这个镜头之前，请指出你需要在这些精选帧之间看到的某个时刻。

**00:53:41** Pause the video here.

请在这里暂停视频。

**00:53:43** Your prediction will give us a concrete test for the methods that follow.

你的预测会为后续方法提供一个具体测试。


## 075 — A video edit must survive the next frame (00:53:51)

**00:53:51** A still image lets us choose a flattering instant.

静态图像让我们挑选一个最好看的瞬间。

**00:53:54** A video makes the object keep its promises in the next frame.

视频则要求物体在下一帧继续信守承诺。

**00:53:58** Look at the source sequence.

看原始序列。

**00:54:00** As viewpoint and visibility change, we continue to recognize the subject.

随着视角和可见性变化，我们仍能认出主体。

**00:54:06** An edit has to preserve that relationship while introducing its requested transformation.

编辑必须在引入所要求变换的同时，保留这种关系。

**00:54:12** Remember our two clocks.

记住两种时钟。

**00:54:14** The sampling steps describe how a model generates a result.

采样步描述模型如何生成结果。

**00:54:17** The frames here describe time passing in the depicted scene.

这里的帧描述被描绘场景中时间的流逝。

**00:54:22** A system may use many sampling steps to produce a short clip, but those steps are not extra seconds of action.

系统可能用很多采样步生成短片，但这些步数不是额外的剧情秒数。

**00:54:29** Now the preservation contract must hold across an event, not just inside a frame.

现在，保留约定必须跨越一个事件成立，而不只是局限于单帧内部。


## 076 — Consistency is not stillness (00:54:35)

**00:54:35** Would a video be perfectly consistent if every frame were identical?

如果每帧完全相同，视频就算完美一致吗？

**00:54:39** Only if the intended scene were perfectly still.

只有目标场景本来就完全静止时才是。

**00:54:43** In our moving shot, pose, viewpoint, illumination, and visibility should change.

在运动镜头中，姿态、视角、光照和可见性都应变化。

**00:54:52** Consistency means that those changes remain coherent with the scene.

一致性意味着这些变化与场景保持协调。

**00:54:56** The blue stem may become hidden during a turn.

蓝色柄可能在转动时被遮住。

**00:55:00** That is different from its color changing without a lighting explanation.

这与没有光照原因却突然变色，是不同的。

**00:55:04** A texture should move with its surface, rather than crawl independently across it.

纹理应该随表面一起移动，而不是独自在表面爬动。

**00:55:09** If there is no convincing explanation, you may have found instability.

如果找不到可信解释，你可能发现了不稳定性。

**00:55:14** This is why the goal cannot simply be to minimize change from one frame to the next.

因此，目标不能只是尽量减少相邻帧之间的变化。


## 077 — A moving image can reveal an idea (00:55:20)

**00:55:20** Here is a prepared storyboard, not a generated video result.

这是准备好的分镜，不是生成视频的结果。

**00:55:24** In the first view, the reflection suggests an object we have not fully seen.

第一个视角中，倒影暗示了一个我们尚未看全的物体。

**00:55:29** A partial view then gives us enough evidence to make a guess.

接着，局部视角提供足够线索，让我们猜测。

**00:55:33** The wide view finally changes our understanding of the setting.

最后的全景改变了我们对环境的理解。

**00:55:36** The audience is doing something during those five seconds: forming an expectation, testing it, and revising it.

在这五秒中，观众一直在行动：建立预期、检验预期，然后修正它。

**00:55:44** Compare this sequence with revealing everything in the first frame.

把这个序列与第一帧就揭示一切的版本比较一下。

**00:55:48** The same sculpture could be present, but the experience would differ.

同一件雕塑可以都在场，但体验会不同。

**00:55:52** For AFTER RAIN, the timing of information can carry fragility as strongly as material or lighting.

对于《雨后》，信息揭示的时机，可以像材质或灯光一样有力地承载脆弱感。

**00:55:58** We should decide that structure before asking a tool to fill in motion.

要求工具补充运动之前，我们应该先决定这种结构。


## 078 — Five seconds can have a structure (00:56:03)

**00:56:03** Five seconds is short, but it can still have a structure.

五秒虽然短，却仍然可以有结构。

**00:56:06** Our proposed first interval shows a reflection, giving the viewer clues.

我们设想的第一段展示倒影，给观众线索。

**00:56:11** The second offers a partial view.

第二段提供局部视角。

**00:56:13** The final interval reveals the wider setting and changes how the object is understood.

最后一段揭示更广阔的环境，改变人们对物体的理解。

**00:56:18** Try another allocation and predict the effect.

试着重新分配时间，并预测效果。

**00:56:21** A longer reflection might create uncertainty; a quick reveal might make the piece feel like a product shot.

更长的倒影段落可能制造不确定感；快速揭示则可能让它像产品广告镜头。

**00:56:28** These timings are artistic proposals, not measured outputs from a video model.

这些时间安排是艺术提案，不是视频模型的实测输出。

**00:56:34** By deciding the experience first, we can evaluate whether the generated motion and cuts support it.

先决定体验，才能评价生成的运动和剪辑是否支持它。

**00:56:40** A technically smooth sequence can still have the wrong rhythm.

技术上流畅的序列，节奏仍可能不对。


## 079 — An image editor can read a contact sheet (00:56:44)

**00:56:44** The research begins with a surprisingly simple experiment: arrange video frames into a contact sheet and give that image to an image editor.

这项研究始于一个出人意料的简单实验：把视频帧排成接触印样，再交给图像编辑器。

**00:56:53** This asks whether an existing image-editing capability can transfer across multiple views presented together.

它想检验：现有图像编辑能力，能否迁移到一起呈现的多个视角上？

**00:57:01** Treat the result as evidence that motivates a research direction.

请把结果当作启发研究方向的证据。

**00:57:05** A grid can make frames available in a shared context, but a successful-looking selection does not prove reliable video editing.

网格让多帧共享上下文，但一组看似成功的精选帧，并不能证明视频编辑可靠。

**00:57:12** There may be flicker between the sampled frames.

采样帧之间可能仍有闪烁。

**00:57:16** The next steps investigate how to represent video more effectively and adapt the model, rather than assuming a contact sheet alone solves time.

后续研究探索更有效的视频表示和模型适配，而不是假设接触印样本身就解决了时间问题。


## 080 — Four good frames can hide a bad video (00:57:24)

**00:57:24** But the missing intervals are where some of the most revealing failures live.

恰恰是在缺失的间隔中，可能藏着最有揭示性的失败。

**00:57:29** A feature can jump, disappear, and return between the frames on this page.

某个特征可能在这一页的两帧之间跳动、消失，再返回。

**00:57:33** Use the contact sheet for what it does well: compare appearance across selected moments and locate regions to inspect.

发挥接触印样的长处：比较选定时刻的外观，并定位需要检查的区域。

**00:57:41** Then review full playback, with closer attention around turns and occlusion.

然后完整播放，特别关注转向和遮挡附近。

**00:57:45** Think of a film review based only on publicity stills.

想象只凭宣传剧照来评论电影。

**00:57:49** You might judge the costume and lighting, but you have not yet seen the performance.

你或许能评价服装和灯光，却还没有看过表演。

**00:57:55** In generative video, that missing performance includes whether the same object continues to exist convincingly from moment to moment.

对生成视频而言，缺失的“表演”包括：同一个物体是否在每个时刻都可信地持续存在。


## 081 — The grid can live in latent space (00:58:04)

**00:58:04** The grid can also be constructed in latent space.

网格也可以在潜在空间中构造。

**00:58:07** Instead of only combining visible pixel images, the method encodes frames and arranges their tokens into a virtual image grid.

这种方法不只是组合可见像素图像，而是先编码各帧，再把 token 排列成虚拟图像网格。

**00:58:16** The published example changes a bus into a graphics card.

论文中的示例把一辆公交车变成显卡。

**00:58:20** The key idea is shared treatment of positions within that constructed representation.

核心思路是，在这种构造的表示中统一处理各位置。

**00:58:25** Do not confuse this step with using a video VAE; they are different design decisions.

不要把这一步与使用视频 VAE 混为一谈；它们是不同的设计决定。

**00:58:31** Putting frames near one another in a representation can help the model exchange information, but it does not impose a guarantee of physical continuity.

在表示中把帧放在一起，有助于模型交换信息，但不会强制保证物理连续性。

**00:58:40** We still need to inspect what survives across changing views.

我们仍需检查视角变化时，哪些信息得以保留。


## 082 — A virtual grid is a representation choice (00:58:45)

**00:58:45** A virtual grid provides a way to arrange information for a model trained around image-like structure.

虚拟网格为围绕图像结构训练的模型，提供了一种信息排列方式。

**00:58:51** It creates shared context and positional relationships among frame representations.

它在各帧表示之间建立共享上下文和位置关系。

**00:58:56** That can be a useful bridge when reusing an image editor.

复用图像编辑器时，这可以成为有用的桥梁。

**00:59:00** But arrangement alone does not enforce the laws of motion or the persistence of a hidden object.

但排列方式本身，不能强制执行运动规律或被遮挡物体的持续存在。

**00:59:05** Those abilities depend on the model, adaptation, data, and other parts of the workflow.

这些能力取决于模型、适配、数据，以及工作流程中的其他部分。

**00:59:12** Separate the representation choice from the capability claim.

请区分表示选择与能力主张。

**00:59:16** A diagram can explain how frames enter a system without proving that the system handles every difficult event.

图表可以解释帧如何进入系统，却不能证明系统处理得了所有困难事件。

**00:59:23** That proof would require appropriate outputs and evaluation.

证明这一点，需要适当的输出和评价。


## 083 — Repurposing an image model for video (00:59:28)

**00:59:28** This architecture asks whether an image editor's abilities can be reused for video.

这个架构检验图像编辑器的能力能否复用于视频。

**00:59:33** The video VAE encodes a clip into a compressed latent representation.

视频 VAE 将片段编码成压缩的潜在表示。

**00:59:39** Learned projections connect those video latents with the image-editing transformer, so the model can operate on a compatible arrangement of information.

学习得到的投影把视频潜变量与图像编辑 Transformer 连接起来，让模型处理兼容的信息排列。

**00:59:48** Trace that main route before reading the smaller branches.

阅读较小分支之前，先沿着这条主路径看。

**00:59:51** Adaptation connects them; the paper explores LoRA and full training choices, with an optional enhancement stage.

适配把它们连接起来；论文探索了 LoRA 和全量训练，并包含可选的增强阶段。

**00:59:59** The attractive idea is reuse of editing knowledge.

吸引人的想法是复用编辑知识。

**01:00:03** The test is whether that reuse preserves the temporal relationships our shot needs.

真正的检验是：这种复用是否保留了镜头需要的时间关系？

**01:00:08** Architecture explains where information flows; the edited clip shows whether the intended event survives.

架构解释信息流向哪里；编辑后的片段展示预期事件是否得以保留。


## 084 — Adaptation connects incompatible representations (01:00:15)

**01:00:15** The image model and video VAE do not necessarily speak the same representational language.

图像模型与视频 VAE 未必使用相同的表示语言。

**01:00:22** Learned projections help connect them.

学习得到的投影帮助连接二者。

**01:00:25** Adaptation then allows the editing transformer to operate usefully with the video representation.

之后，适配让编辑 Transformer 能有效使用视频表示。

**01:00:31** This is a common engineering pattern: reuse a capable component while learning the interface and behavior needed for another task.

这是一种常见工程模式：复用有能力的组件，同时学习另一项任务所需的接口和行为。

**01:00:39** The benefit is not automatic.

收益并不会自动出现。

**01:00:41** A projection must preserve useful information, and the adapted model must learn how the new structure relates to edits.

投影必须保留有用信息，适配后的模型也必须学会新结构如何与编辑相关联。

**01:00:49** In the published pipeline, decoding returns the edited representation to video.

在已发表的流程中，解码将编辑后的表示还原为视频。

**01:00:56** Follow the information through every stage rather than treating the name of the reused model as an explanation by itself.

请追踪信息经过每个阶段，而不要把被复用模型的名字本身当作解释。


## 085 — 45 frames become 12: why? (01:01:04)

**01:01:04** Here is a counting puzzle with a small trap.

这里有一个带小陷阱的计数题。

**01:01:06** In this example, forty-five pixel frames are compressed with a special first frame and a factor of four for the remaining temporal groups.

在这个例子中，四十五个像素帧采用特殊首帧处理，其余时间组按四倍压缩。

**01:01:15** Forty-five is four times eleven, plus one.

四十五等于四乘十一，再加一。

**01:01:19** The latent sequence therefore has eleven plus one, or twelve frames.

因此，潜在序列有十一加一，即十二帧。

**01:01:24** Those twelve latent frames can be arranged in the illustrated three-by-four virtual grid.

这十二个潜在帧可以排成图中的三乘四虚拟网格。

**01:01:29** Simply dividing forty-five by four would miss the first-frame convention.

简单地用四十五除以四，会忽略首帧约定。

**01:01:34** This arithmetic belongs to the representation used in this example; it is not a rule for every video model.

这个算法属于本例采用的表示，并不是所有视频模型的通用规则。

**01:01:41** A small detail in temporal compression changes what the editing model actually receives.

时间压缩中的一个小细节，会改变编辑模型实际收到什么。


## 086 — Temporal compression preserves a special first frame (01:01:47)

**01:01:47** Solve the frame-count relationship step by step.

一步步解出帧数关系。

**01:01:50** Start with four k plus one equals forty-five.

从 4k 加一等于四十五开始。

**01:01:54** Subtract one to obtain forty-four, then divide by four to get k equals eleven.

减一得到四十四，再除以四，得到 k 等于十一。

**01:02:00** The latent sequence contains k plus one frames, so its length is twelve.

潜在序列包含 k 加一帧，因此长度为十二。

**01:02:06** This arithmetic encodes the special handling of the first frame in the stated convention.

这个计算反映了该约定对首帧的特殊处理。

**01:02:11** It is different from simply dividing the total number of pixel frames by four.

它不同于直接把像素帧总数除以四。

**01:02:17** Understanding them can prevent mistakes when arranging latent frames into a virtual grid or comparing the representation with the original clip.

理解这些关系，可以避免在排列潜在帧网格、或与原始片段比较时出错。


## 087 — Guidance in a published video ablation (01:02:26)

**01:02:26** The requested edit turns the scene into a cyberpunk workshop with holographic documents.

编辑要求是把场景变成赛博朋克工作室，并加入全息文档。

**01:02:32** Compare the results before focusing on the setting labels.

先比较结果，再关注设置标签。

**01:02:36** Which version carries out the requested transformation more completely, and what visible evidence supports your answer?

哪个版本更完整地实现了要求的变换？有什么可见证据？

**01:02:43** The paper presents this case as an example where classifier-free guidance improves edit completeness.

论文用这个案例说明，无分类器引导可以改善编辑完整性。

**01:02:50** Its baseline retains source latents, and its implementation includes rescaling.

其基线保留了原始潜变量，实现中还包含重缩放。

**01:02:55** This does not establish one universally best guidance value.

这并不意味着存在一个普遍最优的引导值。

**01:02:59** Guidance also has a computation cost when it requires another model pass.

如果引导需要额外一次模型前向计算，也会带来计算成本。

**01:03:05** Relate the example back to our equation: changing a prediction combination affects how strongly the result follows the condition.

联系之前的方程：改变预测组合，会影响结果遵循条件的力度。


## 088 — What this comparison actually shows (01:03:13)

**01:03:13** An ablation asks what changes when one component is removed.

消融实验追问：移除某个组件后，会发生什么变化？

**01:03:17** In the displayed example, guidance makes more of the requested transformation visible.

在展示的例子中，引导让更多要求的变换可见。

**01:03:22** That is useful evidence about this comparison.

这是关于这次比较的有用证据。

**01:03:26** It is not yet a measurement of reliability across the kinds of shots we might produce for an exhibition.

但它还不是对我们可能用于展览的各类镜头的可靠性测量。

**01:03:32** As a reader, separate the observation from the next question.

作为读者，请区分观察结果与下一步问题。

**01:03:35** We can observe a more complete edit here.

我们可以观察到，这里的编辑更加完整。

**01:03:38** We would still want to know what happens across other clips, how often preservation suffers, and what the computation costs.

但仍想知道：其他片段会怎样、保留效果多常受损，以及计算成本是多少。

**01:03:46** You can learn a mechanism from a selected example while designing a stronger test for the decision you actually need to make.

你可以从精选示例中学习机制，同时为真正需要做出的决定设计更强的测试。


## 089 — Local editing across frames (01:03:55)

**01:03:55** This is a local edit: the sheep's face changes and receives a white star-shaped patch.

这是一次局部编辑：绵羊的脸发生变化，增加了一块白色星形斑纹。

**01:04:00** Look at the patch across poses.

观察不同姿态下的这块斑纹。

**01:04:03** Does it remain attached to the same facial region, with a plausible change in apparent shape as the head turns?

它是否始终附着在同一面部区域，并在头部转动时合理改变表观形状？

**01:04:10** The most attractive frame is not enough.

仅看最漂亮的一帧是不够的。

**01:04:12** An identity marker can slide, disappear, or change shape at another moment.

身份标记可能在另一个时刻滑动、消失或变形。

**01:04:17** Selected frames let us ask the right questions, but the full clip is needed to check continuity between them.

精选帧让我们提出正确问题，但要检查帧间连续性，仍需完整视频。

**01:04:25** For our discussion, identify a small recognizable feature that will make identity drift easier to notice.

请为讨论找一个可辨认的小特征，让身份漂移更容易被发现。


## 090 — Follow the identity marker (01:04:33)

**01:04:33** Choose a distinctive feature and follow it through the shot: its attachment to the subject, its shape, and its reappearance after partial visibility.

选一个独特特征，并在镜头中追踪它：与主体的附着关系、形状，以及部分遮挡后的再次出现。

**01:04:43** For our pear, the blue stem is a useful witness.

对这只梨而言，蓝色柄是有用的见证者。

**01:04:47** It may become hidden as the camera moves; we should allow that.

随着相机移动，它可能被遮住；这是应当允许的。

**01:04:51** When it returns, it should still belong to the same object.

当它返回时，仍应属于同一个物体。

**01:04:54** A marker that slides across the surface or reappears in a new shape tells us something a general impression of smooth motion might miss.

如果标记在表面滑动，或以新形状返回，就揭示了整体流畅感可能掩盖的问题。

**01:05:02** We are using the marker as a diagnostic aid, while still judging the whole object's identity and the shot's intended motion.

我们用标记辅助诊断，同时仍评价整个物体的身份和镜头预期运动。


## 091 — Global stylization across frames (01:05:11)

**01:05:11** Here the instruction changes the whole sequence into a minimal monochrome sketch.

这里的指令把整个序列变成极简单色素描。

**01:05:16** Local object replacement and global stylization allow different degrees of visual freedom.

局部物体替换与整体风格化，允许的视觉自由程度不同。

**01:05:22** Even with a large style change, motion and scene structure should remain readable.

即使风格变化很大，运动和场景结构也应保持可读。

**01:05:28** Look for rules across frames: line density, silhouette treatment, and the handling of depth.

寻找跨帧规则：线条密度、轮廓处理，以及深度的表现方式。

**01:05:35** Do they feel like one visual language?

它们是否像同一种视觉语言？

**01:05:38** This connects directly to our collage discussion.

这与我们的拼贴讨论直接相连。

**01:05:42** A style is more useful to an art director when described through operations that can persist through time.

当风格被描述为能够持续存在于时间中的操作时，它对艺术总监更有用。

**01:05:48** If every frame reinvents those operations, the result may feel unstable even when each still is appealing.

如果每一帧都重新发明这些操作，即使每张静帧好看，整体也可能不稳定。


## 092 — A style can flicker while motion stays correct (01:05:56)

**01:05:56** A video can preserve object motion while its style flickers.

视频可以保留物体运动，风格却仍然闪烁。

**01:06:00** Line density may jump, a paper texture may crawl, or shading may switch between flat and volumetric treatments without a scene explanation.

线条密度可能跳变，纸张纹理可能爬动，明暗处理也可能无故在平面与立体之间切换。

**01:06:09** Review style as a set of temporal rules, just as we reviewed it across surfaces in the collage.

像审查拼贴不同表面的风格那样，把视频风格作为一组时间规则来审查。

**01:06:15** Some variation is appropriate when the camera or light changes.

相机或光线变化时，某些变化是合理的。

**01:06:19** The question is whether the material language remains coherent.

问题是：材质语言是否仍然连贯？

**01:06:23** If your artwork depends on a particular drawing or collage treatment, this check is as important as whether the subject stays in the same place.

如果作品依赖特定绘画或拼贴处理，这项检查与主体位置是否稳定同样重要。

**01:06:32** Motion correctness alone does not establish visual consistency.

运动正确，本身并不能证明视觉一致。


## 093 — Why raw frame difference is misleading (01:06:37)

**01:06:37** If we subtract one frame from the next, a perfectly correct camera movement can create a large difference.

如果将相邻帧相减，完全正确的相机运动也会产生很大差异。

**01:06:44** The same surface has simply moved to another pixel location.

同一个表面，只是移动到了另一个像素位置。

**01:06:49** A more meaningful comparison first aligns corresponding visible content, then measures the remaining appearance difference.

更有意义的比较，是先对齐对应的可见内容，再测量剩余外观差异。

**01:06:57** The teaching equation uses a motion warp for that alignment and a visibility mask to exclude occluded regions.

这个教学公式使用运动变换来对齐，并用可见性蒙版排除被遮挡区域。

**01:07:04** We should not demand agreement for content that is hidden or newly revealed.

对于隐藏或刚被揭示的内容，不应要求直接一致。

**01:07:09** This is an illustrative metric, not a claim about the exact evaluation used by the paper.

这是示意性指标，不是在声称论文使用了这一精确评测。

**01:07:15** A low difference can also reward a video that barely moves.

低差异也可能奖励几乎不动的视频。

**01:07:19** A metric must be checked against the task it is supposed to represent.

必须对照指标所要代表的任务，检查指标本身。


## 094 — The frozen video wins the wrong test (01:07:24)

**01:07:24** Imagine two candidates for our five-second shot.

想象五秒镜头的两个候选版本。

**01:07:27** One shows a coherent camera move around the pear with modest frame differences.

一个展示相机连贯地绕梨移动，帧间差异适中。

**01:07:32** The other repeats a single beautiful frame.

另一个反复显示同一张漂亮图像。

**01:07:35** A naive difference score might prefer the frozen version.

简单的差异分数，可能偏爱冻结版本。

**01:07:39** Our audience would immediately notice that the intended reveal never happened.

观众却会立刻发现，预期的揭示根本没有发生。

**01:07:43** That is a useful example of optimizing the measurement while missing the purpose.

这是优化测量指标却错失目的的一个例子。

**01:07:48** We need evidence for both coherent appearance and requested motion.

我们需要外观连贯和所要求运动两方面的证据。

**01:07:52** Neither can substitute for the other.

两者不能相互替代。

**01:07:55** Each time, one easy-to-observe quality tempted us to stand in for the whole brief.

每一次，我们都容易让某个易于观察的品质，代表整份创作要求。

**01:08:00** Good evaluation keeps the intended result in view, especially when a convenient score seems reassuring.

良好的评测始终关注预期结果，尤其当方便的分数令人安心时。


## 095 — A keyframe can anchor the edit (01:08:08)

**01:08:08** An edited keyframe can make an art direction easier to approve.

编辑后的关键帧，可以让艺术方向更容易获得认可。

**01:08:12** Instead of describing every material and lighting choice in words, we can point to an image and say, this is the appearance we want.

不必用文字描述每个材质和灯光决定，我们可以指着一张图说：这就是想要的外观。

**01:08:21** The workflow then carries that approved look through the source clip.

工作流程再把获批的外观延续到原始片段中。

**01:08:24** Follow the four stages: select a representative frame, edit and inspect its appearance, propagate the look, and review the resulting motion.

请跟随四个阶段：选代表帧，编辑并检查外观，传播外观，再审查运动结果。

**01:08:36** Will the material persist during a turn?

转向时，材质能保持吗？

**01:08:39** Will the identity survive a column passing in front of it?

柱子从前方经过时，身份能保留吗？

**01:08:44** This distinction is central to the Runway reading later: an appealing interaction pattern still needs a shot review suited to the artwork.

这个区别是后面 Runway 阅读的核心：吸引人的交互方式，仍需要适合这件作品的镜头审查。


## 096 — AlignVid: when the image overrules the text (01:08:52)

**01:08:52** What if the approved image is so influential that the requested event never happens?

如果获批图像影响力太强，以至于要求的事件根本没发生，怎么办？

**01:08:57** Our AlignVid collaboration studies that tension in image-to-video generation.

我们合作开展的 AlignVid 研究，探索图生视频中的这种张力。

**01:09:02** In the published examples, a baseline omits a sunflower or leaves a person standing; the corresponding AlignVid results implement more of the requested event.

论文示例中，基线漏掉了向日葵，或让人物一直站着；对应的 AlignVid 结果实现了更多要求的事件。

**01:09:13** The intervention scales queries or keys in selected attention blocks and denoising steps without retraining.

这种干预无需重新训练，只在选定注意力模块和去噪步骤中，缩放查询或键。

**01:09:19** In a scalar form, scaling Q by gamma changes the weights to softmax of gamma times Q K transpose over square root d.

在标量缩放形式下，用 gamma 缩放 Q 后，权重变成 softmax(gamma × QKᵀ / √d)。

**01:09:28** It changes attention concentration.

它改变了注意力的集中程度。

**01:09:31** Classifier-free guidance instead combines predictions with different conditioning; these are distinct operations.

无分类器引导则组合不同条件下的预测；两者是不同操作。

**01:09:38** For an artist, the interesting failure is a faithful-looking image that refuses to do what the scene requires.

对艺术家来说，有意思的失败是：图像看起来很忠实，却拒绝做场景要求的事。

**01:09:44** The published frames illustrate that tension; a finished shot still needs temporal review.

已发表的帧展示了这种张力；完成的镜头仍需进行时间连续性审查。


## 097 — A five-second video edit brief (01:09:50)

**01:09:50** Our five-second brief changes the pear's body to ivory ceramic while retaining the blue stem, camera movement, and position.

我们的五秒要求，是把梨的主体变为象牙白陶瓷，同时保留蓝色柄、相机运动和位置。

**01:09:58** Reflections may adapt to the new material.

倒影可以适应新材质。

**01:10:00** The pear must remain the same object after passing behind a column.

梨经过柱子后面再出现时，必须仍是同一个物体。

**01:10:05** Which clause is hardest to verify?

哪一条最难验证？

**01:10:08** The reappearance is a strong candidate, because the system must maintain identity through a period of invisibility.

再次出现很可能最难，因为系统必须在一段不可见期间保持身份。

**01:10:14** This is more demanding than transferring a visible color from frame to frame.

这比逐帧传递可见颜色更具挑战。

**01:10:19** Begin with one clear transformation in a short clip.

先在短片中做一个明确的变换。

**01:10:23** Then make the critical visibility event part of the review, rather than discovering it only after choosing a favorite result.

然后把关键可见性事件纳入审查，而不是选好最喜欢的结果后才发现它。


## 098 — The moment the pear returns (01:10:31)

**01:10:31** The column is the moment of truth.

柱子就是关键考验。

**01:10:34** Before the pear disappears, we can inspect its silhouette, material, and stem.

梨消失前，我们能检查它的轮廓、材质和柄。

**01:10:39** During occlusion, there may be nothing visible to compare.

遮挡期间，可能没有可见内容可供比较。

**01:10:43** When it returns, the model has to make the same object convincing again.

返回时，模型必须再次让同一个物体显得可信。

**01:10:48** Do not reject correct invisibility as a failure.

不要把正确的不可见状态当作失败。

**01:10:51** Instead, compare the identity before and after the event, including how the object emerges at the boundary.

而应比较事件前后的身份，包括物体如何从边界处显露。

**01:10:58** A smooth-looking clip can still reveal a subtly redesigned pear on the other side.

看起来流畅的视频，也可能在另一侧露出一只被悄悄重新设计的梨。

**01:11:03** We are not merely watching for anything strange; we are testing the preservation requirement at the point where the shot makes it hardest to satisfy.

我们不只是寻找怪异之处，而是在镜头最难满足要求的地方，检验保留约束。


## 099 — A small addition creates new relationships (01:11:12)

**01:11:12** Adding a small object creates several new relationships.

添加一个小物体，会创造多种新关系。

**01:11:16** In this published example, the edit adds a drone.

在这个已发表示例中，编辑添加了一架无人机。

**01:11:20** Even if the drone looks convincing by itself, its scale, placement, and movement must fit the scene.

即使无人机本身很可信，它的尺度、位置和运动也必须适合场景。

**01:11:27** Imagine a camera moving forward while the added object changes size in the wrong direction.

想象相机向前移动，新增物体的大小却朝错误方向变化。

**01:11:32** The object might be beautifully rendered, but its relationship with the camera would expose the edit.

物体可能渲染得很漂亮，但它与相机的关系会暴露编辑。

**01:11:38** Or it might hover at an unintended location relative to other objects.

或者，它可能相对其他物体悬停在错误位置。

**01:11:42** When you review an addition, evaluate those relationships across time.

审查新增物体时，要跨时间评价这些关系。

**01:11:47** A close-up of the inserted object cannot answer every question about whether it belongs.

仅看新增物体的特写，无法回答它是否融入场景的所有问题。


## 100 — An addition must belong to camera motion (01:11:53)

**01:11:53** An added object must fit the camera as well as the scene.

新增物体既要符合场景，也要符合相机。

**01:11:57** A convincing texture and shape in one still do not establish a coherent trajectory.

单张静帧中可信的纹理和形状，不能证明运动轨迹连贯。

**01:12:02** If the camera approaches, the apparent size and position of the addition should evolve in a compatible way.

如果相机靠近，新增物体的表观大小和位置应相应变化。

**01:12:09** The exact expectation depends on whether the object is stationary or moving independently.

具体预期取决于物体是静止，还是独立运动。

**01:12:15** That is why the brief should specify the intended relationship.

因此，创作说明应该明确预期关系。

**01:12:19** Inspect the addition relative to nearby objects and the background, not only in a crop.

要相对附近物体和背景检查新增内容，而不只看裁剪图。

**01:12:26** A production-quality edit is a collection of relationships that survive over time, rather than a new object that looks impressive in isolation.

制作级编辑是一组经得起时间考验的关系，而不是一个单独看很惊艳的新物体。


## 101 — The shot review (01:12:35)

**01:12:35** Review the whole shot at normal speed, then inspect moments where failure is most likely.

先以正常速度看完整个镜头，再检查最可能失败的时刻。

**01:12:41** Pay attention to identity, intended motion, occlusion, boundaries, and the ending.

关注身份、预期运动、遮挡、边界和结尾。

**01:12:46** High-motion content is a limitation the research authors specifically identify.

高运动量内容，是研究作者明确指出的局限。

**01:12:51** If a clip fails, choose a revision based on the failure.

片段失败时，要根据具体失败选择修改。

**01:12:54** You might reduce the transformation, shorten the segment, add a reference frame where supported, or repair part of the result through conventional compositing.

可以减小变换幅度、缩短片段、在支持时添加参考帧，或用传统合成修复部分结果。

**01:13:05** Another generation is useful when it tests a reasoned hypothesis.

当新一轮生成用于检验有理由的假设时，它才有用。

**01:13:09** Do not let repeated attempts distract you from whether the piece still communicates the artistic intention you began with.

不要让反复尝试分散注意力，忘记作品是否仍在传达最初的艺术意图。


## 102 — A failed shot suggests a different workflow (01:13:16)

**01:13:16** A failed shot is information about the workflow.

失败镜头提供了关于工作流程的信息。

**01:13:20** If the appearance is wrong from the beginning, revising the keyframe or the material direction may help.

如果一开始外观就不对，修改关键帧或材质方向可能有帮助。

**01:13:26** If the first frame is convincing but identity breaks after occlusion, the problem calls for temporal review and a more suitable propagation or editing approach.

如果首帧可信，但遮挡后身份破坏，就需要时间连续性审查，以及更合适的传播或编辑方法。

**01:13:36** If a region must remain exact, direct compositing may be part of the solution.

如果某个区域必须完全不变，直接合成可以是解决方案的一部分。

**01:13:41** We can also split a complicated transformation into stages, provided the transitions remain coherent.

也可以把复杂变换拆成多个阶段，但必须保证过渡连贯。

**01:13:48** The important move is to connect the observed failure to the next intervention.

关键是把观察到的失败，与下一步干预联系起来。

**01:13:53** Repeating a vague request with more enthusiasm does not use what the failed shot has taught us.

更热情地重复模糊要求，并没有利用失败镜头教给我们的东西。


## 103 — Discuss: the difficult moment (01:13:59)

**01:13:59** Return to the proposed five-second ceramic-pear shot.

回到拟议的五秒陶瓷梨镜头。

**01:14:03** Which moment would you inspect most closely?

你会最仔细检查哪个时刻？

**01:14:05** What should remain recognizable after the column?

经过柱子后，什么应该仍然可辨认？

**01:14:08** What failure would make you reject an attractive clip?

什么失败会让你拒绝一段好看的视频？

**01:14:12** Discuss these three questions by describing an event and the evidence you would seek.

请通过描述事件和所需证据，讨论这三个问题。

**01:14:17** If your answer is temporal consistency, make it visible: what changes, when, and why would that violate the intended shot?

如果你的答案是“时间一致性”，请让它可见：什么在何时改变，为什么违反了镜头意图？

**01:14:27** Pause the video here.

请在这里暂停视频。

**01:14:29** Compare your acceptance criteria with another person’s before we examine the compact review plan.

在查看简要审查方案之前，先与另一位同学比较验收标准。


## 104 — A five-second test plan (01:14:38)

**01:14:38** Here is a compact test plan.

这是一个简要测试方案。

**01:14:41** The request is an ivory ceramic body with the blue stem retained.

要求是象牙白陶瓷主体，并保留蓝色柄。

**01:14:45** Camera motion and object motion should follow the source.

相机运动和物体运动应遵循原片。

**01:14:49** The clip includes a column so that we can inspect a difficult reappearance.

片段中包含柱子，让我们检查困难的再次出现过程。

**01:14:53** The evidence includes normal-speed playback and a closer look before and after that event.

证据包括正常速度播放，以及事件前后的仔细观察。

**01:14:59** We also inspect reflections because a material change has optical consequences.

还要检查倒影，因为材质变化会产生光学后果。

**01:15:04** This plan does not require a long benchmark suite.

这个方案不需要庞大的基准测试集。

**01:15:08** It simply makes the request, the likely failure, and the acceptance evidence explicit.

它只是明确要求、可能的失败，以及验收证据。

**01:15:14** You can use the same structure to test a different subject or visual transformation.

你可以用同样的结构测试其他主体或视觉变换。


## 105 — Three decisions without a model (01:15:19)

**01:15:19** Let us check whether the mechanisms are now useful without a model in front of us.

现在没有模型在面前，让我们检验这些机制是否真的有用了。

**01:15:24** Consider the three questions on screen.

请考虑屏幕上的三个问题。

**01:15:27** A material change reaches the reflection: explain why.

材质变化会影响倒影：请解释原因。

**01:15:32** A subject reference and LoRA both help represent a concept: explain the difference.

主体参考图和 LoRA 都有助于表达概念：请解释两者区别。

**01:15:39** Four frames look excellent: explain why the clip can still fail.

四帧看起来很出色：请解释整段视频为什么仍可能失败。

**01:15:44** Take a short silent beat before answering.

回答前，先安静思考片刻。

**01:15:47** The next slide names the decisions hiding inside the questions.

下一页会指出这些问题中隐藏的决定。

**01:15:51** If you get stuck, locate the type of problem first: the physical scene, the model's information, or evidence across time.

如果卡住了，先定位问题类型：物理场景、模型获得的信息，还是跨时间的证据。

**01:16:00** That is often enough to begin a clear answer.

通常，做到这一步就足以开始给出清晰答案。


## 106 — Match the problem to the intervention (01:16:03)

**01:16:03** Here are the decisions those questions conceal.

这些问题隐藏着以下决定。

**01:16:06** If the reflection must change with the new material, we need to define the permitted region and consequences of the edit.

如果倒影必须随新材质变化，就需要定义编辑允许的区域和连带后果。

**01:16:14** If we want a concept for this request, a reference can condition generation; if we want learned adaptation, LoRA changes selected weights.

如果只为这次要求表达概念，参考图可提供生成条件；如果需要学习式适配，LoRA 会改变选定权重。

**01:16:25** If continuity matters, the evidence must include the intervals between our chosen stills.

如果重视连续性，证据就必须包含精选静帧之间的间隔。

**01:16:30** Each decision rules out a tempting shortcut.

每个决定，都排除了一条诱人的捷径。

**01:16:34** Freezing every surrounding pixel may freeze the wrong reflection.

冻结周围每一个像素，可能冻结错误的倒影。

**01:16:38** Calling every supplied image training confuses conditioning with adaptation.

把所有输入图像都称为训练，会混淆条件控制与适配。

**01:16:43** Calling four stills a successful video omits time.

把四张静帧称为成功视频，就遗漏了时间。

**01:16:47** Use this slide to repair your explanation, rather than memorize a preferred sentence.

请用这一页修正你的解释，而不是背诵某句标准答案。


## 107 — The mechanisms behind your decisions (01:16:53)

**01:16:53** The material changes how light interacts with the pear, so its visible consequences can extend into the reflection.

材质改变光与梨的相互作用，因此可见后果会延伸到倒影。

**01:17:01** A reference conditions an output; LoRA learns a low-rank update to selected model weights.

参考图为输出提供条件；LoRA 为选定模型权重学习低秩更新。

**01:17:07** Four good frames leave unobserved transitions, including events such as occlusion where identity can fail.

四张好帧仍留下未观察的过渡，包括可能导致身份失败的遮挡事件。

**01:17:14** Those are compact answers, but each has an application.

这些答案很简短，但每个都有实际用途。

**01:17:18** They help us write a better preservation contract, choose between two kinds of intervention, and design a more revealing review.

它们帮助我们写出更好的保留约定、在两类干预中选择，并设计更有揭示性的审查。

**01:17:26** If your answer used different words and preserved those distinctions, it works.

如果你用了不同措辞，但保留了这些区别，答案就成立。

**01:17:31** Tomorrow's interface may rename its controls, but the difference between specifying a request, adapting a model, and checking its result will still matter.

明天的界面可能给控件改名，但提出要求、适配模型与检查结果之间的区别，仍然重要。


## 108 — Three levels of checking an edit (01:17:41)

**01:17:41** We can check an edit at three levels.

我们可以在三个层面检查编辑。

**01:17:43** First, did the requested transformation occur?

第一，所要求的变换发生了吗？

**01:17:47** Second, did the necessary invariants survive?

第二，必要的不变量保留了吗？

**01:17:50** Third, does the result serve the artwork's intention?

第三，结果是否服务于作品意图？

**01:17:54** A result can pass the first level and fail the second, as when the correct material appears on a different object.

结果可能通过第一层却不通过第二层，例如正确材质出现在了另一个物体上。

**01:18:01** It can pass both technical levels and still fail the artistic one, as when the chosen effect undermines the intended mood.

也可能通过两个技术层面，却在艺术层面失败，例如所选效果削弱了预期情绪。

**01:18:10** Keeping the levels separate makes critique more precise.

区分这些层面，会让评论更精确。

**01:18:14** It also explains why neither a generic aesthetic score nor a strict pixel comparison can answer every question we care about.

这也解释了为什么通用审美分数和严格像素比较，都无法回答我们关心的所有问题。


## 109 — Who decides? (01:18:22)

**01:18:22** We have spent the lecture making the editing request more precise.

整堂课，我们都在让编辑要求更精确。

**01:18:27** Now consider who gets to choose the request in the first place.

现在想想：一开始由谁来选择提出什么要求？

**01:18:31** A system can offer convincing alternatives, but somebody still decides which intention matters and which evidence counts as success.

系统可以提供可信方案，但仍需要有人决定什么意图重要，以及什么证据才算成功。

**01:18:41** The three readings let us examine that responsibility from different positions.

三篇阅读让我们从不同位置审视这种责任。

**01:18:46** Pachocki raises questions about capable AI systems and human values.

Pachocki 提出了能力强大的 AI 系统与人类价值观之间的问题。

**01:18:51** Koe asks readers to reconsider their own goals and habits.

Koe 请读者重新审视自己的目标与习惯。

**01:18:55** Runway presents a workflow for turning an approved image into a video edit.

Runway 展示把获批图像转化为视频编辑的工作流程。

**01:18:59** Keep our pear in mind as we read them: we can delegate parts of its production while still arguing about what the artwork should become.

阅读时请记住这只梨：我们可以委托部分制作，同时仍讨论作品应该成为什么。


## 110 — After Topic 3: judge the whole shot (01:19:08)

**01:19:08** We have an approved appearance and a moving shot to judge.

现在，我们有了获批外观，以及需要评价的运动镜头。

**01:19:11** What can an approved keyframe establish, and what can it not?

获批关键帧能确定什么，又不能确定什么？

**01:19:16** Why might a low frame-difference score reward a bad video?

为什么低帧差分数可能奖励一段坏视频？

**01:19:20** Which evidence would convince you that identity survived motion?

什么证据能说服你，身份经受住了运动？

**01:19:24** Discuss these three questions.

请讨论这三个问题。

**01:19:26** Your answer should account for the requested event as well as the subject’s appearance.

答案既要考虑主体外观，也要考虑所要求的事件。

**01:19:30** Use the column, a turn, or the sketch treatment as a concrete example.

可以用柱子、转向或素描处理作为具体例子。

**01:19:35** Pause the video here.

请在这里暂停视频。

**01:19:36** We will next ask how the readings change our view of responsibility for those decisions.

接下来，我们将追问阅读如何改变我们对这些决定所涉及责任的看法。


## 111 — AI goals and human direction (01:19:45)

**01:19:45** These first two readings ask about direction from different sides.

前两篇阅读从不同方面追问方向。

**01:19:49** An Alien Mind considers the relationship between an AI system achieving a goal and respecting human values.

《An Alien Mind》思考 AI 系统实现目标与尊重人类价值观之间的关系。

**01:19:57** Dan Koe's essay invites people to examine the goals and habits directing their own lives.

Dan Koe 的文章邀请人们审视引导自己生活的目标和习惯。

**01:20:02** We will use its proposal as material for critical discussion of creative practice.

我们把他的主张作为批判性讨论创作实践的材料。

**01:20:07** Find a specific claim, explain your interpretation, and connect it to a decision from the lecture.

找出一个具体主张，解释你的理解，并联系课堂中的一个决定。

**01:20:14** A fictional artist is fine for discussing personal direction.

讨论个人方向时，可以使用虚构艺术家的例子。

**01:20:18** The fifth quiz question asks you to reflect on one reading of your choice.

测验第五题要求你从自选的一篇阅读出发进行反思。

**01:20:23** Our class discussion can compare all three perspectives without requiring everyone to disclose personal experiences.

课堂可以比较三种视角，而不要求每个人披露个人经历。


## 112 — Reading 3: image-guided video editing (01:20:30)

**01:20:30** The third reading moves from questions about direction to a concrete editing workflow.

第三篇阅读从方向问题，转向具体编辑流程。

**01:20:36** Runway's announcement of Aleph 2.0 and Edit Studio describes establishing an appearance in an edited image and applying that change through video.

Runway 关于 Aleph 2.0 和 Edit Studio 的公告，描述了先在编辑图像中确定外观，再把变化应用到视频的方法。

**01:20:45** It connects directly to our keyframe discussion.

它与关键帧讨论直接相连。

**01:20:48** Read it with the ceramic pear and column in mind.

阅读时，请想着陶瓷梨和柱子。

**01:20:51** Which decision does the interface make easier to express?

这个界面让哪项决定更容易表达？

**01:20:55** Which result would you still need to watch before approval?

批准之前，哪些结果仍需要亲眼播放检查？

**01:20:59** The page is a vendor's account, so its demonstrations illustrate claimed capabilities rather than supply an independent comparative evaluation.

这是一篇厂商介绍，因此演示说明的是其声称的能力，而不是独立的比较评测。

**01:21:08** We can learn from the interaction design and still propose an occlusion test that asks whether the workflow meets this exhibition's particular needs.

我们可以借鉴交互设计，同时提出遮挡测试，检验流程是否满足这个展览的具体需要。


## 113 — Optional technical reading (01:21:17)

**01:21:17** These technical papers are optional companions to the discussion readings.

这些技术论文，是讨论阅读之外的可选配套材料。

**01:21:21** Choose according to the question you want to investigate.

请根据你想研究的问题选择。

**01:21:24** InstructPix2Pix helps explain edit supervision; Flow Matching develops the generative training framework.

InstructPix2Pix 帮助解释编辑监督；Flow Matching 发展了生成训练框架。

**01:21:32** Qwen-Image-2.0 and Qwen-Video-Edit provide recent architecture examples.

Qwen-Image-2.0 和 Qwen-Video-Edit 提供了近期架构示例。

**01:21:38** You do not need to read all of them before trying the class discussion.

参加课堂讨论之前，不必读完全部论文。

**01:21:42** Start with a question, locate the part of the paper that addresses it, and distinguish the proposed method from the authors' evidence.

从一个问题开始，找到论文中回应它的部分，并区分提出的方法与作者的证据。

**01:21:50** The slide notes retain the other references on multi-image inputs, rewards, composition, and video editing.

幻灯片备注保留了多图输入、奖励、构图和视频编辑方面的其他参考资料。


## 114 — Three source types, three kinds of evidence (01:21:58)

**01:21:58** The three readings provide different kinds of material.

三篇阅读提供不同类型的材料。

**01:22:02** A research leader's perspective develops an argument about alignment and future development.

研究领导者的观点文章，提出关于对齐和未来发展的论证。

**01:22:07** A reflective essay offers a way to examine personal direction.

反思性文章提供一种审视个人方向的方法。

**01:22:12** A vendor announcement presents a workflow and promotes its capabilities.

厂商公告展示工作流程，并推广其能力。

**01:22:16** Read each according to what it can support.

阅读每种材料时，要考虑它能够支持什么。

**01:22:19** Identify forecasts, personal claims, and demonstrations rather than treating every sentence as the same kind of evidence.

识别预测、个人主张和演示，不要把每句话都当作同类证据。

**01:22:27** You may disagree with an author and still find a useful question.

你可以不同意作者，却仍从中找到有用的问题。

**01:22:31** Our synthesis is practical: what do you want to make, what will you delegate, and what evidence will you use to decide whether the collaboration served that intention?

我们的综合问题很实际：你想创作什么、委托什么，以及用什么证据判断合作是否服务于这一意图？


## 115 — Separate image and text guidance (01:22:42)

**01:22:42** This technical extension separates image guidance and text guidance in InstructPix2Pix.

这个技术延伸，区分 InstructPix2Pix 中的图像引导和文本引导。

**01:22:49** Begin with the prediction using neither condition.

从不使用任何条件的预测开始。

**01:22:52** Add a scaled difference for including the image, then another scaled difference for adding the text alongside that image.

加上引入图像后的缩放差值，再加上在图像之外引入文本后的另一个缩放差值。

**01:22:59** Set both scales to one and follow the cancellation: the intermediate terms disappear, leaving the fully conditioned noise prediction.

把两个强度都设为一，观察抵消：中间项消失，只留下完全条件化的噪声预测。

**01:23:07** This is a useful check that you understand the expression.

这是检查你是否理解表达式的有用方法。

**01:23:11** Epsilon here denotes a noise prediction, unlike the velocity in our flow-matching slides.

这里的 epsilon 表示噪声预测，与流匹配页面中的速度不同。

**01:23:18** These controls belong to this formulation and should not be assumed to map directly onto every current editor's interface.

这些控制属于这一公式，不应假设它们直接对应所有当前编辑器的界面。


## 116 — What remains when the terms cancel? (01:23:26)

**01:23:26** Now set both guidance scales to one and follow the cancellation.

现在把两个引导强度都设为一，观察抵消。

**01:23:30** The initial no-condition prediction cancels its negative copy in the first difference.

初始无条件预测，与第一个差值中的负项抵消。

**01:23:36** The image-only prediction then cancels its negative copy in the second difference.

随后，仅图像条件的预测，与第二个差值中的负项抵消。

**01:23:41** The result is the prediction with both image and text conditions.

最终得到同时使用图像和文本条件的预测。

**01:23:45** That gives us a reference point for understanding the two controls.

这为理解两个控制提供了参考点。

**01:23:49** If the text scale were zero while the image scale stayed one, we would instead recover the image-only prediction.

如果文本强度为零、图像强度仍为一，就会得到仅使用图像条件的预测。

**01:23:57** It shows what information each difference adds.

这说明了每个差值添加什么信息。

**01:24:00** Keep epsilon's meaning explicit here: this InstructPix2Pix expression predicts noise, whereas the earlier flow-matching example predicted velocity.

请明确 epsilon 的含义：这里的 InstructPix2Pix 公式预测噪声，而前面的流匹配例子预测速度。


## 117 — Training: the target supplies the lesson (01:24:11)

**01:24:11** Read the pseudocode in two groups.

把伪代码分成两组来读。

**01:24:14** First we prepare an example: source, instruction, target, conditions, and target latent.

首先准备样本：原图、指令、目标、条件，以及目标潜变量。

**01:24:23** Then we sample noise and time, construct the intermediate latent, predict velocity, and compare that prediction with the known training target.

然后采样噪声和时间，构造中间潜变量，预测速度，并将预测与已知训练目标比较。

**01:24:32** The final update changes trainable weights based on the loss.

最后根据损失更新可训练权重。

**01:24:37** At inference, there is no known edited target to supply in this way.

推理时，并没有已知编辑目标可以这样提供。

**01:24:41** We instead follow the learned generative predictions from a starting state.

我们从初始状态出发，遵循学习得到的生成预测。

**01:24:46** This pseudocode explains the conceptual loop; it omits practical engineering such as batches, precision, and scheduling.

这段伪代码解释概念循环，省略了批处理、精度和调度等实际工程细节。

**01:24:55** Its most important distinction is what information is available during learning versus use.

最重要的区别，是学习与使用时分别有哪些信息可用。


## 118 — Inference: where did the target go? (01:25:01)

**01:25:01** Look for the line that disappeared.

请找出消失的那一行。

**01:25:03** The inference loop has no known edited target and no loss that updates weights.

推理循环没有已知编辑目标，也没有用于更新权重的损失。

**01:25:08** Instead, we encode the source and instruction, begin with noise, repeatedly predict a velocity and move the latent, then decode the result.

我们编码原图和指令，从噪声开始，反复预测速度并移动潜变量，最后解码结果。

**01:25:18** Training uses examples to adjust the learned model.

训练用样本调整学习到的模型。

**01:25:21** In this simplified inference loop, we use that model to construct an output we do not yet possess.

在这个简化推理循环中，我们用该模型构造尚未拥有的输出。

**01:25:28** The code is conceptual; practical systems use their own representations and sampling schedules.

代码是概念性的；实际系统使用各自的表示和采样调度。

**01:25:35** But you can now explain why providing another reference changes the information available for a request without automatically becoming a model-training operation.

但你现在可以解释：增加参考图会改变请求可用的信息，却不会自动变成模型训练。

**01:25:44** It is the same distinction we used when choosing between references and LoRA.

这与选择参考图还是 LoRA 时使用的区别相同。


## 119 — Choose the approach by the requirement (01:25:50)

**01:25:50** Before we turn to the readings' discussion questions, use this table as a compact decision aid.

转向阅读讨论问题之前，请把这张表当作简明决策辅助。

**01:25:56** Exact untouched pixels suggest a role for masks and compositing.

如果未编辑像素必须精确不变，可以考虑蒙版与合成。

**01:26:00** A new visual language may benefit from reference conditioning.

新的视觉语言，可能受益于参考图条件控制。

**01:26:04** A reusable learned concept may motivate adaptation.

可复用的学习概念，可能需要模型适配。

**01:26:08** A transformation that must survive motion needs a video workflow and temporal evidence.

必须经受运动的变换，需要视频工作流程和时间证据。

**01:26:14** These approaches can cooperate in the same artwork.

这些方法可以在同一件作品中协作。

**01:26:17** Choose according to the contract, then inspect the failure most likely to undermine it.

根据约定选择，再检查最可能破坏约定的失败。

**01:26:23** That leaves us with a larger question for the readings: once production becomes easier, how do we choose a worthwhile direction and retain meaningful judgment over the result?

这给阅读留下更大的问题：制作更容易之后，如何选择值得追求的方向，并对结果保留有意义的判断？


## 120 — An Alien Mind (01:26:34)

**01:26:34** An Alien Mind distinguishes achieving an assigned goal from generalizing human values in unfamiliar circumstances.

《An Alien Mind》区分了完成指定目标，与在陌生情境中泛化人类价值观。

**01:26:41** Pachocki also raises questions about monitoring increasingly capable systems.

Pachocki 也提出了如何监测日益强大系统的问题。

**01:26:47** Treat the essay as an argument containing claims and forecasts that we can examine.

请把文章视为包含主张与预测的论证，我们可以对其检视。

**01:26:52** For our discussion, imagine an exhibition team delegating production while retaining responsibility for what the work communicates.

讨论时，想象一个展览团队委托制作，但仍对作品传达的内容负责。

**01:27:00** Discuss the three questions on screen: explain the goal-and-values distinction, identify a forecast and the evidence it would need, and defend a boundary for human control.

讨论屏幕上的三个问题：解释目标与价值观的区别，找出一个预测及所需证据，并为人类控制的边界辩护。

**01:27:12** Pause the video here.

请在这里暂停视频。

**01:27:14** Our art-direction example is an analogy; it does not give an image editor the same agency or risk profile as an autonomous research system.

艺术指导只是类比；它并不赋予图像编辑器与自主研究系统相同的能动性或风险特征。


## 121 — How to fix your entire life in 1 day (01:27:27)

**01:27:27** Koe's title makes a dramatic promise: How to fix your entire life in one day.

Koe 的标题作出了强烈承诺：《如何在一天内修复你的整个生活》。

**01:27:32** The essay proposes examining identity and goals, interrupting habitual behavior, and turning reflection into action.

文章建议审视身份与目标、打断习惯行为，并把反思转化为行动。

**01:27:41** We can discuss the usefulness of that proposal without accepting the title as a guarantee.

我们可以讨论这些建议是否有用，而不必把标题当作保证。

**01:27:46** Imagine an artist who can generate a hundred attractive images but cannot choose what to make.

想象一个艺术家，能生成一百张好看的图像，却无法决定要创作什么。

**01:27:52** Does easy production help clarify a direction, or make avoiding the decision easier?

轻松制作会帮助明确方向，还是让逃避决定更容易？

**01:27:58** Discuss the three questions on screen.

请讨论屏幕上的三个问题。

**01:28:01** Choose an idea worth using or challenging, explain the reason, and connect it to creative intention.

选一个值得采用或质疑的观点，解释理由，并把它与创作意图联系起来。

**01:28:08** Pause the video here.

请在这里暂停视频。

**01:28:10** You can use a fictional artist or public example; no personal disclosure is needed.

可以使用虚构艺术家或公开案例，不需要披露个人经历。


## 122 — Introducing Aleph 2.0 and Edit Studio (01:28:20)

**01:28:20** Runway's reading proposes approving an edited image before applying its appearance through a video.

Runway 的阅读提出：先批准编辑图像，再把它的外观应用到整段视频。

**01:28:26** It offers a concrete answer to a communication problem: an art director can point to the desired look, rather than describe every feature in words.

它具体回答了一个沟通问题：艺术总监可以指出目标外观，而不必用文字描述每个特征。

**01:28:35** Now bring back the column.

现在，让柱子重新出现。

**01:28:37** A convincing keyframe does not tell us what the pear will look like after it reappears.

可信关键帧并不能告诉我们，梨再次显露后会是什么样子。

**01:28:42** Discuss the three questions on screen: what the frame establishes, which claim deserves a harder test, and how the workflow serves AFTER RAIN's intention.

讨论屏幕上的三个问题：这一帧确定了什么，哪个主张需要更严格测试，以及流程如何服务于《雨后》的意图。

**01:28:52** Pause the video.

请暂停视频。

**01:28:54** The source is a vendor announcement; our job is to distinguish a useful demonstrated workflow from a reliability claim still needing evaluation.

来源是厂商公告；我们的任务是区分有用的已演示流程，与仍需评价的可靠性主张。


## 123 — Your final judgment (01:29:07)

**01:29:07** Return to your first judgment of the glass pear.

回到你最初对玻璃梨的判断。

**01:29:11** We began with a small request and discovered that it touched the scene's physics, the artwork's intention, and the audience's experience over time.

我们从一个小要求出发，发现它触及场景物理、作品意图，以及观众随时间展开的体验。

**01:29:20** You now have more precise ways to say what should change, what should survive, and how to judge the result.

现在，你有更精确的方法说明什么应改变、什么应保留，以及如何判断结果。

**01:29:27** Finish with the three questions on screen.

最后，请讨论屏幕上的三个问题。

**01:29:30** Where would you place the boundary between control and surprise?

你会把控制与惊喜之间的边界放在哪里？

**01:29:33** How would you balance technical success and artistic purpose?

你会如何平衡技术成功与艺术目的？

**01:29:37** What would you delegate, and what evidence would you require?

你会委托什么，又要求什么证据？

**01:29:41** Pause the video for the final discussion.

请暂停视频，进行最终讨论。

**01:29:44** Connect one mechanism or reading to a concrete artistic decision.

把一种机制或一篇阅读，与具体艺术决定联系起来。

**01:29:49** Listen for an answer that makes you revise your own.

留意一个能让你修正自己看法的回答。

**01:29:52** That revision is a fitting last act for a class about editing.

对一堂关于编辑的课来说，这样的修正，是恰当的最后一笔。
