1
00:00:00,000 --> 00:00:01,825
Imagine you are the art director.
想象一下，你是艺术总监。

2
00:00:02,059 --> 00:00:03,913
Your exhibition opens tomorrow.
你的展览明天开幕。

3
00:00:04,425 --> 00:00:07,695
You love this glass pear, but you want to see it in ivory ceramic.
你很喜欢这只玻璃梨，但想看看它变成象牙白陶瓷的样子。

4
00:00:08,279 --> 00:00:12,762
You give the editor a tiny request: change the material, keep everything else.
你给编辑器一个小小的要求：改变材质，其他一切保持不变。

5
00:00:13,347 --> 00:00:14,980
Then you notice the reflection.
然后，你注意到了倒影。

6
00:00:15,653 --> 00:00:17,492
Should it still look like glass?
它还应该像玻璃一样吗？

7
00:00:18,413 --> 00:00:22,019
Suddenly, a three-word edit contains a whole argument about the world.
突然，一个简单的编辑指令，包含了关于世界如何运作的一整套判断。

8
00:00:22,224 --> 00:00:24,064
That is our starting point.
这就是我们的起点。

9
00:00:24,064 --> 00:00:30,561
We will follow this fictional artwork through image editing, exhibition design, and a moving shot.
我们将跟随这件虚构作品，经历图像编辑、展览设计，再进入一段运动镜头。

10
00:00:31,233 --> 00:00:33,775
The exhibition is called AFTER RAIN.
展览名叫《雨后》，英文是 AFTER RAIN。

11
00:00:34,577 --> 00:00:36,958
The pear is our recurring character.
这只梨会反复登场。

12
00:00:37,630 --> 00:00:46,127
By the end, you should be able to explain a model's mechanism, defend a visual choice, and catch a failure that a beautiful preview can hide.
到课程结束，你应该能够解释模型机制、为视觉选择提供理由，并发现精美预览可能掩盖的失败。

13
00:00:46,798 --> 00:00:49,997
First, look carefully at what you think must survive.
首先，仔细看看你认为必须保留的东西。

14
00:00:50,597 --> 00:00:53,429
We are going to give the same object three different jobs.
我们要让同一个物体承担三种不同的任务。

15
00:00:53,941 --> 00:00:59,256
First it is a technical puzzle: can its material change while its identity survives?
首先，它是一个技术难题：能否改变材质，同时保留它的身份特征？

16
00:00:59,767 --> 00:01:05,243
Then it becomes an artwork: can its placement and lighting make us feel that it is fragile?
接着，它成为艺术作品：能否通过摆放和灯光，让我们感到它的脆弱？

17
00:01:06,045 --> 00:01:12,193
Finally it becomes a character in time: can it disappear behind something and return as the same object?
最后，它成为时间中的角色：能否在被遮挡后重新出现，仍然是同一个物体？

18
00:01:12,777 --> 00:01:15,244
Those jobs demand different judgments.
这些任务需要不同的判断。

19
00:01:15,245 --> 00:01:17,931
A technically clean image may say very little.
技术上干净的图像，可能表达得很少。

20
00:01:18,516 --> 00:01:22,517
An expressive poster may contain an intentional impossibility.
富有表现力的海报，可能故意包含不可能的景象。

21
00:01:23,101 --> 00:01:26,123
A wonderful still may belong to a broken video.
一张精彩的静帧，可能来自一段存在问题的视频。

22
00:01:26,708 --> 00:01:33,293
As its job changes, you may find yourself approving a change that you would have rejected five minutes earlier.
随着任务改变，你可能会接受一个五分钟前还会拒绝的变化。

23
00:01:33,893 --> 00:01:35,864
For a few seconds, ignore the pear.
先暂时别看梨。

24
00:01:36,288 --> 00:01:37,748
Look at the water beneath it.
看看它下面的水。

25
00:01:38,331 --> 00:01:39,864
Now look back at the body.
现在再看回梨的主体。

26
00:01:40,449 --> 00:01:45,356
If the body becomes opaque ceramic, what happens to the light that used to pass through the glass?
如果主体变成不透明的陶瓷，原先穿过玻璃的光会怎样？

27
00:01:46,027 --> 00:01:49,283
The answer cannot live entirely inside the object's outline.
答案不可能完全局限在物体轮廓之内。

28
00:01:49,868 --> 00:01:53,401
This comparison is a prepared classroom illustration.
这组对比是事先准备的课堂示意图。

29
00:01:53,401 --> 00:01:59,825
Use it to locate the consequences of the request: transmission, highlights, and the reflection below.
用它来找出指令引发的影响：透光、高光，以及下方的倒影。

30
00:02:00,410 --> 00:02:03,461
We still want the blue stem and recognizable silhouette.
我们仍然希望保留蓝色的柄和可辨认的轮廓。

31
00:02:04,046 --> 00:02:13,610
But preserving every surrounding pixel could preserve the wrong optics. a successful local edit may require a carefully justified change somewhere else.
但保留周围的每一个像素，可能会保留错误的光学效果。成功的局部编辑，有时需要在别处做出有充分理由的改变。

32
00:02:14,210 --> 00:02:16,969
Let us make the opening comparison more disciplined.
让我们更严谨地审视开场的对比。

33
00:02:17,554 --> 00:02:24,358
The disappearance of transmitted amber light is consistent with changing transparent glass into opaque ceramic.
从透明玻璃变为不透明陶瓷后，透过的琥珀色光消失，是合理的。

34
00:02:25,030 --> 00:02:32,960
A changed blue stem, by contrast, would need a separate justification because the instruction did not request that transformation.
相比之下，如果蓝色的柄也变了，就需要另外解释，因为指令并没有要求这种变化。

35
00:02:33,543 --> 00:02:35,938
The reflection is more interesting.
倒影则更有意思。

36
00:02:35,938 --> 00:02:43,137
A material change can require its appearance to change, while its placement should still agree with the object and the water.
材质改变可能要求倒影外观随之变化，但它的位置仍应与物体和水面一致。

37
00:02:44,056 --> 00:02:48,510
We cannot classify every difference using a simple rule that change is bad.
我们不能用“变化就是坏事”这一简单规则来判断每个差异。

38
00:02:49,095 --> 00:02:51,459
We need a model of the intended scene.
我们需要理解目标场景应当如何运作。

39
00:02:51,971 --> 00:02:57,256
That is why the edit contract includes allowed consequences as well as invariants.
因此，编辑约定既要包含不变量，也要包含允许发生的连带变化。

40
00:02:57,856 --> 00:02:59,798
Here is the route through that puzzle.
下面是我们破解这个难题的路线。

41
00:03:00,383 --> 00:03:08,604
We begin inside an image editor: its representations, sampling process, and controls.
我们先进入图像编辑器内部，研究它的表示方式、采样过程和控制手段。

42
00:03:09,406 --> 00:03:13,948
Then we take the art director's seat and decide what those controls should accomplish.
然后坐到艺术总监的位置，决定这些控制手段应该实现什么。

43
00:03:14,532 --> 00:03:20,563
In the final technical section, the image starts moving and the preservation problem becomes harder.
在最后一个技术部分，图像开始运动，保留原有信息的问题也变得更难。

44
00:03:21,147 --> 00:03:24,315
The spoken script is planned for about ninety minutes.
讲解稿按约九十分钟规划。

45
00:03:24,315 --> 00:03:29,688
The marked discussions, the break, and student presentations need additional class time.
标出的讨论、课间休息和学生展示，需要额外的课堂时间。

46
00:03:30,273 --> 00:03:33,792
You will see the same pear return in different situations.
你会看到同一只梨出现在不同情境中。

47
00:03:34,464 --> 00:03:40,494
Each return asks us to revise our judgment, rather than learn a completely unrelated example.
每次重逢，都要求我们修正判断，而不是学习一个毫不相关的新例子。

48
00:03:41,298 --> 00:03:47,212
At the end, three readings let us question who sets the goal and who decides whether it was achieved.
最后三篇阅读，将让我们追问：谁来设定目标，又由谁判断目标是否达成？

49
00:03:47,811 --> 00:03:51,257
Here are three decisions I want you to be able to defend by the end.
到课程结束，我希望你能为这三类决定提供理由。

50
00:03:51,841 --> 00:03:56,280
When an edit fails, which control would you change, and why?
编辑失败时，你会调整哪一种控制手段，为什么？

51
00:03:56,864 --> 00:04:01,390
When two posters both look good, which belongs in this exhibition?
两张海报都很好看时，哪一张更适合这个展览？

52
00:04:01,814 --> 00:04:06,648
When a video looks convincing, which moment would you inspect before accepting it?
视频看起来可信时，你会在接受它之前检查哪个时刻？

53
00:04:07,231 --> 00:04:10,473
You do need a link between mechanism and consequence.
你确实需要把机制与结果联系起来。

54
00:04:10,473 --> 00:04:17,306
A mask tells us where; a reference can show us what; an artistic brief tells us why the result matters.
蒙版告诉我们改哪里；参考图可以展示改成什么；艺术创作说明告诉我们结果为何重要。

55
00:04:17,979 --> 00:04:20,533
Today we will practice making those links aloud.
今天，我们会练习把这些联系说清楚。

56
00:04:21,337 --> 00:04:26,009
The aim is to leave with reasons you can use when the next tool has a different interface.
目标是让你带走一套理由，即使下一款工具换了界面，也能继续使用。

57
00:04:26,609 --> 00:04:33,546
The pear gives us a concrete technical problem: change its material while retaining the information that makes it recognizable.
这只梨给我们一个具体的技术问题：改变材质，同时保留让人认出它的信息。

58
00:04:34,129 --> 00:04:38,744
We can describe this as conditional generation under preservation constraints.
我们可以把它描述为：在保留约束下进行条件生成。

59
00:04:39,415 --> 00:04:44,030
The inputs specify the change; the contract specifies what should survive.
输入指定要改变什么；编辑约定指定要保留什么。

60
00:04:44,613 --> 00:04:48,716
Throughout this section, connect each mechanism to one part of that problem.
在这一部分，请把每种机制对应到这个问题的某一环节。

61
00:04:49,388 --> 00:04:52,601
We begin by making our expectations explicit.
我们先把预期说清楚。

62
00:04:53,201 --> 00:04:56,530
Before we open the machinery, decide what success should look like.
在打开技术黑箱之前，先确定什么才算成功。

63
00:04:57,041 --> 00:05:00,049
When glass becomes ceramic, what should stay the same?
玻璃变成陶瓷时，哪些东西应该保持不变？

64
00:05:00,414 --> 00:05:03,159
Which parts of the scene must change with the material?
场景的哪些部分必须随材质一起改变？

65
00:05:03,743 --> 00:05:06,561
What could a model misunderstand in that request?
模型可能会如何误解这个要求？

66
00:05:07,146 --> 00:05:11,482
Discuss these three questions and point to visible evidence for your choices.
请讨论这三个问题，并指出支持你们选择的可见证据。

67
00:05:11,482 --> 00:05:18,622
If your partner wants to preserve a detail that you want to change, identify the artistic intention behind each answer.
如果同伴想保留的细节恰好是你想改变的，请找出两种答案背后的艺术意图。

68
00:05:19,207 --> 00:05:20,579
Pause the video here.
请在这里暂停视频。

69
00:05:21,163 --> 00:05:26,084
We will return to these expectations when we write the worked ceramic-edit contract.
等到我们为陶瓷编辑写出完整约定时，再回头检视这些预期。

70
00:05:30,084 --> 00:05:32,187
The gallery has become a forest.
画廊变成了森林。

71
00:05:32,611 --> 00:05:34,567
Let us write the edit contract.
让我们写下编辑约定。

72
00:05:34,655 --> 00:05:37,094
The intended change is the environment.
我们想改变的是环境。

73
00:05:37,604 --> 00:05:42,992
The pear's identity, its pedestal, and the framing should remain recognizable.
梨的身份特征、底座和构图，都应该仍然可辨认。

74
00:05:43,577 --> 00:05:47,737
The ambient light and reflections are allowed to adapt to the forest.
环境光和反射则允许适应森林。

75
00:05:48,321 --> 00:05:50,322
Why include that last category?
为什么要加入最后这一类？

76
00:05:50,643 --> 00:05:55,199
Because preserving every visible relationship would contradict the new setting.
因为保留所有可见关系，会与新环境矛盾。

77
00:05:55,200 --> 00:05:59,405
Imagine retaining a bright gallery-window reflection inside a dark forest.
想象一下，在昏暗森林里还保留着明亮画廊窗户的反光。

78
00:05:59,989 --> 00:06:04,502
The model might preserve the source faithfully and still produce an implausible scene.
模型可能忠实保留了原图，却仍然生成了不合理的场景。

79
00:06:05,173 --> 00:06:11,918
Before generating, separate the requested change, the invariants, and the consequences that should adapt.
生成之前，要区分所要求的变化、不变量，以及应当随之调整的连带效果。

80
00:06:12,503 --> 00:06:15,642
Afterwards, inspect those categories separately.
生成之后，再分别检查这些类别。

81
00:06:16,242 --> 00:06:20,155
Suppose a curator says, keep the pear exactly the same.
假设策展人说：让这只梨完全保持原样。

82
00:06:20,740 --> 00:06:23,660
There are at least two meanings hiding in that sentence.
这句话至少隐藏着两种含义。

83
00:06:24,244 --> 00:06:28,538
One means that a visitor should recognize the same sculpture in a new room.
一种是：观众在新房间里，仍然能认出这是同一件雕塑。

84
00:06:29,121 --> 00:06:31,341
Its highlights may change with the room.
它的高光可以随着房间而变化。

85
00:06:31,852 --> 00:06:35,561
The other means that selected pixel values must remain identical.
另一种是：选定像素的数值必须完全一致。

86
00:06:36,071 --> 00:06:37,780
Those are different requirements.
这是两种不同的要求。

87
00:06:38,101 --> 00:06:41,942
Generative conditioning can encourage recognizable identity.
生成式条件控制可以帮助保留可辨认的身份特征。

88
00:06:41,942 --> 00:06:47,389
Copying source pixels or compositing can enforce exact preservation in a specified region.
复制原图像素或进行合成，可以强制指定区域完全不变。

89
00:06:47,972 --> 00:06:54,631
Neither requirement automatically makes the whole image coherent: a copied region may now have the wrong lighting.
但这两种要求都不会自动保证整幅图像协调：复制过来的区域，光照可能已经不合适。

90
00:06:55,141 --> 00:06:58,412
Before choosing a tool, settle the meaning of same.
选择工具之前，先确定“相同”到底是什么意思。

91
00:06:59,013 --> 00:07:06,956
Read the expression as a distribution of possible edited results, given the source, instruction, and optional references.
把这个表达式理解为：给定原图、指令和可选参考信息后，可能产生的编辑结果的分布。

92
00:07:07,759 --> 00:07:11,030
The symbol theta represents the model's learned parameters.
符号 theta 表示模型学习到的参数。

93
00:07:11,454 --> 00:07:17,366
The vertical bar means 'given.' We are not asking for a unique answer in every case.
竖线表示“给定”。我们并不是在所有情况下都要求唯一答案。

94
00:07:17,878 --> 00:07:21,557
Two different forest scenes might both satisfy the same brief.
两个不同的森林场景，都可能满足同一份创作说明。

95
00:07:21,558 --> 00:07:27,370
However, a source image is a condition rather than a promise that every unmentioned pixel will be copied.
但原图只是条件，并不意味着每一个未提及的像素都会被复制。

96
00:07:27,953 --> 00:07:32,597
To judge success, we need a more precise contract than 'it looks plausible.'
要判断是否成功，我们需要比“看起来合理”更精确的约定。

97
00:07:33,197 --> 00:07:38,132
The probability expression says that the system produces a candidate given conditions.
概率表达式表示：系统在给定条件下生成一个候选结果。

98
00:07:38,644 --> 00:07:42,732
It does not contain a certificate that the candidate passes our checks.
它并不附带证明，保证这个结果通过我们的检查。

99
00:07:43,316 --> 00:07:47,287
The acceptance step is a separate decision that we impose on the result.
是否接受，是我们对结果做出的另一个独立判断。

100
00:07:47,960 --> 00:07:50,019
Imagine two forest outputs.
想象两个森林版本。

101
00:07:50,178 --> 00:07:54,691
Both look plausible, but one changes the pear's stem and one retains it.
两个都看起来合理，但一个改变了梨柄，另一个保留了它。

102
00:07:54,691 --> 00:07:59,217
The model can assign probability to both; our contract can reject one.
模型可以给两者都分配概率，而我们的约定可以拒绝其中一个。

103
00:08:00,021 --> 00:08:06,489
In practice, acceptance may involve human inspection, image comparisons, or explicit production constraints.
实际上，验收可能依赖人工检查、图像对比，或明确的制作约束。

104
00:08:07,089 --> 00:08:14,960
The RGB image has 1024 by 1024 pixels and three channels: just over three million scalar values.
RGB 图像有 1024 乘 1024 个像素，每个像素有三个通道，总共略多于三百万个标量值。

105
00:08:15,383 --> 00:08:28,948
With spatial reduction by sixteen and sixty-four latent channels, the representation becomes 64 by 64 by 64: 262,144 values.
空间尺寸缩小十六倍、潜在通道数为六十四时，表示变成 64 乘 64 乘 64，即 262,144 个数值。

106
00:08:29,459 --> 00:08:30,540
Notice the trap.
注意这里的陷阱。

107
00:08:31,212 --> 00:08:39,447
Sixteen times smaller along each spatial axis does not mean sixteen times fewer values overall, because the number of channels changes too.
每个空间轴缩小十六倍，并不意味着总数值量减少十六倍，因为通道数也变了。

108
00:08:39,447 --> 00:08:42,425
Here the scalar count falls by a factor of twelve.
在这个例子中，标量数量减少了十二倍。

109
00:08:43,010 --> 00:08:45,303
That is specific to this configuration.
这个比例只适用于这一具体配置。

110
00:08:45,580 --> 00:08:51,376
The model works in a learned representation, and its decoder must recover visible detail.
模型在学习得到的表示中工作，解码器必须从中恢复可见细节。

111
00:08:52,297 --> 00:08:57,378
Our next question is what that compression already loses before we request any edit.
接下来要问的是：在我们提出任何编辑要求之前，压缩已经丢失了什么？

112
00:08:57,978 --> 00:09:02,023
Suppose the small exhibition caption is already blurry after an edit.
假设编辑后，展览的小字号说明文字变模糊了。

113
00:09:02,695 --> 00:09:07,396
You might spend an hour rewriting the prompt to say, preserve every letter.
你可能花一个小时重写提示词，强调要保留每一个字母。

114
00:09:08,068 --> 00:09:14,785
There is a quicker diagnostic: encode the original image and decode it again without requesting a change.
有一个更快的诊断方法：不要求任何改变，先把原图编码，再解码回来。

115
00:09:15,589 --> 00:09:23,035
If the letters are already damaged in that round trip, part of the problem lies in representation and reconstruction.
如果这次往返已经损坏了文字，那么问题的一部分就在表示和重建环节。

116
00:09:23,035 --> 00:09:28,335
A more emphatic instruction cannot restore a detail the editing path never retained faithfully.
再强硬的指令，也无法恢复编辑流程从未忠实保留的细节。

117
00:09:29,007 --> 00:09:32,892
Look at fine text, thin edges, and small textures first.
首先检查小字、细线边缘和微小纹理。

118
00:09:33,475 --> 00:09:38,936
If reconstruction is sound but the edit damages them, investigate the editing stage.
如果重建没有问题，而编辑损坏了它们，就调查编辑阶段。

119
00:09:39,520 --> 00:09:44,645
It also explains why exact exhibition typography deserves its own editable layer.
这也说明，要求精确的展览文字排版，应当保留在独立的可编辑图层中。

120
00:09:45,246 --> 00:09:48,750
Now consider the number of tokens a transformer processes.
现在考虑 Transformer 需要处理多少个 token。

121
00:09:49,262 --> 00:09:58,198
In our toy example, one token per position on a 64 by 64 grid gives 4,096 tokens.
在这个简化例子里，64 乘 64 的网格上，每个位置对应一个 token，共有 4,096 个。

122
00:09:58,781 --> 00:10:05,353
Grouping each two-by-two region gives 1,024 tokens, a reduction by four.
把每个二乘二区域组合起来，就得到 1,024 个 token，数量减少四倍。

123
00:10:06,155 --> 00:10:09,121
Dense self-attention compares token pairs.
稠密自注意力会比较 token 对。

124
00:10:09,704 --> 00:10:21,313
Squaring the two counts gives 16,777,216 and 1,048,576 pair scores.
将两个数量分别平方，得到 16,777,216 和 1,048,576 个配对分数。

125
00:10:21,737 --> 00:10:24,087
That is a reduction by sixteen.
数量减少了十六倍。

126
00:10:24,087 --> 00:10:27,635
This is why representation choices matter so much for cost.
这就是为什么表示方式会如此显著地影响成本。

127
00:10:28,220 --> 00:10:31,812
It is not a claim that the whole system runs sixteen times faster.
但这不等于整个系统的运行速度会提高十六倍。

128
00:10:32,395 --> 00:10:38,280
Text tokens, reference tokens, other layers, and implementation details also matter.
文本 token、参考图 token、其他网络层和实现细节，也会影响结果。

129
00:10:38,880 --> 00:10:41,756
Before reading the second number, make a prediction.
看第二个数字之前，先预测一下。

130
00:10:42,180 --> 00:10:46,415
We double the image width and double its height, keeping the patch scheme fixed.
我们把图像宽度和高度都加倍，同时保持分块方式不变。

131
00:10:47,087 --> 00:10:49,028
How many image tokens do we get?
图像 token 会变成多少？

132
00:10:49,539 --> 00:10:50,912
Four times as many.
原来的四倍。

133
00:10:51,496 --> 00:10:55,585
In dense self-attention, each token can compare with every other token.
在稠密自注意力中，每个 token 都可以与其他所有 token 比较。

134
00:10:56,095 --> 00:10:59,615
Four times four gives sixteen times as many pair scores.
四乘四，配对分数的数量就变成十六倍。

135
00:10:59,615 --> 00:11:07,616
This calculation describes the image-token attention component, not a promise that the entire system becomes sixteen times slower.
这个计算只描述图像 token 的注意力部分，并不保证整个系统会慢十六倍。

136
00:11:08,537 --> 00:11:11,325
Other operations and implementations matter.
其他操作和实现方式同样重要。

137
00:11:12,129 --> 00:11:18,700
The useful habit is to ask what grew: pixels, tokens, comparisons, or measured runtime.
一个有用的习惯是追问：增加的究竟是像素、token、比较次数，还是实测运行时间？

138
00:11:19,284 --> 00:11:22,803
They are related quantities, but they are not interchangeable.
这些量相互关联，但不能互换。

139
00:11:23,403 --> 00:11:26,571
Let us walk through this architecture from the input side.
让我们从输入端走一遍这个架构。

140
00:11:27,156 --> 00:11:34,821
The source image contributes semantic information through a vision-language component and visual information through a VAE.
原图通过视觉语言组件提供语义信息，并通过 VAE 提供视觉信息。

141
00:11:35,405 --> 00:11:39,216
The target stream begins with an evolving noisy representation.
目标分支从一个不断演化的带噪表示开始。

142
00:11:39,800 --> 00:11:43,596
The transformer processes that stream together with the conditions.
Transformer 将这个分支与条件信息一起处理。

143
00:11:44,181 --> 00:11:51,029
During training, we possess an example of the desired edited target, so we can measure what the model should learn.
训练时，我们拥有所需编辑结果的样本，因此可以衡量模型应该学习什么。

144
00:11:51,029 --> 00:11:55,235
During inference, we have the source and request but not the desired target.
推理时，我们有原图和要求，却没有所需的目标图像。

145
00:11:55,819 --> 00:11:57,542
The model must generate it.
模型必须把它生成出来。

146
00:11:57,703 --> 00:11:59,499
That difference is essential.
这个区别至关重要。

147
00:12:00,418 --> 00:12:05,689
The architecture does not quietly receive the answer when you ask it to edit your photograph.
当你让模型编辑照片时，架构并没有偷偷收到答案。

148
00:12:06,290 --> 00:12:12,232
There is a piece of information on the left that we do not possess on the right: the desired edited image.
左边有一项信息，是右边没有的：期望得到的编辑图像。

149
00:12:12,817 --> 00:12:19,957
During supervised training, a source, an instruction, and a target show the model what a successful change looks like.
在监督训练中，原图、指令和目标图一起，向模型展示成功的变化是什么样子。

150
00:12:20,629 --> 00:12:25,520
During use, we provide the request because that target does not yet exist.
使用时，我们之所以提出要求，正是因为目标图像还不存在。

151
00:12:26,032 --> 00:12:29,828
This distinction explains a surprisingly common confusion.
这个区别解释了一个相当常见的混淆。

152
00:12:29,828 --> 00:12:33,829
Showing a reference to an editor is not automatically teaching new weights.
给编辑器看一张参考图，并不自动意味着学习新的权重。

153
00:12:34,339 --> 00:12:37,553
It may simply be conditioning this particular generation.
它可能只是在为这一次生成提供条件。

154
00:12:38,136 --> 00:12:43,116
Adaptation, which we will reach with LoRA, changes trainable parameters.
适配则会改变可训练参数，讲到 LoRA 时我们会展开。

155
00:12:43,539 --> 00:12:50,212
It is available when we construct the learning signal, and absent when we ask the trained model to make our ceramic pear.
构造学习信号时，目标图像是已知的；要求训练好的模型生成陶瓷梨时，它则是未知的。

156
00:12:50,812 --> 00:12:53,675
Think of these two representations using our pear.
用这只梨来理解这两种表示。

157
00:12:54,346 --> 00:13:04,129
Semantic features help with questions such as: which object is the pear, what does ceramic mean, and what does 'it' refer to in the instruction?
语义特征帮助回答：哪个物体是梨，陶瓷是什么意思，以及指令中的“它”指什么？

158
00:13:05,049 --> 00:13:12,423
Visual latents provide appearance information, including shapes and textures that help keep the particular source recognizable.
视觉潜变量提供外观信息，包括形状和纹理，帮助保留原图的独特特征。

159
00:13:13,343 --> 00:13:20,308
This is a conceptual distinction, not a claim that the branches have perfectly isolated responsibilities.
这是一种概念上的区分，并不是说两个分支的职责完全隔离。

160
00:13:20,308 --> 00:13:23,549
Learned representations can overlap in what they encode.
学习到的表示所编码的信息可能重叠。

161
00:13:24,134 --> 00:13:31,113
If an edit fails, ask whether the model misunderstood the request or lost important visual information.
编辑失败时，要问：模型是误解了要求，还是丢失了重要的视觉信息？

162
00:13:31,713 --> 00:13:35,626
How do you teach a system the difference between a pear and this particular pear?
如何教会系统区分“一只梨”和“这只特定的梨”？

163
00:13:36,708 --> 00:13:42,577
That question connects to OpenSubject, research I coauthored with Yexin Liu and our collaborators.
这个问题与 OpenSubject 有关，这是我与Yexin Liu及其他合作者共同完成的研究。

164
00:13:43,161 --> 00:13:49,499
A video offers something valuable: repeated observations of a subject as the view changes.
视频提供了一种宝贵资源：随着视角变化，对同一主体进行反复观察。

165
00:13:49,922 --> 00:13:52,580
The shared identity is a learning opportunity.
共同的身份特征提供了学习机会。

166
00:13:53,252 --> 00:13:55,661
Follow the main path in the figure.
请沿着图中的主路径看。

167
00:13:55,661 --> 00:14:01,194
We curate clips, verify subjects across frames, and select diverse pairs.
我们筛选视频片段，跨帧验证主体，并挑选具有多样性的图像对。

168
00:14:01,706 --> 00:14:07,590
Inpainting or outpainting helps synthesize reference inputs, followed by verification.
通过图像修补或扩图来合成参考输入，再进行验证。

169
00:14:08,175 --> 00:14:14,775
The corpus contains 2.5 million samples; the contribution is training data and a benchmark.
数据集包含二百五十万个样本；这项工作的贡献是训练数据和评测基准。

170
00:14:15,359 --> 00:14:19,622
For our exhibition, imagine the same sculpture photographed in several rooms.
对于我们的展览，可以想象在不同房间拍摄同一件雕塑。

171
00:14:20,206 --> 00:14:24,513
We want freedom to change the setting without losing its distinctive features.
我们希望自由改变环境，同时不丢失它的独特特征。

172
00:14:24,514 --> 00:14:31,757
That is a classroom application of the identity problem, rather than a claim that our fictional pear was tested in the paper.
这是身份保持问题的课堂应用，并不意味着论文测试过这只虚构的梨。

173
00:14:32,356 --> 00:14:36,080
The editor needs a way to learn how to move from noise toward an image.
编辑器需要学习如何从噪声走向图像。

174
00:14:36,664 --> 00:14:46,638
In this simple flow-matching setup, we construct an intermediate latent by mixing noise, epsilon, with a known target latent, z one.
在这个简单的流匹配设定中，我们将噪声 epsilon 与已知目标潜变量 z₁ 混合，构造中间潜变量。

175
00:14:47,557 --> 00:14:51,573
The mixing variable t tells us where we are along that training path.
混合变量 t 告诉我们处于这条训练路径的什么位置。

176
00:14:52,492 --> 00:14:57,880
For this straight interpolation, the target velocity is the target latent minus the noise.
对于这种直线插值，目标速度等于目标潜变量减去噪声。

177
00:14:57,881 --> 00:15:03,443
The model sees an intermediate state and learns to predict that direction, given the source and instruction.
模型看到中间状态，并在原图和指令的条件下，学习预测这个方向。

178
00:15:04,115 --> 00:15:09,255
There is also an important clock distinction: t here is the generative process's time.
还要区分两种时钟：这里的 t 是生成过程中的时间。

179
00:15:09,840 --> 00:15:15,111
Later, our video will have another time axis—the seconds that pass in the scene.
稍后的视频还有另一条时间轴：场景中实际流逝的秒数。

180
00:15:15,711 --> 00:15:18,996
Before trusting an equation, check its easiest cases.
相信一个方程之前，先检查最简单的情况。

181
00:15:19,581 --> 00:15:25,874
At t equals zero, the coefficient on the target becomes zero, so we recover the noise.
当 t 等于零时，目标的系数为零，所以我们得到噪声。

182
00:15:26,546 --> 00:15:33,423
At t equals one, the coefficient on noise vanishes, so we recover the target latent.
当 t 等于一时，噪声的系数消失，所以我们得到目标潜变量。

183
00:15:34,344 --> 00:15:41,293
Along this simple straight path, the difference between target and noise gives the direction we train the model to predict.
沿着这条简单直线路径，目标减去噪声，就是我们训练模型预测的方向。

184
00:15:41,293 --> 00:15:46,915
You do not need to imagine a recognizable half-finished image at every intermediate point.
不必想象每个中间点都对应一张能辨认的半成品图像。

185
00:15:47,587 --> 00:15:56,947
The calculation takes place in a learned representation, and this path is a teaching construction rather than the only possible design.
计算发生在学习得到的表示中；这条路径用于教学，并不是唯一可能的设计。

186
00:15:57,547 --> 00:16:00,569
Here is the smallest possible sampling calculation.
这是一个最小的采样计算例子。

187
00:16:01,081 --> 00:16:04,249
One coordinate currently equals 0.20.
当前某个坐标值是 0.20。

188
00:16:04,832 --> 00:16:11,506
Its predicted velocity is 0.60, and the next step has size 0.10.
预测速度为 0.60，下一步的步长为 0.10。

189
00:16:12,762 --> 00:16:16,909
Before looking at the result, say what operation should happen.
看结果之前，先说说应该做什么运算。

190
00:16:17,492 --> 00:16:27,261
We add a small amount of motion to the current state: 0.20 plus 0.10 times 0.60, giving 0.26.
在当前状态上加一小步：0.20 加上 0.10 乘 0.60，得到 0.26。

191
00:16:27,685 --> 00:16:31,189
In a real latent, many coordinates update together.
在真实的潜变量中，很多坐标会同时更新。

192
00:16:31,189 --> 00:16:36,957
These numbers are invented to show an Euler step, and practical samplers can use more elaborate solvers.
这些数字是为了演示欧拉步而编造的，实际采样器可以采用更复杂的求解器。

193
00:16:37,541 --> 00:16:44,974
But the sequence is now less mysterious: predict a direction, take a step, repeat, then decode.
但这个过程现在不再那么神秘：预测方向，迈出一步，重复，然后解码。

194
00:16:45,574 --> 00:16:49,954
Imagine following a beautifully accurate set of directions to the wrong gallery.
想象你沿着极其精确的路线指引，却走向了错误的画廊。

195
00:16:50,378 --> 00:16:53,561
Taking smaller steps will not repair the destination.
把步子迈小一点，并不能修正目的地。

196
00:16:54,145 --> 00:16:56,948
The same distinction helps with generative sampling.
同样的区分也有助于理解生成采样。

197
00:16:57,620 --> 00:17:01,519
A better numerical solver can follow a learned field more accurately.
更好的数值求解器，可以更准确地沿着学到的向量场前进。

198
00:17:02,104 --> 00:17:05,623
That alone does not settle an ambiguous editing request.
但仅凭这一点，无法解决含糊的编辑要求。

199
00:17:06,294 --> 00:17:13,317
If the pear's boundary is unstable across sampling settings, numerical or model behavior may matter.
如果不同采样设置下梨的边界不稳定，数值计算或模型行为可能有影响。

200
00:17:13,317 --> 00:17:19,246
If we never specified whether the glass reflection should become ceramic, we also have a brief problem.
如果我们从未说明玻璃倒影是否应变成陶瓷倒影，那么创作说明本身也有问题。

201
00:17:20,048 --> 00:17:26,678
Ask whether the system struggled to realize a clear instruction, or whether we have not agreed on what success means.
要问：系统难以实现一个清楚的指令，还是我们尚未就成功的含义达成一致？

202
00:17:27,263 --> 00:17:30,431
Those situations call for different next actions.
这两种情况需要不同的下一步行动。

203
00:17:31,031 --> 00:17:34,185
Where does the ability to follow an edit instruction come from?
遵循编辑指令的能力从何而来？

204
00:17:35,104 --> 00:17:41,398
A training example can contain a source image, an instruction, and the corresponding edited target.
一个训练样本可以包含原图、指令，以及对应的编辑目标图像。

205
00:17:42,202 --> 00:17:48,202
InstructPix2Pix is an early example that used synthetic editing data to teach this relationship.
InstructPix2Pix 是一个较早的例子，它用合成编辑数据来学习这种关系。

206
00:17:48,874 --> 00:17:56,320
Suppose a training pair says 'make the object ceramic,' but its target also moves the camera and replaces the background.
假设某个训练对的指令是“把物体变成陶瓷”，但目标图还移动了相机并替换了背景。

207
00:17:56,321 --> 00:18:00,921
The learning signal no longer cleanly identifies the requested transformation.
这时，学习信号就不再能清晰地指向所要求的变换。

208
00:18:01,505 --> 00:18:04,484
The model may learn unwanted associations.
模型可能学到不希望出现的关联。

209
00:18:05,155 --> 00:18:13,274
This gives us an important data-quality question: did the target accomplish the requested change while preserving what should remain?
因此，一个重要的数据质量问题是：目标图是否完成了要求的变化，同时保留了应该保留的内容？

210
00:18:14,077 --> 00:18:18,224
Attractive targets alone are not enough to teach dependable editing.
仅有好看的目标图，还不足以教会可靠的编辑。

211
00:18:18,824 --> 00:18:22,869
A training pair is a lesson, and lessons can accidentally teach the wrong thing.
一个训练对就是一堂课，而课程可能无意中教错东西。

212
00:18:23,452 --> 00:18:28,491
Imagine a source glass pear and a target ceramic pear that also has a different background.
想象原图是一只玻璃梨，目标图是一只陶瓷梨，但背景也变了。

213
00:18:29,162 --> 00:18:32,096
The written request says only to change the material.
文字要求却只说改变材质。

214
00:18:32,608 --> 00:18:36,404
Which visible changes should the model associate with that request?
模型应当把哪些可见变化与这个要求联系起来？

215
00:18:36,405 --> 00:18:44,304
If this mismatch recurs in the training data, the model can learn an unwanted association between a material change and a scene change.
如果训练数据中反复出现这种不匹配，模型可能把材质变化与场景变化错误地联系起来。

216
00:18:44,976 --> 00:18:48,787
Compare the instruction with the actual difference between source and target.
请将指令与原图、目标图之间的实际差异进行比较。

217
00:18:49,458 --> 00:18:53,913
Ask what supervision is rewarding, including changes nobody meant to label.
要问监督信号在奖励什么，包括那些没人打算标注的变化。

218
00:18:54,497 --> 00:18:58,775
The preservation contract begins in the examples used to teach the editor.
保留约定，应当从用来教会编辑器的样本开始。

219
00:18:59,375 --> 00:19:03,565
Sometimes a conditioned prediction moves in the right direction but too weakly.
有时，条件预测的方向是对的，但力度不足。

220
00:19:04,149 --> 00:19:11,465
Classifier-free guidance uses the difference between a conditioned prediction and a baseline prediction to steer the result.
无分类器引导利用条件预测与基线预测之间的差异，来引导结果。

221
00:19:12,137 --> 00:19:16,093
Read the formula as baseline, plus a scaled change in direction.
可以把公式读成：基线，加上经过缩放的方向变化。

222
00:19:16,678 --> 00:19:22,212
When s equals one, the baseline terms cancel and we recover the conditioned prediction.
当 s 等于一时，基线项抵消，我们得到条件预测。

223
00:19:22,796 --> 00:19:25,891
Above one, we extrapolate beyond it.
大于一时，我们会外推到条件预测之外。

224
00:19:25,892 --> 00:19:30,111
That can strengthen the requested change, but it can strengthen errors as well.
这可以强化所要求的变化，但也可能强化错误。

225
00:19:30,783 --> 00:19:38,596
In an editing system, the baseline may still retain image information; baseline does not always mean no information at all.
在编辑系统中，基线可能仍保留图像信息；基线并不总意味着完全没有信息。

226
00:19:39,398 --> 00:19:45,692
Think of guidance as a specific operation on predictions, rather than a universal quality dial.
把引导理解为对预测值的一种具体运算，而不是万能的质量旋钮。

227
00:19:46,292 --> 00:19:51,958
The baseline predicts 0.2 and the conditioned version predicts 0.5.
基线预测是 0.2，条件预测是 0.5。

228
00:19:52,381 --> 00:19:56,776
With guidance scale two, do we get a value somewhere between them?
引导强度设为二，结果会落在两者之间吗？

229
00:19:57,448 --> 00:19:58,062
No.
不会。

230
00:19:58,572 --> 00:20:05,027
We take 0.2 plus twice the difference of 0.3, which gives 0.8.
我们用 0.2 加上两倍的差值 0.3，得到 0.8。

231
00:20:05,537 --> 00:20:08,093
We have moved beyond the conditioned prediction.
我们已经越过了条件预测。

232
00:20:08,517 --> 00:20:13,043
If the useful direction includes a slight mistake, we can amplify both.
如果有效方向中包含一点错误，那么两者都可能被放大。

233
00:20:13,043 --> 00:20:19,175
For the pear, the new material might become clearer while the stem or boundary becomes less faithful.
对于这只梨，新材质可能更明显，但柄或边界可能变得不那么忠实。

234
00:20:19,979 --> 00:20:23,454
Compare the requested change and preservation separately.
请分别比较所要求的变化和保留效果。

235
00:20:24,038 --> 00:20:31,266
One slider can move those judgments in opposite directions; a single overall impression can hide the tradeoff.
同一个滑块可能让两项评价朝相反方向变化；单一的总体印象会掩盖这种取舍。

236
00:20:31,866 --> 00:20:35,093
Different controls communicate different kinds of information.
不同控制手段传达不同类型的信息。

237
00:20:35,765 --> 00:20:38,086
Text describes a requested change.
文本描述所要求的变化。

238
00:20:38,758 --> 00:20:40,685
A mask specifies a region.
蒙版指定区域。

239
00:20:41,197 --> 00:20:45,606
A spatial condition can describe pose, depth, or edges.
空间条件可以描述姿态、深度或边缘。

240
00:20:46,190 --> 00:20:51,797
An image reference can supply appearance information that would be difficult to express precisely in words.
参考图可以提供难以用语言精确表达的外观信息。

241
00:20:52,309 --> 00:20:56,601
These are not interchangeable knobs, and every product does not expose all of them.
这些控制并不能互相替代，也不是每款产品都提供全部选项。

242
00:20:57,186 --> 00:21:01,829
ControlNet and IP-Adapter are examples of distinct architectural approaches.
ControlNet 和 IP-Adapter 就是两种不同架构思路的例子。

243
00:21:01,829 --> 00:21:11,394
If exact untouched pixels are essential, a generation mask alone may be insufficient; explicit copying or compositing can enforce that requirement.
如果必须精确保留未编辑像素，仅有生成蒙版可能不够；明确复制像素或合成，可以强制实现这一要求。

244
00:21:12,065 --> 00:21:18,066
Choose a control by asking what information is missing, then check whether the resulting image actually respected it.
选择控制手段时，先问缺少什么信息，再检查结果是否真正遵循了它。

245
00:21:18,738 --> 00:21:25,659
A mask indicates a permitted region; it does not by itself solve seams or the reflection of a changed material.
蒙版指出允许编辑的区域，但它本身不能解决接缝，也不能处理材质改变后的倒影。

246
00:21:25,659 --> 00:21:29,806
The next slide names an architecture that learns to use a spatial condition.
下一页介绍一种学习使用空间条件的架构。

247
00:21:30,406 --> 00:21:34,027
Suppose your sentence is understood, but the pear keeps changing shape.
假设模型理解了你的话，但梨的形状总在变化。

248
00:21:34,612 --> 00:21:39,079
You can describe its outline with more adjectives, or supply spatial evidence.
你可以用更多形容词描述轮廓，也可以直接提供空间证据。

249
00:21:39,883 --> 00:21:45,695
ControlNet gives an edge map, depth map, or pose a learned route into a compatible diffusion model.
ControlNet 为边缘图、深度图或姿态信息，提供了一条进入兼容扩散模型的可学习路径。

250
00:21:46,278 --> 00:21:48,936
Trace the two paths in this original architecture.
请追踪原始架构中的两条路径。

251
00:21:49,141 --> 00:21:51,550
The pretrained path stays frozen.
预训练主路径保持冻结。

252
00:21:51,550 --> 00:21:59,975
A trainable copy processes the added condition, and zero-initialized one-by-one convolutions connect its features to the main path.
一个可训练副本处理新增条件，零初始化的一乘一卷积把它的特征连接到主路径。

253
00:22:00,559 --> 00:22:05,816
Initially those connections contribute zero; training learns their contribution.
初始时，这些连接的贡献为零；训练会学出它们应有的贡献。

254
00:22:06,326 --> 00:22:10,605
The copied branch itself starts from pretrained weights, not all zeros.
副本分支本身从预训练权重开始，而不是全部从零开始。

255
00:22:11,189 --> 00:22:12,942
Now return to our request.
现在回到我们的要求。

256
00:22:13,526 --> 00:22:16,139
Edges can help specify the outline.
边缘可以帮助指定轮廓。

257
00:22:16,139 --> 00:22:23,732
They cannot, by themselves, tell us whether the reflected ceramic looks convincing or the blue stem retains its identity.
但仅凭边缘，无法判断陶瓷倒影是否可信，也无法保证蓝色梨柄保留身份特征。

258
00:22:24,332 --> 00:22:29,881
With several references, the system must know which reference contributes which information.
当有多张参考图时，系统必须知道每张图分别提供什么信息。

259
00:22:30,684 --> 00:22:36,262
Imagine asking for the sculpture from image one and the atmosphere from image two.
想象你要求使用图一的雕塑，以及图二的氛围。

260
00:22:37,065 --> 00:22:42,526
If the references become confused, you may get the wrong object with the right lighting.
如果参考图的角色混淆，可能会得到光照正确、物体却错误的结果。

261
00:22:43,110 --> 00:22:48,980
The illustrated research approach uses separators and image-index information to distinguish inputs.
图示研究方法使用分隔符和图像索引信息来区分输入。

262
00:22:48,980 --> 00:22:54,704
In an art brief, name the role of each image rather than presenting a pile of vaguely related inspiration.
在艺术创作说明中，要说明每张图的角色，而不是堆放一批关系模糊的灵感图。

263
00:22:55,376 --> 00:23:01,042
Then inspect for leakage: did a composition reference accidentally replace the subject?
然后检查信息是否串用：构图参考是否意外替换了主体？

264
00:23:02,122 --> 00:23:07,846
This figure describes one proposed mechanism, not a universal design used by every editor.
这张图描述的是一种提出的机制，并不是所有编辑器通用的设计。

265
00:23:08,446 --> 00:23:12,549
Look at these published multi-image examples with reference roles in mind.
请带着“参考图角色”的概念，观察这些已发表的多图示例。

266
00:23:13,133 --> 00:23:17,835
Before judging the output, identify what each input was supposed to contribute.
评价输出之前，先确定每个输入原本应该贡献什么。

267
00:23:18,420 --> 00:23:22,566
Then trace the subject and the requested transformation into the result.
然后在结果中追踪主体和所要求的变换。

268
00:23:23,370 --> 00:23:26,699
A reference-role failure can look superficially attractive.
参考图角色混淆的结果，表面上也可能很好看。

269
00:23:27,283 --> 00:23:33,094
The system may borrow the wrong object's appearance or import a background that was never intended.
系统可能借用了错误物体的外观，或引入了原本不需要的背景。

270
00:23:33,094 --> 00:23:39,723
We are examining qualitative examples from a particular paper, not conducting a broad comparison between products.
我们正在查看某篇论文的定性示例，并不是对产品做全面比较。

271
00:23:40,395 --> 00:23:49,039
The figure helps us practice an inspection method: follow the intended contribution of each input and look for unintended transfers between them.
这张图帮助我们练习一种检查方法：追踪每个输入的预期贡献，寻找意外的信息转移。

272
00:23:49,639 --> 00:23:54,458
Attention lets a token gather information from other tokens with different weights.
注意力让一个 token 以不同权重，从其他 token 收集信息。

273
00:23:55,261 --> 00:24:04,387
Our simplified scalar example gives weight 0.8 to a value of 0.9, and weight 0.2 to a value of 0.1.
在这个简化的标量例子中，数值 0.9 的权重为 0.8，数值 0.1 的权重为 0.2。

274
00:24:05,191 --> 00:24:08,125
The weighted result is 0.74.
加权结果是 0.74。

275
00:24:08,797 --> 00:24:17,558
The real mechanism operates on learned vectors, so this is arithmetic intuition rather than a literal description of artistic decision-making.
真实机制处理的是学习到的向量，因此这里只是在建立算术直觉，并非字面描述艺术决策。

276
00:24:17,558 --> 00:24:22,990
The important structure is selective combination: every available source need not contribute equally.
关键结构是选择性组合：不是每个可用来源都必须贡献同样多的信息。

277
00:24:23,662 --> 00:24:28,583
In a multi-reference edit, the system also needs to distinguish where information came from.
在多参考图编辑中，系统还需要区分信息来自哪里。

278
00:24:29,386 --> 00:24:36,249
Otherwise, gathering information successfully can still produce the wrong mixture of subject identity and visual style.
否则，即使成功收集了信息，也可能错误地混合主体身份与视觉风格。

279
00:24:36,849 --> 00:24:40,441
The weights in this simplified attention example sum to one.
在这个简化注意力例子中，权重之和为一。

280
00:24:41,024 --> 00:24:44,777
That makes the result a weighted combination of the values.
因此，结果是各个数值的加权组合。

281
00:24:45,362 --> 00:24:50,574
A larger weight makes the associated value contribute more to this particular calculation.
权重越大，对应数值在这次计算中的贡献就越大。

282
00:24:51,495 --> 00:24:56,605
Now be cautious about jumping from the calculation to an explanation of the final image.
但要谨慎，不能直接从这个计算跳到对最终图像的解释。

283
00:24:57,408 --> 00:25:01,847
Real networks contain many layers, heads, and transformations.
真实网络包含许多层、注意力头和变换。

284
00:25:01,847 --> 00:25:06,358
A high weight at one point may not tell us why a visible feature ultimately appeared.
某处权重很高，未必能解释某个可见特征最终为何出现。

285
00:25:06,943 --> 00:25:17,617
For practical reference editing, the useful question remains whether the intended information survives in the output, regardless of how compelling an attention visualization looks.
在实际参考图编辑中，关键仍是预期信息是否出现在输出中，无论注意力可视化看起来多么有说服力。

286
00:25:18,217 --> 00:25:21,648
What if a useful concept must recur across many requests?
如果一个有用的概念需要在许多请求中反复出现，怎么办？

287
00:25:22,320 --> 00:25:28,876
A reference can condition a generation, while LoRA adapts selected learned weights using a low-rank update.
参考图可以为一次生成提供条件；LoRA 则通过低秩更新，调整选定的模型权重。

288
00:25:29,461 --> 00:25:32,483
The update is a product of two smaller matrices.
更新量是两个较小矩阵的乘积。

289
00:25:33,067 --> 00:25:41,230
Instead of learning every entry of a large update matrix independently, we express that update as a product of two smaller matrices.
我们不独立学习大型更新矩阵的每个元素，而是用两个较小矩阵的乘积来表示更新。

290
00:25:41,230 --> 00:25:52,108
For a 4096 by 4096 weight matrix and rank sixteen, the full matrix has 16,777,216 entries.
对于一个 4096 乘 4096、秩设为十六的例子，完整矩阵有 16,777,216 个元素。

291
00:25:52,663 --> 00:26:00,548
The two factors together have 131,072, a factor of 128 fewer for this update.
两个因子合计只有 131,072 个元素，这次更新的参数量减少了一百二十八倍。

292
00:26:01,191 --> 00:26:05,090
That is not a claim that the entire model shrinks by 128.
这并不意味着整个模型缩小了一百二十八倍。

293
00:26:05,644 --> 00:26:07,572
The base weights still exist.
基础权重仍然存在。

294
00:26:08,346 --> 00:26:14,931
Low rank makes adaptation economical, but does not certify that the learned subject survives a new pose.
低秩让适配更经济，但不能证明学到的主体在新姿态下仍能保持一致。

295
00:26:14,931 --> 00:26:19,137
That question requires examples the adaptation did not already see.
这个问题需要用适配时未见过的样例来检验。

296
00:26:19,736 --> 00:26:23,503
An adapted model reproduces your favorite training portrait perfectly.
适配后的模型完美复现了你最喜欢的训练肖像。

297
00:26:24,015 --> 00:26:26,527
Is that enough to use the character in a new film?
这就足以让这个角色出演新电影了吗？

298
00:26:27,110 --> 00:26:35,098
Consider what else it may have learned: the familiar camera angle, the background, even the lighting that always accompanied the subject.
想一想它还可能学到了什么：熟悉的机位、背景，甚至总与主体一起出现的光照。

299
00:26:35,769 --> 00:26:39,229
A held-out check deliberately changes those circumstances.
留出测试会有意改变这些条件。

300
00:26:39,901 --> 00:26:45,085
Ask for a new view or a different setting and inspect the distinctive features.
要求一个新视角或不同环境，并检查独特特征。

301
00:26:45,085 --> 00:26:50,838
We want a reusable concept, rather than a narrow ability to reproduce familiar combinations.
我们需要的是可复用的概念，而不只是复现熟悉组合的狭窄能力。

302
00:26:51,350 --> 00:26:56,621
For the pear, keep the blue stem recognizable while moving the exhibition outdoors.
对于这只梨，要在把展览移到户外时，仍保留可辨认的蓝色柄。

303
00:26:57,204 --> 00:27:02,433
If identity collapses there, more praise for the training examples will not solve the problem.
如果身份特征在那里崩溃，再称赞训练样例也解决不了问题。

304
00:27:03,032 --> 00:27:06,902
A reward tells a learning process what kinds of outputs to favor.
奖励告诉学习过程应当偏好哪些输出。

305
00:27:07,486 --> 00:27:13,560
The figure shows a recent approach that separates task-specific reward training and then distills what is learned.
图中是一种近期方法：分别训练任务特定的奖励，再蒸馏学到的能力。

306
00:27:14,363 --> 00:27:20,219
Look at the categories: editing quality involves more than a single judgment of attractiveness.
看这些类别：编辑质量不只是单一的美观判断。

307
00:27:21,022 --> 00:27:25,285
Now imagine an artwork whose purpose is to feel awkward or disturbing.
现在想象一件艺术作品，其目的就是让人感到别扭或不安。

308
00:27:25,285 --> 00:27:29,155
A generic preference for polished images could work against that purpose.
对精致图像的普遍偏好，可能反而违背这个目的。

309
00:27:29,827 --> 00:27:35,302
Human preference signals are useful, but they do not define artistic merit for every project.
人类偏好信号很有用，但不能为所有项目定义艺术价值。

310
00:27:35,814 --> 00:27:45,611
This is the bridge to the next section: technical optimization can help produce candidates, while an artist still needs to decide which properties serve the work.
这正是通向下一部分的桥梁：技术优化可以帮助生成候选作品，而艺术家仍要决定哪些属性服务于作品。

311
00:27:46,211 --> 00:27:50,051
These edit examples invite a question about evaluation.
这些编辑示例引出了一个评价问题。

312
00:27:50,563 --> 00:27:53,205
Which properties would you reward separately?
你会分别奖励哪些属性？

313
00:27:53,790 --> 00:28:01,149
You might ask whether the instruction was carried out, whether identity survived, and whether the image is visually convincing.
可以问：指令是否完成、身份是否保留，以及图像在视觉上是否可信。

314
00:28:01,734 --> 00:28:03,676
Those judgments can disagree.
这些判断可能彼此冲突。

315
00:28:04,259 --> 00:28:09,458
A highly polished output might erase an awkward feature that is central to the artwork.
高度精致的输出，可能抹掉作品中至关重要的别扭特征。

316
00:28:09,458 --> 00:28:14,598
An unusual composition might serve the brief while attracting a lower generic preference score.
不寻常的构图可能符合创作要求，却得到较低的通用偏好分数。

317
00:28:15,269 --> 00:28:20,293
We should understand what a reward encourages before treating a high score as artistic approval.
在把高分视为艺术认可之前，我们应当理解奖励究竟鼓励什么。

318
00:28:20,876 --> 00:28:30,076
The published examples illustrate the research setting; our classroom task is to articulate the intention against which we would judge a particular result.
已发表示例展示的是研究情境；课堂任务则是明确创作意图，并据此判断具体结果。

319
00:28:30,675 --> 00:28:32,865
Which failure would you notice first?
你会先注意到哪种失败？

320
00:28:33,669 --> 00:28:40,254
On one side, the output is beautiful, but it has quietly replaced our particular sculpture with a generic decorative pear.
一边的结果很美，却悄悄把我们的特定雕塑换成了普通的装饰梨。

321
00:28:40,839 --> 00:28:45,949
On the other, the right sculpture is present, but its reflection still behaves like the old material.
另一边保留了正确雕塑，但倒影仍像原来的材质。

322
00:28:46,870 --> 00:28:48,971
The first may win a quick aesthetic vote.
第一种可能赢得快速的审美投票。

323
00:28:49,001 --> 00:28:53,922
The second may pass a checklist that only asks whether the requested object changed.
第二种可能通过只检查“目标物体是否改变”的清单。

324
00:28:53,922 --> 00:28:56,009
Neither satisfies the full brief.
两者都没有满足完整要求。

325
00:28:56,681 --> 00:29:01,909
Has the editor lost identity, broken the scene's optical logic, or missed the work's intention?
编辑器是丢失了身份、破坏了场景光学逻辑，还是偏离了作品意图？

326
00:29:02,712 --> 00:29:05,296
Naming the mismatch gives us a next action.
说清不匹配之处，才能决定下一步。

327
00:29:05,968 --> 00:29:10,801
Saying only that an output looks a little wrong leaves the diagnosis unfinished.
只说输出“看起来有点不对”，诊断还没有完成。

328
00:29:11,401 --> 00:29:14,876
A diagnosis becomes useful when it changes the next action.
只有能改变下一步行动的诊断，才有用。

329
00:29:15,548 --> 00:29:21,375
If the wrong object changes, first suspect an ambiguous reference or region specification.
如果改错了物体，首先怀疑参考对象或区域描述有歧义。

330
00:29:22,177 --> 00:29:27,391
If the correct object becomes another instance, appearance grounding may be the issue.
如果改对了物体，却变成另一个个体，问题可能在外观信息的约束。

331
00:29:28,062 --> 00:29:33,567
If the material looks disconnected from its reflection, the dependent change may be missing.
如果材质与倒影脱节，可能缺少必要的连带变化。

332
00:29:34,151 --> 00:29:38,211
These are candidate explanations, not automatic conclusions.
这些只是候选解释，不是自动成立的结论。

333
00:29:38,211 --> 00:29:40,620
Choose a small revision that tests one of them.
选择一个小调整，检验其中一种解释。

334
00:29:40,678 --> 00:29:43,788
If it does not help, reconsider the hypothesis.
如果没有帮助，就重新考虑假设。

335
00:29:44,299 --> 00:29:54,390
This is more informative than changing the prompt, seed, reference, and model all at once, because then a successful result would leave us unsure which intervention mattered.
这比同时改提示词、随机种子、参考图和模型更有信息量；否则，即使成功，也不知道哪个改动起了作用。

336
00:29:54,989 --> 00:29:57,559
We now have several ways to communicate an edit.
现在，我们有了几种表达编辑要求的方法。

337
00:29:58,144 --> 00:30:01,911
Would you preserve the reflection or let it change, and why?
你会保留倒影，还是允许它变化？为什么？

338
00:30:02,714 --> 00:30:05,372
When would a mask help more than a longer prompt?
什么时候蒙版比更长的提示词更有帮助？

339
00:30:06,043 --> 00:30:09,386
What would edges or depth control still leave uncertain?
边缘或深度控制仍会留下哪些不确定性？

340
00:30:10,307 --> 00:30:13,899
Discuss these three questions using the pear already on screen.
请用屏幕上的这只梨，讨论这三个问题。

341
00:30:14,484 --> 00:30:19,754
For each control you favor, name the failure it addresses and something it cannot settle.
对于你支持的每种控制手段，指出它解决什么失败，以及它不能决定什么。

342
00:30:19,754 --> 00:30:21,084
Pause the video here.
请在这里暂停视频。

343
00:30:21,594 --> 00:30:29,567
The next slide offers one possible contract; compare it with your reasoning rather than treating it as the only artistic answer.
下一页给出一种可能的约定；请与你的推理比较，不要把它当成唯一的艺术答案。

344
00:30:33,567 --> 00:30:36,618
Here is one defensible answer to the ceramic puzzle.
对于陶瓷难题，这是一种有理有据的答案。

345
00:30:37,203 --> 00:30:40,415
Preserve the blue stem and recognizable silhouette.
保留蓝色柄和可辨认的轮廓。

346
00:30:40,999 --> 00:30:46,373
Allow transmission, highlights, and the reflection to change with the new material.
允许透光、高光和倒影随着新材质改变。

347
00:30:47,175 --> 00:30:52,782
Then inspect the boundary and water, because that is where the request reaches beyond the object.
然后检查边界和水面，因为要求的影响会在那里超出物体本身。

348
00:30:53,367 --> 00:30:58,682
Notice the wording: allow a justified consequence, rather than permit arbitrary drift.
注意措辞：允许有理由的连带变化，而不是允许任意漂移。

349
00:30:58,682 --> 00:31:02,011
We have not given the model permission to redesign the gallery.
我们没有授权模型重新设计画廊。

350
00:31:02,595 --> 00:31:07,092
We have named the changes necessary to make this material transformation coherent.
我们明确指出了，让这次材质变换协调一致所必需的变化。

351
00:31:07,764 --> 00:31:14,510
A mask could help localize work, while compositing might protect a region that truly must remain exact.
蒙版可以帮助限定工作区域；合成则可以保护真正必须完全不变的区域。

352
00:31:15,314 --> 00:31:18,365
The contract tells us how to use those tools.
编辑约定告诉我们如何使用这些工具。

353
00:31:18,965 --> 00:31:21,901
We can now explain our opening puzzle at three levels.
现在，我们可以在三个层面解释开场难题。

354
00:31:22,412 --> 00:31:25,697
The brief decides what should change and what should survive.
创作说明决定什么应改变、什么应保留。

355
00:31:26,281 --> 00:31:31,188
The model uses representations, conditioning, and sampling to propose a result.
模型通过表示、条件控制和采样，提出候选结果。

356
00:31:31,859 --> 00:31:35,451
Our review checks whether that proposal actually satisfies the brief.
我们的审查检验这个候选结果是否真正满足要求。

357
00:31:36,123 --> 00:31:38,986
Keep those levels separate when you diagnose a failure.
诊断失败时，要区分这三个层面。

358
00:31:39,788 --> 00:31:43,380
Stronger guidance will not choose the exhibition's purpose.
更强的引导不会替你选择展览的目的。

359
00:31:43,380 --> 00:31:47,643
A more eloquent artistic statement will not enforce identical pixels.
更动人的艺术陈述不会强制像素完全一致。

360
00:31:48,154 --> 00:31:50,944
A beautiful preview will not prove preservation.
精美预览不会证明保留成功。

361
00:31:51,454 --> 00:31:55,733
The useful skill is connecting the right intervention to the observed problem.
有用的能力，是把正确的干预对应到观察到的问题。

362
00:31:56,244 --> 00:32:02,712
If your partner can tell when to use it and what it leaves uncertain, you have understood more than its name.
如果同伴能说清何时使用它、它还留下什么不确定性，你就不只是记住了它的名字。

363
00:32:03,313 --> 00:32:06,277
We can now explain our controls rather than just name them.
现在，我们能够解释控制手段，而不只是说出名称。

364
00:32:06,861 --> 00:32:10,205
How does ControlNet differ from a mask or a text prompt?
ControlNet 与蒙版或文本提示词有什么区别？

365
00:32:10,788 --> 00:32:13,476
Why can stronger guidance make an edit worse?
为什么更强的引导可能让编辑变差？

366
00:32:14,060 --> 00:32:17,141
What evidence would show that an edit preserved identity?
什么证据能够表明编辑保留了身份特征？

367
00:32:18,061 --> 00:32:19,828
Discuss these three questions.
请讨论这三个问题。

368
00:32:20,002 --> 00:32:27,902
Choose a concrete requirement—a silhouette, an untouched region, or a distinctive stem—and connect your explanation to it.
选一个具体要求——轮廓、未编辑区域，或独特的梨柄——并把解释与它联系起来。

369
00:32:27,902 --> 00:32:32,429
Then consider whether the control guarantees that requirement or only helps express it.
然后考虑：这种控制是保证实现要求，还是仅仅帮助表达要求？

370
00:32:32,939 --> 00:32:35,320
Pause the video here before the break.
休息之前，请在这里暂停视频。

371
00:32:39,320 --> 00:32:42,532
We will take ten minutes and resume when the class is ready.
我们休息十分钟，等大家准备好再继续。

372
00:32:43,117 --> 00:32:48,315
When you come back, think about a series of images you would recognize as belonging to one artwork.
回来时，请想一组你能认出属于同一件艺术作品的图像。

373
00:32:48,899 --> 00:32:56,259
What makes them belong together: the subject, the palette, the treatment of space, or something else?
是什么让它们属于一体：主体、色彩、空间处理，还是别的因素？

374
00:32:56,842 --> 00:32:59,120
Leave that question open for now.
暂时先保留这个问题。

375
00:32:59,120 --> 00:33:04,552
We will use it to move from individual editing operations to a coherent visual language.
我们将用它，从单次编辑操作转向连贯的视觉语言。

376
00:33:08,552 --> 00:33:12,100
Look again at the artwork we have been using as a technical test.
再看看这件一直被我们用来做技术测试的作品。

377
00:33:12,685 --> 00:33:14,948
After the break, it has a different job.
休息之后，它有了不同的任务。

378
00:33:15,459 --> 00:33:19,197
We are no longer asking only whether the edit obeys a request.
我们不再只问编辑是否遵循了要求。

379
00:33:19,781 --> 00:33:25,389
We are asking what a visitor might feel, and which visual decisions create that feeling.
我们要问：观众可能感受到什么，哪些视觉决定创造了这种感受？

380
00:33:26,061 --> 00:33:30,397
Two images can have different surfaces and still belong to one exhibition.
两张图像的表面可以不同，却仍属于同一个展览。

381
00:33:30,397 --> 00:33:33,361
Two images can share a palette and feel unrelated.
两张图像可以使用相同色彩，却感觉毫不相关。

382
00:33:33,945 --> 00:33:35,959
The difference is worth arguing about.
这种差别值得讨论。

383
00:33:36,631 --> 00:33:40,939
In the next section, I will show prepared alternatives for AFTER RAIN.
下一部分，我会展示为《雨后》准备的不同方案。

384
00:33:41,742 --> 00:33:45,831
Choose a direction in your mind, and be ready to explain a visible reason.
请在心里选择一个方向，并准备说出可见的理由。

385
00:33:46,415 --> 00:33:48,562
Your neighbor may choose the other one.
你的邻座可能会选另一个。

386
00:33:49,161 --> 00:33:50,913
Now take the art director’s seat.
现在，请坐到艺术总监的位置。

387
00:33:51,030 --> 00:33:57,163
The model can produce many plausible outputs; we decide which differences matter for AFTER RAIN.
模型能生成许多合理的结果；我们来决定哪些差异对《雨后》重要。

388
00:33:57,967 --> 00:34:02,975
As we compare the prepared images, name the visible relationship that carries the idea.
比较这些准备好的图像时，请指出承载创意的可见关系。

389
00:34:03,558 --> 00:34:07,924
That could be scale, light, or the treatment of space.
它可能是尺度、光线，或空间处理。

390
00:34:09,005 --> 00:34:13,342
Your explanation will be more useful than simply calling an image impressive.
这样的解释，比简单地说图像很惊艳更有用。

391
00:34:13,942 --> 00:34:18,250
Before seeing the poster alternatives, consider what you want a visitor to experience.
看海报方案之前，先想想你希望观众经历什么。

392
00:34:18,760 --> 00:34:22,222
What makes an image feel fragile rather than merely attractive?
什么让图像显得脆弱，而不仅仅是好看？

393
00:34:22,805 --> 00:34:26,470
Can very different images belong to the same artwork, and why?
差别很大的图像，能否属于同一件作品？为什么？

394
00:34:27,142 --> 00:34:30,121
Which artistic decisions would you keep for yourself?
哪些艺术决定，你会留给自己？

395
00:34:31,378 --> 00:34:33,086
Discuss these three questions.
请讨论这三个问题。

396
00:34:33,203 --> 00:34:37,831
You might disagree about the effect of empty space or the meaning of a material.
你们可能对留白的效果或材质的含义有不同看法。

397
00:34:37,831 --> 00:34:40,663
Locate the visual evidence behind that disagreement.
请找出分歧背后的视觉证据。

398
00:34:41,248 --> 00:34:42,591
Pause the video here.
请在这里暂停视频。

399
00:34:43,263 --> 00:34:47,584
Keep your starting position in mind when we compare the prepared directions.
比较预先准备的方向时，请记住你的初始立场。

400
00:34:51,585 --> 00:34:59,542
Here is our commission: AFTER RAIN, a fictional exhibition about fragile objects in changing environments.
这是我们的委托：《雨后》，一个关于变化环境中脆弱物体的虚构展览。

401
00:35:00,127 --> 00:35:06,728
Imagine a visitor seeing its poster across a corridor before encountering a five-second moving image inside.
想象观众先在走廊另一端看到海报，再进入展厅看到一段五秒的动态图像。

402
00:35:07,311 --> 00:35:09,662
What should that visitor expect to feel?
你希望这位观众期待怎样的感受？

403
00:35:10,172 --> 00:35:14,480
The prepared examples keep the amber pear and blue stem recognizable.
准备好的示例都保留了可辨认的琥珀色梨和蓝色柄。

404
00:35:14,992 --> 00:35:22,482
We will examine photographic and collage directions, then consider how a moving reveal changes the experience.
我们将研究摄影和拼贴两个方向，再考虑运动中的揭示如何改变体验。

405
00:35:22,483 --> 00:35:25,359
Nobody needs to generate an image during this lecture.
这节课不需要任何人现场生成图像。

406
00:35:25,944 --> 00:35:29,841
Fragile is a useful beginning, but it is not yet an art direction.
“脆弱”是一个有用的起点，但还不算艺术方向。

407
00:35:30,513 --> 00:35:36,384
We need to translate that word into something visible enough to compare and precise enough to revise.
我们需要把这个词转化为足够可见、能够比较，又足够精确、能够修改的东西。

408
00:35:36,984 --> 00:35:39,875
Try replacing fragile with expensive in the brief.
试着把创作说明中的“脆弱”换成“昂贵”。

409
00:35:40,459 --> 00:35:43,993
You might still choose glass, but would you use the same composition?
你可能仍然选择玻璃，但还会用同样的构图吗？

410
00:35:44,664 --> 00:35:48,986
A large centered object and assertive lighting might suggest a luxury product.
居中放大的物体配合强势光照，可能让人联想到奢侈品。

411
00:35:49,571 --> 00:35:53,951
A small object surrounded by quiet space might instead seem exposed.
安静空间中一个小小的物体，则可能显得无所庇护。

412
00:35:54,272 --> 00:35:56,594
Scale can make the pear feel vulnerable.
尺度可以让梨显得易受伤害。

413
00:35:57,265 --> 00:35:59,821
Restrained light can make us look closely.
克制的光线可以引导我们细看。

414
00:35:59,821 --> 00:36:03,954
A reflection can suggest a world less stable than the object itself.
倒影可以暗示一个比物体本身更不稳定的世界。

415
00:36:04,625 --> 00:36:12,086
These are artistic hypotheses to test against an audience's reading, not a formula that makes every image fragile.
这些是需要通过观众解读来检验的艺术假设，不是让所有图像显得脆弱的公式。

416
00:36:12,890 --> 00:36:14,949
Point to the feature doing the work.
请指出真正起作用的特征。

417
00:36:15,620 --> 00:36:20,672
If nobody can locate it in the image, the intention may still be living only in the prompt.
如果没人能在图像中找到它，意图可能仍然只存在于提示词里。

418
00:36:21,272 --> 00:36:24,208
In Hokusai's Great Wave, look first for Mount Fuji.
看葛饰北斋的《神奈川冲浪里》，先找富士山。

419
00:36:24,630 --> 00:36:26,106
It is small and distant.
它很小，也很远。

420
00:36:26,777 --> 00:36:29,815
Now follow the curves of the wave and the boats beneath it.
现在沿着巨浪的曲线，以及下方的小船看。

421
00:36:30,398 --> 00:36:35,698
Scale and rhythm create a tension we can discuss without turning the artwork into a style label.
尺度和节奏创造了张力；我们可以讨论这种张力，而不是把作品变成一个风格标签。

422
00:36:36,283 --> 00:36:43,116
For AFTER RAIN, we might borrow a relationship: a small, vulnerable form facing a much larger environment.
对《雨后》来说，我们可以借鉴一种关系：一个小而脆弱的形体，面对远大于自身的环境。

423
00:36:43,117 --> 00:36:47,643
We do not need to reproduce the wave or ask for a generic imitation of the artist.
不需要复制巨浪，也不需要要求泛泛地模仿这位艺术家。

424
00:36:48,227 --> 00:36:52,550
The reference becomes useful when we can say which visual decision we are studying.
当我们能够说明正在研究哪个视觉决定时，参考作品才真正有用。

425
00:36:53,133 --> 00:36:55,222
Then we must test its translation.
然后，还必须检验这种转化。

426
00:36:55,806 --> 00:37:00,171
Our quiet flooded gallery has a different subject and emotional register.
我们安静而积水的画廊，有着不同的主体和情感基调。

427
00:37:00,772 --> 00:37:05,868
Instead of using a reference as a label, describe a relationship you can observe.
不要把参考作品当作标签，而要描述你能观察到的关系。

428
00:37:06,452 --> 00:37:13,520
You might want a small stable form beneath a large dynamic curve, or a rhythm that moves the eye through the composition.
你可能想要一个巨大动态曲线下的小型稳定形体，或引导视线穿过构图的节奏。

429
00:37:14,192 --> 00:37:17,768
Then translate that relationship into the new subject and brief.
然后，把这种关系转化到新的主体和创作要求中。

430
00:37:18,440 --> 00:37:23,784
The result should be judged as its own artwork, not as a contest to resemble the reference.
应把结果作为独立作品评价，而不是比赛谁更像参考图。

431
00:37:23,784 --> 00:37:29,188
This approach makes references more useful to collaborators because they can understand what you are borrowing.
这种方法让参考资料对合作者更有用，因为他们能理解你借鉴的是什么。

432
00:37:29,771 --> 00:37:40,898
It also makes iteration more focused: if the intended tension is missing, you can revise scale or rhythm rather than vaguely asking for more influence from the source.
它也让迭代更聚焦：如果缺少预期张力，可以调整尺度或节奏，而不是含糊地要求“更受原作影响”。

433
00:37:41,498 --> 00:37:45,222
Before admiring the surface detail, decide where your eye arrives.
欣赏表面细节之前，先判断你的视线最先落在哪里。

434
00:37:45,806 --> 00:37:49,865
Is it the pear, the reflection, or the space waiting above them?
是梨、倒影，还是它们上方留出的空间？

435
00:37:50,668 --> 00:37:53,515
A poster has to organize that first encounter.
海报必须组织这第一次相遇。

436
00:37:54,187 --> 00:37:59,619
If every region is equally busy, the audience has to invent a hierarchy we have not provided.
如果每个区域都同样繁忙，观众就得自行建立我们没有提供的层次。

437
00:38:00,204 --> 00:38:05,211
Here, the open area can make the object feel small and give the title somewhere to live.
这里的开放区域能让物体显得小，也给标题留下位置。

438
00:38:05,211 --> 00:38:12,921
We can ruin the sense of quiet by filling every available gap with decorative detail, even if each addition is attractive on its own.
如果用装饰细节填满每一处空隙，就可能破坏安静感，即使每个新增细节单独看都很漂亮。

439
00:38:13,842 --> 00:38:19,886
When revising this image, I would first ask whether the composition expresses the intended scale and silence.
修改这张图时，我会先问：构图是否表达了预期的尺度感与寂静？

440
00:38:20,690 --> 00:38:23,668
More texture comes later, if the work needs it.
如果作品需要，之后再增加纹理。

441
00:38:24,268 --> 00:38:27,189
Negative space does two jobs in this poster.
留白在这张海报中承担两项任务。

442
00:38:27,699 --> 00:38:33,613
It changes how large or isolated the object feels, and it creates a place for typography.
它改变物体显得多大、多孤立，也为文字排版创造位置。

443
00:38:34,285 --> 00:38:38,534
These are connected design decisions rather than separate finishing steps.
这是相互关联的设计决定，而不是彼此独立的收尾步骤。

444
00:38:39,045 --> 00:38:47,047
If you fill the pale area with dramatic detail, the image may become more visually active but leave no calm route for reading the title.
如果在浅色区域填满戏剧性细节，画面可能更活跃，却没有平静的路径让人阅读标题。

445
00:38:47,047 --> 00:38:51,077
If you reserve too much space, the subject might lose the presence you wanted.
如果留白过多，主体又可能失去你想要的存在感。

446
00:38:51,662 --> 00:38:54,713
Evaluate the space in relation to the finished use.
请结合最终用途来评价空间。

447
00:38:55,298 --> 00:39:02,934
A generative background is part of a larger composition when it must carry text, branding, or other exact elements.
当背景需要承载文字、品牌或其他精确元素时，生成背景只是更大构图的一部分。

448
00:39:03,534 --> 00:39:06,293
The collage direction changes the rules of the world.
拼贴方向改变了这个世界的规则。

449
00:39:06,878 --> 00:39:09,360
Torn edges replace smooth contours.
撕裂边缘取代了平滑轮廓。

450
00:39:09,872 --> 00:39:12,120
Flat layers replace optical depth.
平面层次取代了光学纵深。

451
00:39:12,704 --> 00:39:16,515
Amber and indigo keep a connection to our recurring subject.
琥珀色和靛蓝色，让它仍与反复出现的主体保持联系。

452
00:39:17,318 --> 00:39:21,670
Look at how those decisions affect the pear's apparent weight and vulnerability.
看看这些决定如何影响梨看起来的重量与脆弱感。

453
00:39:22,254 --> 00:39:28,474
If the body looks like paper but the water behaves like a photograph, do we accept that collision?
如果主体像纸，水却像照片，我们是否接受这种碰撞？

454
00:39:28,474 --> 00:39:31,278
We might, if it is deliberate and supports the work.
如果它是有意的，并服务于作品，我们可能接受。

455
00:39:31,862 --> 00:39:34,549
We might reject it as an unresolved mixture.
我们也可能认为它是尚未解决的混杂，因而拒绝。

456
00:39:35,220 --> 00:39:38,009
The answer depends on the intended visual language.
答案取决于预期的视觉语言。

457
00:39:38,594 --> 00:39:44,011
The next comparison asks you to locate precisely where that language holds together or breaks.
下一组比较，请你精确指出：这种语言在哪里成立，又在哪里断裂。

458
00:39:44,611 --> 00:39:51,401
A photographic reflection inside a paper collage can be a mistake, or the most interesting decision in the image.
纸质拼贴中的摄影式倒影，可能是错误，也可能是整张图最有意思的决定。

459
00:39:51,984 --> 00:39:58,351
We need to know whether the mismatch is an intentional disruption and whether it produces the intended effect.
我们需要知道：这种不匹配是否是有意的打破，以及是否产生了预期效果。

460
00:39:59,153 --> 00:40:02,089
Suppose the exhibition explores unstable memories.
假设展览探索的是不稳定的记忆。

461
00:40:02,513 --> 00:40:05,695
An impossibly photographic reflection could make sense.
一个不可能的摄影式倒影，就可能说得通。

462
00:40:06,279 --> 00:40:10,601
Suppose the brief calls for a coherent world assembled from torn paper.
假设要求是构建一个由撕纸组成的统一世界。

463
00:40:10,602 --> 00:40:12,836
The same reflection might weaken it.
同样的倒影可能削弱它。

464
00:40:13,420 --> 00:40:17,509
This does not mean every accident deserves an explanation after the fact.
这并不意味着每个意外都值得事后找理由。

465
00:40:18,181 --> 00:40:24,780
Make the intention specific, examine what the viewer can actually see, and compare alternatives.
请明确意图，检查观众实际能看见什么，并比较不同方案。

466
00:40:25,452 --> 00:40:31,570
A strong critique can distinguish a productive contradiction from an excuse for an unresolved result.
有力的评论能区分富有成效的矛盾，与为未完成结果寻找的借口。

467
00:40:32,170 --> 00:40:35,689
A reference becomes easier to use when we assign it a role.
为参考图分配角色后，它就更容易使用。

468
00:40:36,200 --> 00:40:42,363
One image might define the subject, another the composition, and another the material treatment.
一张图可以定义主体，另一张定义构图，再一张定义材质处理。

469
00:40:43,034 --> 00:40:45,239
Name the property you want from each.
请说清你想从每张图中得到什么属性。

470
00:40:45,823 --> 00:40:48,831
The model may not isolate those roles perfectly.
模型未必能完美隔离这些角色。

471
00:40:49,415 --> 00:40:56,700
That is why you should look for unwanted transfers, such as importing a reference's background when you only wanted its palette.
因此，要检查意外的信息转移，例如只想借用配色，却把参考图的背景也引入了。

472
00:40:56,701 --> 00:41:00,046
Try removing one reference and observing what disappears.
试着移除一张参考图，观察什么随之消失。

473
00:41:00,629 --> 00:41:02,747
Start with the smallest useful set.
从最小但有用的参考集合开始。

474
00:41:03,331 --> 00:41:10,895
If five references contradict one another, adding a sixth may make the problem harder to diagnose rather than solve it.
如果五张参考图互相矛盾，再加第六张可能让问题更难诊断，而不是解决问题。

475
00:41:11,494 --> 00:41:13,991
We have added a reference because we think it helps.
我们添加参考图，是因为认为它有帮助。

476
00:41:14,503 --> 00:41:16,794
How would we discover whether it actually does?
怎样才能知道它是否真的有帮助？

477
00:41:16,882 --> 00:41:18,328
Remove it and compare.
移除它，再比较。

478
00:41:19,131 --> 00:41:25,308
First state its intended role: perhaps the pear's silhouette, perhaps the composition's sense of scale.
先说明它的预期作用：可能是梨的轮廓，也可能是构图的尺度感。

479
00:41:26,112 --> 00:41:29,688
Then look for that quality in outputs with and without the reference.
然后，在有无这张参考图的输出中寻找这种品质。

480
00:41:30,200 --> 00:41:33,047
Also inspect what came along uninvited.
也要检查哪些信息不请自来。

481
00:41:33,047 --> 00:41:37,734
A useful subject reference can carry a background or lighting scheme we did not want.
有用的主体参考，也可能携带我们不想要的背景或光照。

482
00:41:38,406 --> 00:41:43,575
This is an ablation: change one component to understand its contribution.
这就是消融实验：改变一个组件，以理解它的贡献。

483
00:41:44,655 --> 00:41:50,803
Random generation makes one lucky comparison weak evidence, so several samples can help.
随机生成会让一次幸运的对比缺乏说服力，因此多个样本会有帮助。

484
00:41:51,475 --> 00:41:59,388
If we cannot explain what a reference contributes, we may be making the request more complicated without making the direction clearer.
如果说不清一张参考图的贡献，我们可能只是把要求变复杂，却没有让方向更清楚。

485
00:41:59,989 --> 00:42:02,763
Listen to how this prompt distributes responsibility.
听听这个提示词如何分配职责。

486
00:42:03,274 --> 00:42:06,355
The pear reference supplies silhouette and the blue stem.
梨的参考图提供轮廓和蓝色柄。

487
00:42:06,939 --> 00:42:09,932
Placement low on the right organizes the composition.
将主体放在右下方，用来组织构图。

488
00:42:10,517 --> 00:42:13,524
A pale upper-left area reserves space for the title.
左上方的浅色区域，为标题留出空间。

489
00:42:14,036 --> 00:42:17,233
Soft daylight and restrained reflections support the mood.
柔和日光与克制的倒影，支撑整体情绪。

490
00:42:17,905 --> 00:42:21,512
Every phrase points toward something we could inspect in an output.
每个短语，都指向输出中可以检查的东西。

491
00:42:22,183 --> 00:42:26,447
Compare that with asking for a stunning masterpiece with beautiful lighting.
与“用漂亮光照创作一幅惊艳杰作”这样的要求比较一下。

492
00:42:26,447 --> 00:42:30,886
This prompt is still a proposal, not a binding contract enforced by the model.
这个提示词仍然只是提议，不是模型强制执行的约束协议。

493
00:42:31,471 --> 00:42:43,327
If an output fails, we can now name the failed relationship: the title area is crowded, the light is too assertive, or the reference's background has leaked into the scene.
输出失败时，我们现在可以指出哪种关系失败了：标题区太拥挤、光线太强势，或参考图背景渗入了场景。

494
00:42:43,927 --> 00:42:46,337
A single poster can look coherent by itself.
单张海报本身可能看起来很统一。

495
00:42:47,009 --> 00:42:52,703
A series exposes a harder problem: does the same material decision survive across viewpoints?
系列作品暴露出更难的问题：同样的材质决定，能否在不同视角下保持？

496
00:42:53,375 --> 00:42:58,091
Our Group Editing collaboration studies related images that should be edited consistently.
我们合作开展的 Group Editing 研究，关注应当保持一致编辑的关联图像。

497
00:42:58,763 --> 00:43:03,362
Treating each image independently can produce slightly different costumes or materials.
独立处理每张图，可能生成略有不同的服装或材质。

498
00:43:03,946 --> 00:43:09,875
The method arranges related images as pseudo-video frames to use a video model's consistency prior.
该方法把相关图像排列成伪视频帧，利用视频模型的一致性先验。

499
00:43:10,679 --> 00:43:14,533
VGGT provides geometric correspondences.
VGGT 提供几何对应关系。

500
00:43:14,533 --> 00:43:23,513
Geometry-enhanced rotary positional embeddings connect geometry features with image latents, while Identity-RoPE supports identity preservation.
几何增强的旋转位置编码，将几何特征与图像潜变量联系起来；Identity-RoPE 则支持身份保持。

501
00:43:24,184 --> 00:43:28,171
Follow the penguin examples through the figure before reading every label.
阅读每个标签之前，先沿着图中的企鹅示例看。

502
00:43:28,843 --> 00:43:39,590
Reliable correspondence is part of the technical problem: when views cannot be matched well, repeating the same instruction alone does not establish a coherent series.
可靠对应关系是技术问题的一部分：视角无法准确匹配时，仅重复相同指令，并不能保证系列统一。

503
00:43:40,190 --> 00:43:44,102
Iteration becomes informative when we know what changed between attempts.
当我们知道两次尝试之间改了什么，迭代才有信息量。

504
00:43:44,687 --> 00:43:47,213
Suppose we compare two lighting treatments.
假设我们比较两种灯光处理。

505
00:43:47,725 --> 00:43:53,755
Keep the subject and composition instructions stable, and record the references and model version.
保持主体和构图指令不变，并记录参考图和模型版本。

506
00:43:54,427 --> 00:43:58,792
If a seed control exists, holding it fixed can help the first comparison.
如果可以控制随机种子，固定种子有助于初步比较。

507
00:43:59,464 --> 00:44:04,779
A shared seed still does not guarantee identical composition after a prompt change.
但即使种子相同，改变提示词后，也不能保证构图完全一致。

508
00:44:04,779 --> 00:44:08,707
If results vary widely, examine several outputs for each condition.
如果结果差异很大，就检查每种条件下的多个输出。

509
00:44:09,379 --> 00:44:12,635
Otherwise, one lucky sample may decide the whole direction.
否则，一个幸运样本就可能决定整个方向。

510
00:44:13,220 --> 00:44:16,490
The table is a proposed experiment, not measured results.
这张表是拟议实验，不是实测结果。

511
00:44:17,075 --> 00:44:21,308
Its purpose is to make each generation answer a question about the artwork.
它的目的，是让每次生成都回答一个关于作品的问题。

512
00:44:21,909 --> 00:44:26,932
Imagine you compare two prompts once, and the second gives a wonderful poster.
想象你只比较了两条提示词各一次，第二条生成了精彩海报。

513
00:44:28,187 --> 00:44:33,020
Did its wording cause the improvement, or did it receive a favorable random sample?
改进是措辞造成的，还是它碰巧抽到了有利的随机样本？

514
00:44:33,692 --> 00:44:36,263
From one output, it can be difficult to tell.
只看一个输出，往往很难判断。

515
00:44:36,934 --> 00:44:41,854
Several outputs per condition reveal whether the direction is stable or merely fortunate.
每种条件生成多个结果，才能看出这个方向是稳定有效，还是仅仅幸运。

516
00:44:42,526 --> 00:44:45,462
For art direction, we can ask two different questions.
在艺术指导中，我们可以问两个不同的问题。

517
00:44:45,740 --> 00:44:48,134
Would I exhibit this particular image?
我愿意展出这张具体图像吗？

518
00:44:48,134 --> 00:44:51,375
Could I reliably develop a series using this direction?
我能用这个方向，可靠地发展出一个系列吗？

519
00:44:52,046 --> 00:44:56,486
One exceptional output may answer the first while leaving the second unresolved.
一张出色的输出可能回答了第一个问题，却没有解决第二个。

520
00:44:57,288 --> 00:45:07,935
Repeating generations is helpful when it reveals a pattern relevant to the brief; it becomes a distraction when we keep browsing alternatives to avoid deciding what we value.
重复生成若能揭示与创作要求相关的规律，就有帮助；若只是为了逃避价值判断而不断浏览方案，就成了干扰。

521
00:45:08,534 --> 00:45:10,462
Now the image becomes a poster.
现在，图像变成了海报。

522
00:45:11,045 --> 00:45:19,149
The title is an editable typographic layer, which lets us choose its wording, line breaks, and placement exactly.
标题是可编辑的文字图层，让我们精确选择措辞、换行和位置。

523
00:45:20,069 --> 00:45:27,328
That matters when a work has to carry an exhibition name rather than merely resemble a poster in a generated preview.
作品需要准确呈现展览名称时，这很重要，而不是只在生成预览中“看起来像海报”。

524
00:45:28,130 --> 00:45:31,824
Look at how the title occupies the area we deliberately left open.
看看标题如何占据我们有意留出的区域。

525
00:45:32,628 --> 00:45:35,460
The image and type were designed to cooperate.
图像与文字是协同设计的。

526
00:45:35,460 --> 00:45:40,614
We can adjust hierarchy without asking the model to regenerate the sculpture and risk changing it.
我们可以调整层次，而不必让模型重新生成雕塑、冒着改变它的风险。

527
00:45:41,285 --> 00:45:49,098
This is a useful division of labor: generation develops the visual material, and direct layout controls exact communication.
这是一种有用的分工：生成负责发展视觉素材，直接排版负责精确传达信息。

528
00:45:49,770 --> 00:45:51,902
The audience sees one composition.
观众看到的是一个整体构图。

529
00:45:52,414 --> 00:45:55,728
It does not need to know which parts came from which tool.
他们不需要知道哪个部分来自哪种工具。

530
00:45:56,327 --> 00:45:59,190
Typography establishes a sequence of attention.
文字排版建立了注意力的顺序。

531
00:45:59,700 --> 00:46:04,505
Ask what the viewer should read first, what comes next, and where their eye returns to the artwork.
要问：观众应先读什么、接着读什么，以及视线在哪里回到作品？

532
00:46:05,307 --> 00:46:10,286
In our poster, reserved space allows the title to be clear without covering the sculpture.
在我们的海报中，预留空间让标题清楚可读，又不会遮住雕塑。

533
00:46:10,870 --> 00:46:12,871
This is also a production decision.
这也是一个制作决定。

534
00:46:13,382 --> 00:46:18,756
Exact text can remain editable, while the image carries the material and atmosphere.
精确文字可以保持可编辑，而图像承载材质和氛围。

535
00:46:18,756 --> 00:46:25,078
If the title feels too dominant, change size, placement, or contrast and inspect the whole composition again.
如果标题太抢眼，就调整大小、位置或对比度，再检查整个构图。

536
00:46:25,663 --> 00:46:28,612
Do not judge the text in isolation from the image.
不要脱离图像，单独评价文字。

537
00:46:29,123 --> 00:46:35,284
The poster is the relationship between them, including the empty space that lets each element do its job.
海报是两者之间的关系，也包括让各元素发挥作用的空白。

538
00:46:35,885 --> 00:46:39,097
Take a few seconds with both directions before I describe them.
在我描述之前，先花几秒看看两个方向。

539
00:46:40,018 --> 00:46:42,207
Which would you put outside AFTER RAIN?
你会把哪一个放在《雨后》展厅外？

540
00:46:42,631 --> 00:46:46,895
Choose privately first, so your answer is not just a response to mine.
先自己选择，避免答案只是对我的回应。

541
00:46:47,479 --> 00:46:50,486
Then identify the visible feature that made you choose.
然后指出促使你选择的可见特征。

542
00:46:51,289 --> 00:46:55,480
The photographic direction can invite attention to material and atmosphere.
摄影方向可以引导观众关注材质与氛围。

543
00:46:56,152 --> 00:47:01,584
The collage direction can make fragility feel constructed through paper edges and layers.
拼贴方向可以通过纸边和层次，让脆弱感显得被建构出来。

544
00:47:01,584 --> 00:47:05,454
Both can serve the brief, but they make different promises to a visitor.
两者都可能符合要求，但向观众许下的承诺不同。

545
00:47:06,038 --> 00:47:09,747
A useful defense goes beyond realism versus abstraction.
有用的辩护，不应止于写实与抽象的区别。

546
00:47:10,330 --> 00:47:16,230
Tell us what the audience is likely to notice or feel, and how the composition produces that reading.
请告诉我们，观众可能注意或感受到什么，以及构图如何产生这种解读。

547
00:47:17,033 --> 00:47:23,604
We will use the disagreement to decide what to revise, rather than vote for a universally better image.
我们将用分歧来决定修改方向，而不是投票选出普遍更好的图像。

548
00:47:24,204 --> 00:47:26,526
Selection is an artistic decision.
选择本身就是艺术决定。

549
00:47:27,109 --> 00:47:33,549
Once generation gives us many plausible alternatives, choosing one determines what the work becomes.
当生成提供许多合理方案时，选中哪一个，就决定作品最终成为什么。

550
00:47:34,134 --> 00:47:43,100
The reason should connect to the brief: perhaps the small object feels more exposed, or the torn edge makes fragility physically legible.
理由应与创作要求相连：可能是小物体更显得无所庇护，或撕裂边缘让脆弱感变得可触可见。

551
00:47:43,771 --> 00:47:46,736
The rejected direction is useful evidence too.
被放弃的方向也是有用证据。

552
00:47:47,538 --> 00:47:52,181
Explain what it does well and why you are choosing something else for this exhibition.
请解释它哪里做得好，以及为什么为这个展览选择了别的方案。

553
00:47:52,181 --> 00:47:57,744
An unexpected model output can also change the direction, if you choose to develop it deliberately.
如果你决定有意识地发展一个意外输出，它也可以改变创作方向。

554
00:47:58,329 --> 00:48:04,695
The key question is whether you can now articulate the intention and carry it through subsequent decisions.
关键是：你现在能否说清意图，并在后续决定中贯彻它？

555
00:48:05,498 --> 00:48:10,404
Surprise can begin a work; it does not finish the artist's judgment.
惊喜可以开启作品，却不能替艺术家完成判断。

556
00:48:11,004 --> 00:48:15,750
A useful critique names a visible feature, explains its effect, and proposes a revision.
有用的评论会指出可见特征、解释效果，并提出修改建议。

557
00:48:16,335 --> 00:48:23,927
For example: the crisp reflection makes the collage feel photographic, so I would simplify its shape to match the flatter layers.
例如：清晰倒影让拼贴显得像摄影，所以我会简化它的形状，以配合更平面的层次。

558
00:48:24,511 --> 00:48:33,549
Compare that with 'I do not like the reflection.' The first comment gives the artist a relationship to inspect and a possible next move.
与“我不喜欢倒影”相比，前一种评论给艺术家提供了可检查的关系，以及可能的下一步。

559
00:48:33,550 --> 00:48:38,602
You can disagree about the intended effect, but make the disagreement specific.
你们可以对预期效果有不同看法，但要把分歧说具体。

560
00:48:39,274 --> 00:48:43,683
This is a classroom critique framework rather than an official grading rubric.
这是课堂评论框架，不是正式评分标准。

561
00:48:44,487 --> 00:48:49,378
Use it to connect evidence in the image to the idea the work is trying to communicate.
用它把图像中的证据，与作品试图传达的观念联系起来。

562
00:48:49,978 --> 00:48:54,856
A curator asks, can the same sculpture feel more vulnerable without changing its shape?
策展人问：能否不改变形状，却让同一件雕塑显得更脆弱？

563
00:48:55,658 --> 00:48:57,498
Lighting is one way to answer.
灯光是一种回答方式。

564
00:48:58,170 --> 00:49:07,107
Our LightCtrl collaboration studies controllable relighting from a single image, including light direction, intensity, and color temperature.
我们合作开展的 LightCtrl 研究，探索从单张图像进行可控重光照，包括光线方向、强度和色温。

565
00:49:07,690 --> 00:49:09,502
Follow the chair through the figure.
请沿着图中的椅子示例看。

566
00:49:09,778 --> 00:49:13,692
The latent proxy encoder extracts compact physical cues.
潜在代理编码器提取紧凑的物理线索。

567
00:49:13,692 --> 00:49:23,825
A lighting-aware mask guides the denoiser toward regions affected by the change, and preference optimization in the proxy branch supports physical consistency.
光照感知蒙版引导去噪器关注受变化影响的区域，代理分支中的偏好优化则支持物理一致性。

568
00:49:24,746 --> 00:49:29,126
The method connects an interpretable lighting request with image generation.
这种方法把可解释的光照要求与图像生成连接起来。

569
00:49:29,798 --> 00:49:40,384
For our pear, look for consequences rather than the word dramatic: where do highlights move, what happens to shadow, and how does the material read?
对于梨，不要只看“戏剧性”这个词，而要看后果：高光移向哪里、阴影如何变化、材质呈现如何？

570
00:49:40,384 --> 00:49:46,605
Inferring geometry and material from one photograph is ambiguous, so the result still needs inspection.
从单张照片推断几何和材质存在歧义，因此仍需检查结果。

571
00:49:47,188 --> 00:49:53,760
The exhibition example is our application of the idea, not an additional result reported by the paper.
展览示例是我们对这一思路的应用，不是论文另外报告的实验结果。

572
00:49:54,359 --> 00:49:57,587
Choose between the two prepared directions for AFTER RAIN.
请在为《雨后》准备的两个方向中做出选择。

573
00:49:58,097 --> 00:50:01,806
Which poster communicates fragility more clearly, and why?
哪张海报更清楚地传达脆弱感？为什么？

574
00:50:02,317 --> 00:50:05,617
What does the rejected direction reveal about your choice?
被放弃的方向，揭示了你选择中的什么？

575
00:50:06,421 --> 00:50:09,896
Which single revision would most change the audience’s reading?
哪一个单独的修改，会最大程度改变观众的解读？

576
00:50:10,567 --> 00:50:14,495
Discuss these three questions with a specific visual feature in view.
请看着一个具体视觉特征，讨论这三个问题。

577
00:50:15,006 --> 00:50:22,219
You can prefer quiet space or the instability of torn paper, but explain how that choice serves the exhibition.
你可以偏爱安静空间，也可以偏爱撕纸的不稳定感，但要解释这种选择如何服务于展览。

578
00:50:22,220 --> 00:50:23,549
Pause the video here.
请在这里暂停视频。

579
00:50:23,972 --> 00:50:28,352
Listen for a persuasive reason to choose the direction you initially rejected.
试着听取一个有说服力的理由，支持你最初拒绝的方向。

580
00:50:32,353 --> 00:50:35,419
Consider what a process record lets us discuss.
想一想，过程记录能让我们讨论什么。

581
00:50:35,843 --> 00:50:44,063
An artist can record the intention and invariants, the prompt and reference roles, and one observation about a rejected result.
艺术家可以记录意图与不变量、提示词与参考图角色，以及对一个被拒绝结果的观察。

582
00:50:44,648 --> 00:50:47,509
That record explains what an attempt was testing.
这些记录说明了每次尝试在检验什么。

583
00:50:48,020 --> 00:50:52,678
Without that context, a folder of attractive images can be difficult to interpret.
缺少背景信息，一文件夹好看的图像也可能很难解读。

584
00:50:53,262 --> 00:50:57,409
We may not know which input changed or why a version was rejected.
我们可能不知道改了哪个输入，或为什么拒绝某个版本。

585
00:50:57,409 --> 00:51:00,798
A few precise notes make a comparison more informative.
几条精确笔记，就能让比较更有信息量。

586
00:51:01,381 --> 00:51:08,930
In a critique, this lets us ask about the artist’s decisions and the evidence behind them rather than guessing only from the final image.
在评论中，这让我们能够追问艺术家的决定及其证据，而不只是从最终图像猜测。

587
00:51:09,602 --> 00:51:14,289
It also helps distinguish an intentional departure from an accidental change.
它也有助于区分有意偏离与意外变化。

588
00:51:14,889 --> 00:51:18,919
Imagine we are standing between the two finished posters in a gallery review.
想象我们站在画廊评审现场，面前是两张完成的海报。

589
00:51:19,591 --> 00:51:28,425
I would begin with a visible decision that carries the idea, then locate an unintended change, then propose one revision worth discussing.
我会先指出一个承载观念的可见决定，再找出一个意外变化，最后提出一个值得讨论的修改。

590
00:51:29,096 --> 00:51:34,689
The order matters: the critique begins by understanding the work before prescribing a repair.
顺序很重要：评论应先理解作品，再给出修正方案。

591
00:51:35,273 --> 00:51:38,719
We can hear two contrasting readings of the same poster.
对同一张海报，我们可能听到两种相反的解读。

592
00:51:38,720 --> 00:51:43,641
One person may find the empty space quiet; another may find it emotionally distant.
一个人觉得留白安静，另一个人却觉得它情感疏离。

593
00:51:44,152 --> 00:51:46,546
Ask which details support each reading.
请问，哪些细节支持各自的解读？

594
00:51:46,970 --> 00:51:49,306
There is no image-making task here.
这里没有图像制作任务。

595
00:51:49,729 --> 00:51:53,467
We are practicing how to make feedback useful to the next decision.
我们是在练习如何让反馈有助于下一次决定。

596
00:51:54,052 --> 00:52:02,024
A good proposal names what would change and what effect we expect, so that a later version could confirm or challenge the reasoning.
好建议会说明改什么、预期产生什么效果，让后续版本能够证实或挑战这套推理。

597
00:52:02,024 --> 00:52:07,573
Discuss these three questions, then pause the video for the gallery conversation.
请讨论这三个问题，然后暂停视频，进行画廊对话。

598
00:52:11,573 --> 00:52:16,421
Before the artwork begins to move, defend a decision about the images.
在作品开始运动之前，请为一个图像决定提供理由。

599
00:52:17,005 --> 00:52:20,144
How can a technically correct image fail as an artwork?
技术上正确的图像，为什么可能在艺术上失败？

600
00:52:20,816 --> 00:52:23,780
Which visual rule should remain across a series?
一个系列中，哪条视觉规则应该保持？

601
00:52:24,364 --> 00:52:27,737
What evidence makes a critique useful for the next revision?
什么证据能让评论对下一轮修改有用？

602
00:52:28,409 --> 00:52:31,956
Discuss these three questions through one of the prepared posters.
请结合其中一张准备好的海报，讨论这三个问题。

603
00:52:31,957 --> 00:52:40,689
Make words such as coherent or expressive concrete: locate the feature, describe its effect, and explain what changing it would do.
把“统一”“有表现力”等词说具体：指出特征、描述效果，并解释改变它会怎样。

604
00:52:41,361 --> 00:52:42,675
Pause the video here.
请在这里暂停视频。

605
00:52:43,259 --> 00:52:46,793
We will carry those artistic rules into the video section.
我们会把这些艺术规则带入视频部分。

606
00:52:50,793 --> 00:52:52,940
Our image now has to survive time.
现在，图像必须经受时间的考验。

607
00:52:53,523 --> 00:52:56,735
The subject can move, disappear, and return.
主体会运动、消失，然后返回。

608
00:52:57,992 --> 00:53:08,972
We will examine selected frames from published research, follow the technical representations, and use a prepared AFTER RAIN scenario to decide what a convincing edit requires.
我们将查看已发表研究的精选帧、理解技术表示，并用准备好的《雨后》情境判断可信编辑需要什么。

609
00:53:09,775 --> 00:53:13,747
Keep the preservation contract, but add an event that could break it.
保留编辑约定，再加入一个可能让它失效的事件。

610
00:53:14,347 --> 00:53:17,311
A moving image creates new ways to break our contract.
动态图像会以新的方式破坏约定。

611
00:53:17,895 --> 00:53:21,049
What new failures become possible when an image moves?
图像运动起来后，会出现哪些新的失败？

612
00:53:21,560 --> 00:53:24,583
What should happen when the pear disappears behind a column?
梨消失在柱子后面时，应该发生什么？

613
00:53:25,167 --> 00:53:28,306
Can four convincing frames prove that a video works?
四张可信的帧，能证明整段视频成立吗？

614
00:53:28,890 --> 00:53:30,803
Discuss these three questions.
请讨论这三个问题。

615
00:53:31,606 --> 00:53:36,425
Separate a correct change in visibility from an unwanted change in identity.
区分正确的可见性变化，与不希望发生的身份变化。

616
00:53:36,425 --> 00:53:40,894
Name a moment you would need to see between the selected frames before accepting the shot.
接受这个镜头之前，请指出你需要在这些精选帧之间看到的某个时刻。

617
00:53:41,477 --> 00:53:42,864
Pause the video here.
请在这里暂停视频。

618
00:53:43,375 --> 00:53:47,230
Your prediction will give us a concrete test for the methods that follow.
你的预测会为后续方法提供一个具体测试。

619
00:53:51,230 --> 00:53:53,961
A still image lets us choose a flattering instant.
静态图像让我们挑选一个最好看的瞬间。

620
00:53:54,544 --> 00:53:57,904
A video makes the object keep its promises in the next frame.
视频则要求物体在下一帧继续信守承诺。

621
00:53:58,575 --> 00:54:00,239
Look at the source sequence.
看原始序列。

622
00:54:00,824 --> 00:54:05,832
As viewpoint and visibility change, we continue to recognize the subject.
随着视角和可见性变化，我们仍能认出主体。

623
00:54:06,504 --> 00:54:11,366
An edit has to preserve that relationship while introducing its requested transformation.
编辑必须在引入所要求变换的同时，保留这种关系。

624
00:54:12,169 --> 00:54:13,819
Remember our two clocks.
记住两种时钟。

625
00:54:14,331 --> 00:54:17,980
The sampling steps describe how a model generates a result.
采样步描述模型如何生成结果。

626
00:54:17,981 --> 00:54:21,471
The frames here describe time passing in the depicted scene.
这里的帧描述被描绘场景中时间的流逝。

627
00:54:22,055 --> 00:54:28,728
A system may use many sampling steps to produce a short clip, but those steps are not extra seconds of action.
系统可能用很多采样步生成短片，但这些步数不是额外的剧情秒数。

628
00:54:29,400 --> 00:54:35,035
Now the preservation contract must hold across an event, not just inside a frame.
现在，保留约定必须跨越一个事件成立，而不只是局限于单帧内部。

629
00:54:35,636 --> 00:54:39,316
Would a video be perfectly consistent if every frame were identical?
如果每帧完全相同，视频就算完美一致吗？

630
00:54:39,988 --> 00:54:42,732
Only if the intended scene were perfectly still.
只有目标场景本来就完全静止时才是。

631
00:54:43,653 --> 00:54:51,289
In our moving shot, pose, viewpoint, illumination, and visibility should change.
在运动镜头中，姿态、视角、光照和可见性都应变化。

632
00:54:52,209 --> 00:54:56,384
Consistency means that those changes remain coherent with the scene.
一致性意味着这些变化与场景保持协调。

633
00:54:56,969 --> 00:54:59,830
The blue stem may become hidden during a turn.
蓝色柄可能在转动时被遮住。

634
00:55:00,342 --> 00:55:04,664
That is different from its color changing without a lighting explanation.
这与没有光照原因却突然变色，是不同的。

635
00:55:04,665 --> 00:55:09,264
A texture should move with its surface, rather than crawl independently across it.
纹理应该随表面一起移动，而不是独自在表面爬动。

636
00:55:09,936 --> 00:55:14,040
If there is no convincing explanation, you may have found instability.
如果找不到可信解释，你可能发现了不稳定性。

637
00:55:14,842 --> 00:55:20,026
This is why the goal cannot simply be to minimize change from one frame to the next.
因此，目标不能只是尽量减少相邻帧之间的变化。

638
00:55:20,626 --> 00:55:24,043
Here is a prepared storyboard, not a generated video result.
这是准备好的分镜，不是生成视频的结果。

639
00:55:24,554 --> 00:55:28,671
In the first view, the reflection suggests an object we have not fully seen.
第一个视角中，倒影暗示了一个我们尚未看全的物体。

640
00:55:29,182 --> 00:55:32,541
A partial view then gives us enough evidence to make a guess.
接着，局部视角提供足够线索，让我们猜测。

641
00:55:33,052 --> 00:55:36,439
The wide view finally changes our understanding of the setting.
最后的全景改变了我们对环境的理解。

642
00:55:36,951 --> 00:55:44,105
The audience is doing something during those five seconds: forming an expectation, testing it, and revising it.
在这五秒中，观众一直在行动：建立预期、检验预期，然后修正它。

643
00:55:44,106 --> 00:55:47,567
Compare this sequence with revealing everything in the first frame.
把这个序列与第一帧就揭示一切的版本比较一下。

644
00:55:48,150 --> 00:55:51,612
The same sculpture could be present, but the experience would differ.
同一件雕塑可以都在场，但体验会不同。

645
00:55:52,531 --> 00:55:58,123
For AFTER RAIN, the timing of information can carry fragility as strongly as material or lighting.
对于《雨后》，信息揭示的时机，可以像材质或灯光一样有力地承载脆弱感。

646
00:55:58,708 --> 00:56:02,679
We should decide that structure before asking a tool to fill in motion.
要求工具补充运动之前，我们应该先决定这种结构。

647
00:56:03,279 --> 00:56:06,171
Five seconds is short, but it can still have a structure.
五秒虽然短，却仍然可以有结构。

648
00:56:06,842 --> 00:56:10,901
Our proposed first interval shows a reflection, giving the viewer clues.
我们设想的第一段展示倒影，给观众线索。

649
00:56:11,325 --> 00:56:13,296
The second offers a partial view.
第二段提供局部视角。

650
00:56:13,662 --> 00:56:18,173
The final interval reveals the wider setting and changes how the object is understood.
最后一段揭示更广阔的环境，改变人们对物体的理解。

651
00:56:18,685 --> 00:56:21,794
Try another allocation and predict the effect.
试着重新分配时间，并预测效果。

652
00:56:21,795 --> 00:56:28,278
A longer reflection might create uncertainty; a quick reveal might make the piece feel like a product shot.
更长的倒影段落可能制造不确定感；快速揭示则可能让它像产品广告镜头。

653
00:56:28,950 --> 00:56:33,593
These timings are artistic proposals, not measured outputs from a video model.
这些时间安排是艺术提案，不是视频模型的实测输出。

654
00:56:34,265 --> 00:56:39,726
By deciding the experience first, we can evaluate whether the generated motion and cuts support it.
先决定体验，才能评价生成的运动和剪辑是否支持它。

655
00:56:40,309 --> 00:56:44,048
A technically smooth sequence can still have the wrong rhythm.
技术上流畅的序列，节奏仍可能不对。

656
00:56:44,648 --> 00:56:53,102
The research begins with a surprisingly simple experiment: arrange video frames into a contact sheet and give that image to an image editor.
这项研究始于一个出人意料的简单实验：把视频帧排成接触印样，再交给图像编辑器。

657
00:56:53,906 --> 00:57:00,257
This asks whether an existing image-editing capability can transfer across multiple views presented together.
它想检验：现有图像编辑能力，能否迁移到一起呈现的多个视角上？

658
00:57:01,061 --> 00:57:05,032
Treat the result as evidence that motivates a research direction.
请把结果当作启发研究方向的证据。

659
00:57:05,032 --> 00:57:12,216
A grid can make frames available in a shared context, but a successful-looking selection does not prove reliable video editing.
网格让多帧共享上下文，但一组看似成功的精选帧，并不能证明视频编辑可靠。

660
00:57:12,888 --> 00:57:15,589
There may be flicker between the sampled frames.
采样帧之间可能仍有闪烁。

661
00:57:16,013 --> 00:57:24,189
The next steps investigate how to represent video more effectively and adapt the model, rather than assuming a contact sheet alone solves time.
后续研究探索更有效的视频表示和模型适配，而不是假设接触印样本身就解决了时间问题。

662
00:57:24,790 --> 00:57:28,455
But the missing intervals are where some of the most revealing failures live.
恰恰是在缺失的间隔中，可能藏着最有揭示性的失败。

663
00:57:29,040 --> 00:57:33,317
A feature can jump, disappear, and return between the frames on this page.
某个特征可能在这一页的两帧之间跳动、消失，再返回。

664
00:57:33,989 --> 00:57:40,896
Use the contact sheet for what it does well: compare appearance across selected moments and locate regions to inspect.
发挥接触印样的长处：比较选定时刻的外观，并定位需要检查的区域。

665
00:57:41,407 --> 00:57:45,875
Then review full playback, with closer attention around turns and occlusion.
然后完整播放，特别关注转向和遮挡附近。

666
00:57:45,875 --> 00:57:49,116
Think of a film review based only on publicity stills.
想象只凭宣传剧照来评论电影。

667
00:57:49,540 --> 00:57:54,344
You might judge the costume and lighting, but you have not yet seen the performance.
你或许能评价服装和灯光，却还没有看过表演。

668
00:57:55,264 --> 00:58:03,572
In generative video, that missing performance includes whether the same object continues to exist convincingly from moment to moment.
对生成视频而言，缺失的“表演”包括：同一个物体是否在每个时刻都可信地持续存在。

669
00:58:04,172 --> 00:58:07,093
The grid can also be constructed in latent space.
网格也可以在潜在空间中构造。

670
00:58:07,676 --> 00:58:15,372
Instead of only combining visible pixel images, the method encodes frames and arranges their tokens into a virtual image grid.
这种方法不只是组合可见像素图像，而是先编码各帧，再把 token 排列成虚拟图像网格。

671
00:58:16,044 --> 00:58:19,504
The published example changes a bus into a graphics card.
论文中的示例把一辆公交车变成显卡。

672
00:58:20,308 --> 00:58:25,141
The key idea is shared treatment of positions within that constructed representation.
核心思路是，在这种构造的表示中统一处理各位置。

673
00:58:25,141 --> 00:58:30,602
Do not confuse this step with using a video VAE; they are different design decisions.
不要把这一步与使用视频 VAE 混为一谈；它们是不同的设计决定。

674
00:58:31,112 --> 00:58:39,933
Putting frames near one another in a representation can help the model exchange information, but it does not impose a guarantee of physical continuity.
在表示中把帧放在一起，有助于模型交换信息，但不会强制保证物理连续性。

675
00:58:40,604 --> 00:58:44,473
We still need to inspect what survives across changing views.
我们仍需检查视角变化时，哪些信息得以保留。

676
00:58:45,074 --> 00:58:50,462
A virtual grid provides a way to arrange information for a model trained around image-like structure.
虚拟网格为围绕图像结构训练的模型，提供了一种信息排列方式。

677
00:58:51,045 --> 00:58:55,996
It creates shared context and positional relationships among frame representations.
它在各帧表示之间建立共享上下文和位置关系。

678
00:58:56,580 --> 00:58:59,661
That can be a useful bridge when reusing an image editor.
复用图像编辑器时，这可以成为有用的桥梁。

679
00:59:00,172 --> 00:59:05,808
But arrangement alone does not enforce the laws of motion or the persistence of a hidden object.
但排列方式本身，不能强制执行运动规律或被遮挡物体的持续存在。

680
00:59:05,808 --> 00:59:12,043
Those abilities depend on the model, adaptation, data, and other parts of the workflow.
这些能力取决于模型、适配、数据，以及工作流程中的其他部分。

681
00:59:12,628 --> 00:59:16,147
Separate the representation choice from the capability claim.
请区分表示选择与能力主张。

682
00:59:16,818 --> 00:59:23,112
A diagram can explain how frames enter a system without proving that the system handles every difficult event.
图表可以解释帧如何进入系统，却不能证明系统处理得了所有困难事件。

683
00:59:23,695 --> 00:59:27,477
That proof would require appropriate outputs and evaluation.
证明这一点，需要适当的输出和评价。

684
00:59:28,078 --> 00:59:32,867
This architecture asks whether an image editor's abilities can be reused for video.
这个架构检验图像编辑器的能力能否复用于视频。

685
00:59:33,539 --> 00:59:38,430
The video VAE encodes a clip into a compressed latent representation.
视频 VAE 将片段编码成压缩的潜在表示。

686
00:59:39,015 --> 00:59:47,528
Learned projections connect those video latents with the image-editing transformer, so the model can operate on a compatible arrangement of information.
学习得到的投影把视频潜变量与图像编辑 Transformer 连接起来，让模型处理兼容的信息排列。

687
00:59:48,200 --> 00:59:51,616
Trace that main route before reading the smaller branches.
阅读较小分支之前，先沿着这条主路径看。

688
00:59:51,616 --> 00:59:58,990
Adaptation connects them; the paper explores LoRA and full training choices, with an optional enhancement stage.
适配把它们连接起来；论文探索了 LoRA 和全量训练，并包含可选的增强阶段。

689
00:59:59,573 --> 01:00:02,611
The attractive idea is reuse of editing knowledge.
吸引人的想法是复用编辑知识。

690
01:00:03,196 --> 01:00:07,868
The test is whether that reuse preserves the temporal relationships our shot needs.
真正的检验是：这种复用是否保留了镜头需要的时间关系？

691
01:00:08,380 --> 01:00:15,388
Architecture explains where information flows; the edited clip shows whether the intended event survives.
架构解释信息流向哪里；编辑后的片段展示预期事件是否得以保留。

692
01:00:15,988 --> 01:00:21,581
The image model and video VAE do not necessarily speak the same representational language.
图像模型与视频 VAE 未必使用相同的表示语言。

693
01:00:22,252 --> 01:00:24,340
Learned projections help connect them.
学习得到的投影帮助连接二者。

694
01:00:25,260 --> 01:00:30,809
Adaptation then allows the editing transformer to operate usefully with the video representation.
之后，适配让编辑 Transformer 能有效使用视频表示。

695
01:00:31,480 --> 01:00:39,073
This is a common engineering pattern: reuse a capable component while learning the interface and behavior needed for another task.
这是一种常见工程模式：复用有能力的组件，同时学习另一项任务所需的接口和行为。

696
01:00:39,584 --> 01:00:41,920
The benefit is not automatic.
收益并不会自动出现。

697
01:00:41,921 --> 01:00:49,061
A projection must preserve useful information, and the adapted model must learn how the new structure relates to edits.
投影必须保留有用信息，适配后的模型也必须学会新结构如何与编辑相关联。

698
01:00:49,732 --> 01:00:55,369
In the published pipeline, decoding returns the edited representation to video.
在已发表的流程中，解码将编辑后的表示还原为视频。

699
01:00:56,172 --> 01:01:03,473
Follow the information through every stage rather than treating the name of the reused model as an explanation by itself.
请追踪信息经过每个阶段，而不要把被复用模型的名字本身当作解释。

700
01:01:04,073 --> 01:01:06,512
Here is a counting puzzle with a small trap.
这里有一个带小陷阱的计数题。

701
01:01:06,936 --> 01:01:15,098
In this example, forty-five pixel frames are compressed with a special first frame and a factor of four for the remaining temporal groups.
在这个例子中，四十五个像素帧采用特殊首帧处理，其余时间组按四倍压缩。

702
01:01:15,770 --> 01:01:18,733
Forty-five is four times eleven, plus one.
四十五等于四乘十一，再加一。

703
01:01:19,537 --> 01:01:23,800
The latent sequence therefore has eleven plus one, or twelve frames.
因此，潜在序列有十一加一，即十二帧。

704
01:01:24,311 --> 01:01:29,305
Those twelve latent frames can be arranged in the illustrated three-by-four virtual grid.
这十二个潜在帧可以排成图中的三乘四虚拟网格。

705
01:01:29,305 --> 01:01:33,496
Simply dividing forty-five by four would miss the first-frame convention.
简单地用四十五除以四，会忽略首帧约定。

706
01:01:34,168 --> 01:01:40,812
This arithmetic belongs to the representation used in this example; it is not a rule for every video model.
这个算法属于本例采用的表示，并不是所有视频模型的通用规则。

707
01:01:41,484 --> 01:01:46,594
A small detail in temporal compression changes what the editing model actually receives.
时间压缩中的一个小细节，会改变编辑模型实际收到什么。

708
01:01:47,194 --> 01:01:50,143
Solve the frame-count relationship step by step.
一步步解出帧数关系。

709
01:01:50,728 --> 01:01:53,969
Start with four k plus one equals forty-five.
从 4k 加一等于四十五开始。

710
01:01:54,553 --> 01:02:00,409
Subtract one to obtain forty-four, then divide by four to get k equals eleven.
减一得到四十四，再除以四，得到 k 等于十一。

711
01:02:00,992 --> 01:02:06,059
The latent sequence contains k plus one frames, so its length is twelve.
潜在序列包含 k 加一帧，因此长度为十二。

712
01:02:06,644 --> 01:02:11,842
This arithmetic encodes the special handling of the first frame in the stated convention.
这个计算反映了该约定对首帧的特殊处理。

713
01:02:11,842 --> 01:02:16,412
It is different from simply dividing the total number of pixel frames by four.
它不同于直接把像素帧总数除以四。

714
01:02:17,332 --> 01:02:26,386
Understanding them can prevent mistakes when arranging latent frames into a virtual grid or comparing the representation with the original clip.
理解这些关系，可以避免在排列潜在帧网格、或与原始片段比较时出错。

715
01:02:26,985 --> 01:02:32,067
The requested edit turns the scene into a cyberpunk workshop with holographic documents.
编辑要求是把场景变成赛博朋克工作室，并加入全息文档。

716
01:02:32,738 --> 01:02:35,908
Compare the results before focusing on the setting labels.
先比较结果，再关注设置标签。

717
01:02:36,579 --> 01:02:43,281
Which version carries out the requested transformation more completely, and what visible evidence supports your answer?
哪个版本更完整地实现了要求的变换？有什么可见证据？

718
01:02:43,953 --> 01:02:49,808
The paper presents this case as an example where classifier-free guidance improves edit completeness.
论文用这个案例说明，无分类器引导可以改善编辑完整性。

719
01:02:50,393 --> 01:02:55,547
Its baseline retains source latents, and its implementation includes rescaling.
其基线保留了原始潜变量，实现中还包含重缩放。

720
01:02:55,547 --> 01:02:59,154
This does not establish one universally best guidance value.
这并不意味着存在一个普遍最优的引导值。

721
01:02:59,825 --> 01:03:04,337
Guidance also has a computation cost when it requires another model pass.
如果引导需要额外一次模型前向计算，也会带来计算成本。

722
01:03:05,008 --> 01:03:12,952
Relate the example back to our equation: changing a prediction combination affects how strongly the result follows the condition.
联系之前的方程：改变预测组合，会影响结果遵循条件的力度。

723
01:03:13,552 --> 01:03:16,940
An ablation asks what changes when one component is removed.
消融实验追问：移除某个组件后，会发生什么变化？

724
01:03:17,524 --> 01:03:22,240
In the displayed example, guidance makes more of the requested transformation visible.
在展示的例子中，引导让更多要求的变换可见。

725
01:03:22,824 --> 01:03:25,439
That is useful evidence about this comparison.
这是关于这次比较的有用证据。

726
01:03:26,022 --> 01:03:31,410
It is not yet a measurement of reliability across the kinds of shots we might produce for an exhibition.
但它还不是对我们可能用于展览的各类镜头的可靠性测量。

727
01:03:32,082 --> 01:03:35,922
As a reader, separate the observation from the next question.
作为读者，请区分观察结果与下一步问题。

728
01:03:35,922 --> 01:03:38,244
We can observe a more complete edit here.
我们可以观察到，这里的编辑更加完整。

729
01:03:38,755 --> 01:03:46,378
We would still want to know what happens across other clips, how often preservation suffers, and what the computation costs.
但仍想知道：其他片段会怎样、保留效果多常受损，以及计算成本是多少。

730
01:03:46,961 --> 01:03:54,437
You can learn a mechanism from a selected example while designing a stronger test for the decision you actually need to make.
你可以从精选示例中学习机制，同时为真正需要做出的决定设计更强的测试。

731
01:03:55,038 --> 01:04:00,017
This is a local edit: the sheep's face changes and receives a white star-shaped patch.
这是一次局部编辑：绵羊的脸发生变化，增加了一块白色星形斑纹。

732
01:04:00,600 --> 01:04:02,616
Look at the patch across poses.
观察不同姿态下的这块斑纹。

733
01:04:03,127 --> 01:04:09,580
Does it remain attached to the same facial region, with a plausible change in apparent shape as the head turns?
它是否始终附着在同一面部区域，并在头部转动时合理改变表观形状？

734
01:04:10,091 --> 01:04:12,311
The most attractive frame is not enough.
仅看最漂亮的一帧是不够的。

735
01:04:12,896 --> 01:04:17,933
An identity marker can slide, disappear, or change shape at another moment.
身份标记可能在另一个时刻滑动、消失或变形。

736
01:04:17,933 --> 01:04:24,592
Selected frames let us ask the right questions, but the full clip is needed to check continuity between them.
精选帧让我们提出正确问题，但要检查帧间连续性，仍需完整视频。

737
01:04:25,176 --> 01:04:32,652
For our discussion, identify a small recognizable feature that will make identity drift easier to notice.
请为讨论找一个可辨认的小特征，让身份漂移更容易被发现。

738
01:04:33,252 --> 01:04:42,904
Choose a distinctive feature and follow it through the shot: its attachment to the subject, its shape, and its reappearance after partial visibility.
选一个独特特征，并在镜头中追踪它：与主体的附着关系、形状，以及部分遮挡后的再次出现。

739
01:04:43,488 --> 01:04:46,642
For our pear, the blue stem is a useful witness.
对这只梨而言，蓝色柄是有用的见证者。

740
01:04:47,226 --> 01:04:50,818
It may become hidden as the camera moves; we should allow that.
随着相机移动，它可能被遮住；这是应当允许的。

741
01:04:51,402 --> 01:04:54,629
When it returns, it should still belong to the same object.
当它返回时，仍应属于同一个物体。

742
01:04:54,629 --> 01:05:02,105
A marker that slides across the surface or reappears in a new shape tells us something a general impression of smooth motion might miss.
如果标记在表面滑动，或以新形状返回，就揭示了整体流畅感可能掩盖的问题。

743
01:05:02,777 --> 01:05:11,158
We are using the marker as a diagnostic aid, while still judging the whole object's identity and the shot's intended motion.
我们用标记辅助诊断，同时仍评价整个物体的身份和镜头预期运动。

744
01:05:11,759 --> 01:05:16,139
Here the instruction changes the whole sequence into a minimal monochrome sketch.
这里的指令把整个序列变成极简单色素描。

745
01:05:16,943 --> 01:05:22,287
Local object replacement and global stylization allow different degrees of visual freedom.
局部物体替换与整体风格化，允许的视觉自由程度不同。

746
01:05:22,958 --> 01:05:27,924
Even with a large style change, motion and scene structure should remain readable.
即使风格变化很大，运动和场景结构也应保持可读。

747
01:05:28,595 --> 01:05:34,962
Look for rules across frames: line density, silhouette treatment, and the handling of depth.
寻找跨帧规则：线条密度、轮廓处理，以及深度的表现方式。

748
01:05:35,764 --> 01:05:38,188
Do they feel like one visual language?
它们是否像同一种视觉语言？

749
01:05:38,992 --> 01:05:42,014
This connects directly to our collage discussion.
这与我们的拼贴讨论直接相连。

750
01:05:42,014 --> 01:05:48,219
A style is more useful to an art director when described through operations that can persist through time.
当风格被描述为能够持续存在于时间中的操作时，它对艺术总监更有用。

751
01:05:48,891 --> 01:05:56,032
If every frame reinvents those operations, the result may feel unstable even when each still is appealing.
如果每一帧都重新发明这些操作，即使每张静帧好看，整体也可能不稳定。

752
01:05:56,632 --> 01:05:59,903
A video can preserve object motion while its style flickers.
视频可以保留物体运动，风格却仍然闪烁。

753
01:06:00,486 --> 01:06:09,058
Line density may jump, a paper texture may crawl, or shading may switch between flat and volumetric treatments without a scene explanation.
线条密度可能跳变，纸张纹理可能爬动，明暗处理也可能无故在平面与立体之间切换。

754
01:06:09,729 --> 01:06:15,088
Review style as a set of temporal rules, just as we reviewed it across surfaces in the collage.
像审查拼贴不同表面的风格那样，把视频风格作为一组时间规则来审查。

755
01:06:15,760 --> 01:06:19,819
Some variation is appropriate when the camera or light changes.
相机或光线变化时，某些变化是合理的。

756
01:06:19,819 --> 01:06:23,060
The question is whether the material language remains coherent.
问题是：材质语言是否仍然连贯？

757
01:06:23,645 --> 01:06:32,056
If your artwork depends on a particular drawing or collage treatment, this check is as important as whether the subject stays in the same place.
如果作品依赖特定绘画或拼贴处理，这项检查与主体位置是否稳定同样重要。

758
01:06:32,859 --> 01:06:36,874
Motion correctness alone does not establish visual consistency.
运动正确，本身并不能证明视觉一致。

759
01:06:37,474 --> 01:06:44,177
If we subtract one frame from the next, a perfectly correct camera movement can create a large difference.
如果将相邻帧相减，完全正确的相机运动也会产生很大差异。

760
01:06:44,760 --> 01:06:48,440
The same surface has simply moved to another pixel location.
同一个表面，只是移动到了另一个像素位置。

761
01:06:49,361 --> 01:06:56,807
A more meaningful comparison first aligns corresponding visible content, then measures the remaining appearance difference.
更有意义的比较，是先对齐对应的可见内容，再测量剩余外观差异。

762
01:06:57,230 --> 01:07:04,633
The teaching equation uses a motion warp for that alignment and a visibility mask to exclude occluded regions.
这个教学公式使用运动变换来对齐，并用可见性蒙版排除被遮挡区域。

763
01:07:04,634 --> 01:07:08,752
We should not demand agreement for content that is hidden or newly revealed.
对于隐藏或刚被揭示的内容，不应要求直接一致。

764
01:07:09,423 --> 01:07:14,826
This is an illustrative metric, not a claim about the exact evaluation used by the paper.
这是示意性指标，不是在声称论文使用了这一精确评测。

765
01:07:15,629 --> 01:07:19,061
A low difference can also reward a video that barely moves.
低差异也可能奖励几乎不动的视频。

766
01:07:19,572 --> 01:07:23,616
A metric must be checked against the task it is supposed to represent.
必须对照指标所要代表的任务，检查指标本身。

767
01:07:24,216 --> 01:07:27,049
Imagine two candidates for our five-second shot.
想象五秒镜头的两个候选版本。

768
01:07:27,720 --> 01:07:32,087
One shows a coherent camera move around the pear with modest frame differences.
一个展示相机连贯地绕梨移动，帧间差异适中。

769
01:07:32,759 --> 01:07:35,197
The other repeats a single beautiful frame.
另一个反复显示同一张漂亮图像。

770
01:07:35,781 --> 01:07:38,877
A naive difference score might prefer the frozen version.
简单的差异分数，可能偏爱冻结版本。

771
01:07:39,549 --> 01:07:43,272
Our audience would immediately notice that the intended reveal never happened.
观众却会立刻发现，预期的揭示根本没有发生。

772
01:07:43,784 --> 01:07:48,163
That is a useful example of optimizing the measurement while missing the purpose.
这是优化测量指标却错失目的的一个例子。

773
01:07:48,163 --> 01:07:52,106
We need evidence for both coherent appearance and requested motion.
我们需要外观连贯和所要求运动两方面的证据。

774
01:07:52,690 --> 01:07:54,691
Neither can substitute for the other.
两者不能相互替代。

775
01:07:55,202 --> 01:08:00,137
Each time, one easy-to-observe quality tempted us to stand in for the whole brief.
每一次，我们都容易让某个易于观察的品质，代表整份创作要求。

776
01:08:00,808 --> 01:08:07,672
Good evaluation keeps the intended result in view, especially when a convenient score seems reassuring.
良好的评测始终关注预期结果，尤其当方便的分数令人安心时。

777
01:08:08,272 --> 01:08:11,952
An edited keyframe can make an art direction easier to approve.
编辑后的关键帧，可以让艺术方向更容易获得认可。

778
01:08:12,462 --> 01:08:20,479
Instead of describing every material and lighting choice in words, we can point to an image and say, this is the appearance we want.
不必用文字描述每个材质和灯光决定，我们可以指着一张图说：这就是想要的外观。

779
01:08:21,062 --> 01:08:24,800
The workflow then carries that approved look through the source clip.
工作流程再把获批的外观延续到原始片段中。

780
01:08:24,801 --> 01:08:35,782
Follow the four stages: select a representative frame, edit and inspect its appearance, propagate the look, and review the resulting motion.
请跟随四个阶段：选代表帧，编辑并检查外观，传播外观，再审查运动结果。

781
01:08:36,366 --> 01:08:38,819
Will the material persist during a turn?
转向时，材质能保持吗？

782
01:08:39,403 --> 01:08:42,746
Will the identity survive a column passing in front of it?
柱子从前方经过时，身份能保留吗？

783
01:08:44,002 --> 01:08:52,077
This distinction is central to the Runway reading later: an appealing interaction pattern still needs a shot review suited to the artwork.
这个区别是后面 Runway 阅读的核心：吸引人的交互方式，仍需要适合这件作品的镜头审查。

784
01:08:52,677 --> 01:08:57,204
What if the approved image is so influential that the requested event never happens?
如果获批图像影响力太强，以至于要求的事件根本没发生，怎么办？

785
01:08:57,787 --> 01:09:02,358
Our AlignVid collaboration studies that tension in image-to-video generation.
我们合作开展的 AlignVid 研究，探索图生视频中的这种张力。

786
01:09:02,943 --> 01:09:12,433
In the published examples, a baseline omits a sunflower or leaves a person standing; the corresponding AlignVid results implement more of the requested event.
论文示例中，基线漏掉了向日葵，或让人物一直站着；对应的 AlignVid 结果实现了更多要求的事件。

787
01:09:13,105 --> 01:09:19,310
The intervention scales queries or keys in selected attention blocks and denoising steps without retraining.
这种干预无需重新训练，只在选定注意力模块和去噪步骤中，缩放查询或键。

788
01:09:19,311 --> 01:09:27,940
In a scalar form, scaling Q by gamma changes the weights to softmax of gamma times Q K transpose over square root d.
在标量缩放形式下，用 gamma 缩放 Q 后，权重变成 softmax(gamma × QKᵀ / √d)。

789
01:09:28,525 --> 01:09:30,847
It changes attention concentration.
它改变了注意力的集中程度。

790
01:09:31,357 --> 01:09:38,148
Classifier-free guidance instead combines predictions with different conditioning; these are distinct operations.
无分类器引导则组合不同条件下的预测；两者是不同操作。

791
01:09:38,731 --> 01:09:44,602
For an artist, the interesting failure is a faithful-looking image that refuses to do what the scene requires.
对艺术家来说，有意思的失败是：图像看起来很忠实，却拒绝做场景要求的事。

792
01:09:44,602 --> 01:09:49,625
The published frames illustrate that tension; a finished shot still needs temporal review.
已发表的帧展示了这种张力；完成的镜头仍需进行时间连续性审查。

793
01:09:50,225 --> 01:09:57,614
Our five-second brief changes the pear's body to ivory ceramic while retaining the blue stem, camera movement, and position.
我们的五秒要求，是把梨的主体变为象牙白陶瓷，同时保留蓝色柄、相机运动和位置。

794
01:09:58,197 --> 01:10:00,592
Reflections may adapt to the new material.
倒影可以适应新材质。

795
01:10:00,957 --> 01:10:04,784
The pear must remain the same object after passing behind a column.
梨经过柱子后面再出现时，必须仍是同一个物体。

796
01:10:05,368 --> 01:10:07,733
Which clause is hardest to verify?
哪一条最难验证？

797
01:10:08,536 --> 01:10:14,931
The reappearance is a strong candidate, because the system must maintain identity through a period of invisibility.
再次出现很可能最难，因为系统必须在一段不可见期间保持身份。

798
01:10:14,931 --> 01:10:19,327
This is more demanding than transferring a visible color from frame to frame.
这比逐帧传递可见颜色更具挑战。

799
01:10:19,910 --> 01:10:23,036
Begin with one clear transformation in a short clip.
先在短片中做一个明确的变换。

800
01:10:23,619 --> 01:10:31,227
Then make the critical visibility event part of the review, rather than discovering it only after choosing a favorite result.
然后把关键可见性事件纳入审查，而不是选好最喜欢的结果后才发现它。

801
01:10:31,827 --> 01:10:33,930
The column is the moment of truth.
柱子就是关键考验。

802
01:10:34,732 --> 01:10:39,303
Before the pear disappears, we can inspect its silhouette, material, and stem.
梨消失前，我们能检查它的轮廓、材质和柄。

803
01:10:39,888 --> 01:10:42,968
During occlusion, there may be nothing visible to compare.
遮挡期间，可能没有可见内容可供比较。

804
01:10:43,480 --> 01:10:47,465
When it returns, the model has to make the same object convincing again.
返回时，模型必须再次让同一个物体显得可信。

805
01:10:48,050 --> 01:10:51,335
Do not reject correct invisibility as a failure.
不要把正确的不可见状态当作失败。

806
01:10:51,335 --> 01:10:57,979
Instead, compare the identity before and after the event, including how the object emerges at the boundary.
而应比较事件前后的身份，包括物体如何从边界处显露。

807
01:10:58,490 --> 01:11:03,381
A smooth-looking clip can still reveal a subtly redesigned pear on the other side.
看起来流畅的视频，也可能在另一侧露出一只被悄悄重新设计的梨。

808
01:11:03,966 --> 01:11:12,246
We are not merely watching for anything strange; we are testing the preservation requirement at the point where the shot makes it hardest to satisfy.
我们不只是寻找怪异之处，而是在镜头最难满足要求的地方，检验保留约束。

809
01:11:12,845 --> 01:11:16,043
Adding a small object creates several new relationships.
添加一个小物体，会创造多种新关系。

810
01:11:16,627 --> 01:11:19,576
In this published example, the edit adds a drone.
在这个已发表示例中，编辑添加了一架无人机。

811
01:11:20,161 --> 01:11:26,687
Even if the drone looks convincing by itself, its scale, placement, and movement must fit the scene.
即使无人机本身很可信，它的尺度、位置和运动也必须适合场景。

812
01:11:27,359 --> 01:11:32,265
Imagine a camera moving forward while the added object changes size in the wrong direction.
想象相机向前移动，新增物体的大小却朝错误方向变化。

813
01:11:32,937 --> 01:11:38,311
The object might be beautifully rendered, but its relationship with the camera would expose the edit.
物体可能渲染得很漂亮，但它与相机的关系会暴露编辑。

814
01:11:38,311 --> 01:11:42,486
Or it might hover at an unintended location relative to other objects.
或者，它可能相对其他物体悬停在错误位置。

815
01:11:42,998 --> 01:11:47,160
When you review an addition, evaluate those relationships across time.
审查新增物体时，要跨时间评价这些关系。

816
01:11:47,671 --> 01:11:52,679
A close-up of the inserted object cannot answer every question about whether it belongs.
仅看新增物体的特写，无法回答它是否融入场景的所有问题。

817
01:11:53,279 --> 01:11:56,477
An added object must fit the camera as well as the scene.
新增物体既要符合场景，也要符合相机。

818
01:11:57,061 --> 01:12:02,171
A convincing texture and shape in one still do not establish a coherent trajectory.
单张静帧中可信的纹理和形状，不能证明运动轨迹连贯。

819
01:12:02,844 --> 01:12:09,122
If the camera approaches, the apparent size and position of the addition should evolve in a compatible way.
如果相机靠近，新增物体的表观大小和位置应相应变化。

820
01:12:09,926 --> 01:12:15,999
The exact expectation depends on whether the object is stationary or moving independently.
具体预期取决于物体是静止，还是独立运动。

821
01:12:16,000 --> 01:12:19,417
That is why the brief should specify the intended relationship.
因此，创作说明应该明确预期关系。

822
01:12:19,927 --> 01:12:25,549
Inspect the addition relative to nearby objects and the background, not only in a crop.
要相对附近物体和背景检查新增内容，而不只看裁剪图。

823
01:12:26,352 --> 01:12:35,391
A production-quality edit is a collection of relationships that survive over time, rather than a new object that looks impressive in isolation.
制作级编辑是一组经得起时间考验的关系，而不是一个单独看很惊艳的新物体。

824
01:12:35,991 --> 01:12:40,429
Review the whole shot at normal speed, then inspect moments where failure is most likely.
先以正常速度看完整个镜头，再检查最可能失败的时刻。

825
01:12:41,014 --> 01:12:45,847
Pay attention to identity, intended motion, occlusion, boundaries, and the ending.
关注身份、预期运动、遮挡、边界和结尾。

826
01:12:46,431 --> 01:12:50,900
High-motion content is a limitation the research authors specifically identify.
高运动量内容，是研究作者明确指出的局限。

827
01:12:51,410 --> 01:12:54,915
If a clip fails, choose a revision based on the failure.
片段失败时，要根据具体失败选择修改。

828
01:12:54,915 --> 01:13:04,450
You might reduce the transformation, shorten the segment, add a reference frame where supported, or repair part of the result through conventional compositing.
可以减小变换幅度、缩短片段、在支持时添加参考帧，或用传统合成修复部分结果。

829
01:13:05,253 --> 01:13:09,079
Another generation is useful when it tests a reasoned hypothesis.
当新一轮生成用于检验有理由的假设时，它才有用。

830
01:13:09,502 --> 01:13:16,350
Do not let repeated attempts distract you from whether the piece still communicates the artistic intention you began with.
不要让反复尝试分散注意力，忘记作品是否仍在传达最初的艺术意图。

831
01:13:16,951 --> 01:13:19,550
A failed shot is information about the workflow.
失败镜头提供了关于工作流程的信息。

832
01:13:20,222 --> 01:13:25,873
If the appearance is wrong from the beginning, revising the keyframe or the material direction may help.
如果一开始外观就不对，修改关键帧或材质方向可能有帮助。

833
01:13:26,675 --> 01:13:35,918
If the first frame is convincing but identity breaks after occlusion, the problem calls for temporal review and a more suitable propagation or editing approach.
如果首帧可信，但遮挡后身份破坏，就需要时间连续性审查，以及更合适的传播或编辑方法。

834
01:13:36,503 --> 01:13:41,189
If a region must remain exact, direct compositing may be part of the solution.
如果某个区域必须完全不变，直接合成可以是解决方案的一部分。

835
01:13:41,190 --> 01:13:47,994
We can also split a complicated transformation into stages, provided the transitions remain coherent.
也可以把复杂变换拆成多个阶段，但必须保证过渡连贯。

836
01:13:48,798 --> 01:13:52,915
The important move is to connect the observed failure to the next intervention.
关键是把观察到的失败，与下一步干预联系起来。

837
01:13:53,718 --> 01:13:59,063
Repeating a vague request with more enthusiasm does not use what the failed shot has taught us.
更热情地重复模糊要求，并没有利用失败镜头教给我们的东西。

838
01:13:59,662 --> 01:14:02,641
Return to the proposed five-second ceramic-pear shot.
回到拟议的五秒陶瓷梨镜头。

839
01:14:03,153 --> 01:14:05,459
Which moment would you inspect most closely?
你会最仔细检查哪个时刻？

840
01:14:05,824 --> 01:14:08,380
What should remain recognizable after the column?
经过柱子后，什么应该仍然可辨认？

841
01:14:08,803 --> 01:14:11,607
What failure would make you reject an attractive clip?
什么失败会让你拒绝一段好看的视频？

842
01:14:12,410 --> 01:14:16,834
Discuss these three questions by describing an event and the evidence you would seek.
请通过描述事件和所需证据，讨论这三个问题。

843
01:14:17,505 --> 01:14:27,230
If your answer is temporal consistency, make it visible: what changes, when, and why would that violate the intended shot?
如果你的答案是“时间一致性”，请让它可见：什么在何时改变，为什么违反了镜头意图？

844
01:14:27,231 --> 01:14:28,618
Pause the video here.
请在这里暂停视频。

845
01:14:29,290 --> 01:14:34,766
Compare your acceptance criteria with another person’s before we examine the compact review plan.
在查看简要审查方案之前，先与另一位同学比较验收标准。

846
01:14:38,765 --> 01:14:40,517
Here is a compact test plan.
这是一个简要测试方案。

847
01:14:41,029 --> 01:14:44,708
The request is an ivory ceramic body with the blue stem retained.
要求是象牙白陶瓷主体，并保留蓝色柄。

848
01:14:45,293 --> 01:14:48,461
Camera motion and object motion should follow the source.
相机运动和物体运动应遵循原片。

849
01:14:49,045 --> 01:14:53,120
The clip includes a column so that we can inspect a difficult reappearance.
片段中包含柱子，让我们检查困难的再次出现过程。

850
01:14:53,791 --> 01:14:58,523
The evidence includes normal-speed playback and a closer look before and after that event.
证据包括正常速度播放，以及事件前后的仔细观察。

851
01:14:59,194 --> 01:15:04,173
We also inspect reflections because a material change has optical consequences.
还要检查倒影，因为材质变化会产生光学后果。

852
01:15:04,173 --> 01:15:07,356
This plan does not require a long benchmark suite.
这个方案不需要庞大的基准测试集。

853
01:15:08,028 --> 01:15:13,678
It simply makes the request, the likely failure, and the acceptance evidence explicit.
它只是明确要求、可能的失败，以及验收证据。

854
01:15:14,482 --> 01:15:19,241
You can use the same structure to test a different subject or visual transformation.
你可以用同样的结构测试其他主体或视觉变换。

855
01:15:19,842 --> 01:15:24,017
Let us check whether the mechanisms are now useful without a model in front of us.
现在没有模型在面前，让我们检验这些机制是否真的有用了。

856
01:15:24,602 --> 01:15:26,807
Consider the three questions on screen.
请考虑屏幕上的三个问题。

857
01:15:27,391 --> 01:15:31,625
A material change reaches the reflection: explain why.
材质变化会影响倒影：请解释原因。

858
01:15:32,297 --> 01:15:38,182
A subject reference and LoRA both help represent a concept: explain the difference.
主体参考图和 LoRA 都有助于表达概念：请解释两者区别。

859
01:15:39,263 --> 01:15:43,629
Four frames look excellent: explain why the clip can still fail.
四帧看起来很出色：请解释整段视频为什么仍可能失败。

860
01:15:44,548 --> 01:15:47,293
Take a short silent beat before answering.
回答前，先安静思考片刻。

861
01:15:47,293 --> 01:15:50,870
The next slide names the decisions hiding inside the questions.
下一页会指出这些问题中隐藏的决定。

862
01:15:51,455 --> 01:15:59,472
If you get stuck, locate the type of problem first: the physical scene, the model's information, or evidence across time.
如果卡住了，先定位问题类型：物理场景、模型获得的信息，还是跨时间的证据。

863
01:16:00,055 --> 01:16:02,888
That is often enough to begin a clear answer.
通常，做到这一步就足以开始给出清晰答案。

864
01:16:03,488 --> 01:16:06,146
Here are the decisions those questions conceal.
这些问题隐藏着以下决定。

865
01:16:06,729 --> 01:16:13,871
If the reflection must change with the new material, we need to define the permitted region and consequences of the edit.
如果倒影必须随新材质变化，就需要定义编辑允许的区域和连带后果。

866
01:16:14,673 --> 01:16:24,836
If we want a concept for this request, a reference can condition generation; if we want learned adaptation, LoRA changes selected weights.
如果只为这次要求表达概念，参考图可提供生成条件；如果需要学习式适配，LoRA 会改变选定权重。

867
01:16:25,420 --> 01:16:30,998
If continuity matters, the evidence must include the intervals between our chosen stills.
如果重视连续性，证据就必须包含精选静帧之间的间隔。

868
01:16:30,998 --> 01:16:33,641
Each decision rules out a tempting shortcut.
每个决定，都排除了一条诱人的捷径。

869
01:16:34,225 --> 01:16:37,875
Freezing every surrounding pixel may freeze the wrong reflection.
冻结周围每一个像素，可能冻结错误的倒影。

870
01:16:38,547 --> 01:16:43,133
Calling every supplied image training confuses conditioning with adaptation.
把所有输入图像都称为训练，会混淆条件控制与适配。

871
01:16:43,804 --> 01:16:47,091
Calling four stills a successful video omits time.
把四张静帧称为成功视频，就遗漏了时间。

872
01:16:47,893 --> 01:16:53,310
Use this slide to repair your explanation, rather than memorize a preferred sentence.
请用这一页修正你的解释，而不是背诵某句标准答案。

873
01:16:53,910 --> 01:17:00,612
The material changes how light interacts with the pear, so its visible consequences can extend into the reflection.
材质改变光与梨的相互作用，因此可见后果会延伸到倒影。

874
01:17:01,416 --> 01:17:07,140
A reference conditions an output; LoRA learns a low-rank update to selected model weights.
参考图为输出提供条件；LoRA 为选定模型权重学习低秩更新。

875
01:17:07,723 --> 01:17:14,251
Four good frames leave unobserved transitions, including events such as occlusion where identity can fail.
四张好帧仍留下未观察的过渡，包括可能导致身份失败的遮挡事件。

876
01:17:14,922 --> 01:17:18,382
Those are compact answers, but each has an application.
这些答案很简短，但每个都有实际用途。

877
01:17:18,383 --> 01:17:26,267
They help us write a better preservation contract, choose between two kinds of intervention, and design a more revealing review.
它们帮助我们写出更好的保留约定、在两类干预中选择，并设计更有揭示性的审查。

878
01:17:26,852 --> 01:17:31,160
If your answer used different words and preserved those distinctions, it works.
如果你用了不同措辞，但保留了这些区别，答案就成立。

879
01:17:31,743 --> 01:17:41,103
Tomorrow's interface may rename its controls, but the difference between specifying a request, adapting a model, and checking its result will still matter.
明天的界面可能给控件改名，但提出要求、适配模型与检查结果之间的区别，仍然重要。

880
01:17:41,703 --> 01:17:43,762
We can check an edit at three levels.
我们可以在三个层面检查编辑。

881
01:17:43,923 --> 01:17:46,901
First, did the requested transformation occur?
第一，所要求的变换发生了吗？

882
01:17:47,486 --> 01:17:50,259
Second, did the necessary invariants survive?
第二，必要的不变量保留了吗？

883
01:17:50,932 --> 01:17:54,231
Third, does the result serve the artwork's intention?
第三，结果是否服务于作品意图？

884
01:17:54,816 --> 01:18:01,971
A result can pass the first level and fail the second, as when the correct material appears on a different object.
结果可能通过第一层却不通过第二层，例如正确材质出现在了另一个物体上。

885
01:18:01,971 --> 01:18:09,579
It can pass both technical levels and still fail the artistic one, as when the chosen effect undermines the intended mood.
也可能通过两个技术层面，却在艺术层面失败，例如所选效果削弱了预期情绪。

886
01:18:10,250 --> 01:18:13,594
Keeping the levels separate makes critique more precise.
区分这些层面，会让评论更精确。

887
01:18:14,178 --> 01:18:22,122
It also explains why neither a generic aesthetic score nor a strict pixel comparison can answer every question we care about.
这也解释了为什么通用审美分数和严格像素比较，都无法回答我们关心的所有问题。

888
01:18:22,721 --> 01:18:26,532
We have spent the lecture making the editing request more precise.
整堂课，我们都在让编辑要求更精确。

889
01:18:27,204 --> 01:18:30,548
Now consider who gets to choose the request in the first place.
现在想想：一开始由谁来选择提出什么要求？

890
01:18:31,132 --> 01:18:40,303
A system can offer convincing alternatives, but somebody still decides which intention matters and which evidence counts as success.
系统可以提供可信方案，但仍需要有人决定什么意图重要，以及什么证据才算成功。

891
01:18:41,105 --> 01:18:45,399
The three readings let us examine that responsibility from different positions.
三篇阅读让我们从不同位置审视这种责任。

892
01:18:46,071 --> 01:18:51,107
Pachocki raises questions about capable AI systems and human values.
Pachocki 提出了能力强大的 AI 系统与人类价值观之间的问题。

893
01:18:51,107 --> 01:18:54,554
Koe asks readers to reconsider their own goals and habits.
Koe 请读者重新审视自己的目标与习惯。

894
01:18:55,064 --> 01:18:59,255
Runway presents a workflow for turning an approved image into a video edit.
Runway 展示把获批图像转化为视频编辑的工作流程。

895
01:18:59,840 --> 01:19:07,578
Keep our pear in mind as we read them: we can delegate parts of its production while still arguing about what the artwork should become.
阅读时请记住这只梨：我们可以委托部分制作，同时仍讨论作品应该成为什么。

896
01:19:08,178 --> 01:19:11,245
We have an approved appearance and a moving shot to judge.
现在，我们有了获批外观，以及需要评价的运动镜头。

897
01:19:11,916 --> 01:19:15,552
What can an approved keyframe establish, and what can it not?
获批关键帧能确定什么，又不能确定什么？

898
01:19:16,224 --> 01:19:19,670
Why might a low frame-difference score reward a bad video?
为什么低帧差分数可能奖励一段坏视频？

899
01:19:20,094 --> 01:19:23,554
Which evidence would convince you that identity survived motion?
什么证据能说服你，身份经受住了运动？

900
01:19:24,357 --> 01:19:26,125
Discuss these three questions.
请讨论这三个问题。

901
01:19:26,284 --> 01:19:30,841
Your answer should account for the requested event as well as the subject’s appearance.
答案既要考虑主体外观，也要考虑所要求的事件。

902
01:19:30,841 --> 01:19:34,651
Use the column, a turn, or the sketch treatment as a concrete example.
可以用柱子、转向或素描处理作为具体例子。

903
01:19:35,236 --> 01:19:36,506
Pause the video here.
请在这里暂停视频。

904
01:19:36,666 --> 01:19:41,296
We will next ask how the readings change our view of responsibility for those decisions.
接下来，我们将追问阅读如何改变我们对这些决定所涉及责任的看法。

905
01:19:45,296 --> 01:19:49,020
These first two readings ask about direction from different sides.
前两篇阅读从不同方面追问方向。

906
01:19:49,530 --> 01:19:56,948
An Alien Mind considers the relationship between an AI system achieving a goal and respecting human values.
《An Alien Mind》思考 AI 系统实现目标与尊重人类价值观之间的关系。

907
01:19:57,532 --> 01:20:02,365
Dan Koe's essay invites people to examine the goals and habits directing their own lives.
Dan Koe 的文章邀请人们审视引导自己生活的目标和习惯。

908
01:20:02,526 --> 01:20:07,373
We will use its proposal as material for critical discussion of creative practice.
我们把他的主张作为批判性讨论创作实践的材料。

909
01:20:07,374 --> 01:20:13,331
Find a specific claim, explain your interpretation, and connect it to a decision from the lecture.
找出一个具体主张，解释你的理解，并联系课堂中的一个决定。

910
01:20:14,003 --> 01:20:17,624
A fictional artist is fine for discussing personal direction.
讨论个人方向时，可以使用虚构艺术家的例子。

911
01:20:18,295 --> 01:20:22,398
The fifth quiz question asks you to reflect on one reading of your choice.
测验第五题要求你从自选的一篇阅读出发进行反思。

912
01:20:23,071 --> 01:20:29,860
Our class discussion can compare all three perspectives without requiring everyone to disclose personal experiences.
课堂可以比较三种视角，而不要求每个人披露个人经历。

913
01:20:30,461 --> 01:20:35,221
The third reading moves from questions about direction to a concrete editing workflow.
第三篇阅读从方向问题，转向具体编辑流程。

914
01:20:36,023 --> 01:20:45,295
Runway's announcement of Aleph 2.0 and Edit Studio describes establishing an appearance in an edited image and applying that change through video.
Runway 关于 Aleph 2.0 和 Edit Studio 的公告，描述了先在编辑图像中确定外观，再把变化应用到视频的方法。

915
01:20:45,880 --> 01:20:48,421
It connects directly to our keyframe discussion.
它与关键帧讨论直接相连。

916
01:20:48,581 --> 01:20:51,297
Read it with the ceramic pear and column in mind.
阅读时，请想着陶瓷梨和柱子。

917
01:20:51,968 --> 01:20:55,576
Which decision does the interface make easier to express?
这个界面让哪项决定更容易表达？

918
01:20:55,576 --> 01:20:58,554
Which result would you still need to watch before approval?
批准之前，哪些结果仍需要亲眼播放检查？

919
01:20:59,139 --> 01:21:08,309
The page is a vendor's account, so its demonstrations illustrate claimed capabilities rather than supply an independent comparative evaluation.
这是一篇厂商介绍，因此演示说明的是其声称的能力，而不是独立的比较评测。

920
01:21:08,980 --> 01:21:17,012
We can learn from the interaction design and still propose an occlusion test that asks whether the workflow meets this exhibition's particular needs.
我们可以借鉴交互设计，同时提出遮挡测试，检验流程是否满足这个展览的具体需要。

921
01:21:17,611 --> 01:21:21,320
These technical papers are optional companions to the discussion readings.
这些技术论文，是讨论阅读之外的可选配套材料。

922
01:21:21,831 --> 01:21:24,402
Choose according to the question you want to investigate.
请根据你想研究的问题选择。

923
01:21:24,985 --> 01:21:31,703
InstructPix2Pix helps explain edit supervision; Flow Matching develops the generative training framework.
InstructPix2Pix 帮助解释编辑监督；Flow Matching 发展了生成训练框架。

924
01:21:32,374 --> 01:21:37,718
Qwen-Image-2.0 and Qwen-Video-Edit provide recent architecture examples.
Qwen-Image-2.0 和 Qwen-Video-Edit 提供了近期架构示例。

925
01:21:38,230 --> 01:21:42,026
You do not need to read all of them before trying the class discussion.
参加课堂讨论之前，不必读完全部论文。

926
01:21:42,026 --> 01:21:49,561
Start with a question, locate the part of the paper that addresses it, and distinguish the proposed method from the authors' evidence.
从一个问题开始，找到论文中回应它的部分，并区分提出的方法与作者的证据。

927
01:21:50,480 --> 01:21:58,205
The slide notes retain the other references on multi-image inputs, rewards, composition, and video editing.
幻灯片备注保留了多图输入、奖励、构图和视频编辑方面的其他参考资料。

928
01:21:58,805 --> 01:22:01,608
The three readings provide different kinds of material.
三篇阅读提供不同类型的材料。

929
01:22:02,280 --> 01:22:07,055
A research leader's perspective develops an argument about alignment and future development.
研究领导者的观点文章，提出关于对齐和未来发展的论证。

930
01:22:07,727 --> 01:22:11,392
A reflective essay offers a way to examine personal direction.
反思性文章提供一种审视个人方向的方法。

931
01:22:12,063 --> 01:22:15,932
A vendor announcement presents a workflow and promotes its capabilities.
厂商公告展示工作流程，并推广其能力。

932
01:22:16,604 --> 01:22:18,795
Read each according to what it can support.
阅读每种材料时，要考虑它能够支持什么。

933
01:22:19,466 --> 01:22:27,249
Identify forecasts, personal claims, and demonstrations rather than treating every sentence as the same kind of evidence.
识别预测、个人主张和演示，不要把每句话都当作同类证据。

934
01:22:27,249 --> 01:22:30,739
You may disagree with an author and still find a useful question.
你可以不同意作者，却仍从中找到有用的问题。

935
01:22:31,250 --> 01:22:42,318
Our synthesis is practical: what do you want to make, what will you delegate, and what evidence will you use to decide whether the collaboration served that intention?
我们的综合问题很实际：你想创作什么、委托什么，以及用什么证据判断合作是否服务于这一意图？

936
01:22:42,918 --> 01:22:48,423
This technical extension separates image guidance and text guidance in InstructPix2Pix.
这个技术延伸，区分 InstructPix2Pix 中的图像引导和文本引导。

937
01:22:49,095 --> 01:22:51,752
Begin with the prediction using neither condition.
从不使用任何条件的预测开始。

938
01:22:52,336 --> 01:22:59,113
Add a scaled difference for including the image, then another scaled difference for adding the text alongside that image.
加上引入图像后的缩放差值，再加上在图像之外引入文本后的另一个缩放差值。

939
01:22:59,623 --> 01:23:07,566
Set both scales to one and follow the cancellation: the intermediate terms disappear, leaving the fully conditioned noise prediction.
把两个强度都设为一，观察抵消：中间项消失，只留下完全条件化的噪声预测。

940
01:23:07,566 --> 01:23:10,735
This is a useful check that you understand the expression.
这是检查你是否理解表达式的有用方法。

941
01:23:11,406 --> 01:23:17,437
Epsilon here denotes a noise prediction, unlike the velocity in our flow-matching slides.
这里的 epsilon 表示噪声预测，与流匹配页面中的速度不同。

942
01:23:18,108 --> 01:23:25,965
These controls belong to this formulation and should not be assumed to map directly onto every current editor's interface.
这些控制属于这一公式，不应假设它们直接对应所有当前编辑器的界面。

943
01:23:26,565 --> 01:23:30,011
Now set both guidance scales to one and follow the cancellation.
现在把两个引导强度都设为一，观察抵消。

944
01:23:30,815 --> 01:23:35,560
The initial no-condition prediction cancels its negative copy in the first difference.
初始无条件预测，与第一个差值中的负项抵消。

945
01:23:36,143 --> 01:23:40,787
The image-only prediction then cancels its negative copy in the second difference.
随后，仅图像条件的预测，与第二个差值中的负项抵消。

946
01:23:41,458 --> 01:23:45,227
The result is the prediction with both image and text conditions.
最终得到同时使用图像和文本条件的预测。

947
01:23:45,737 --> 01:23:49,519
That gives us a reference point for understanding the two controls.
这为理解两个控制提供了参考点。

948
01:23:49,519 --> 01:23:56,484
If the text scale were zero while the image scale stayed one, we would instead recover the image-only prediction.
如果文本强度为零、图像强度仍为一，就会得到仅使用图像条件的预测。

949
01:23:57,068 --> 01:23:59,901
It shows what information each difference adds.
这说明了每个差值添加什么信息。

950
01:24:00,485 --> 01:24:10,896
Keep epsilon's meaning explicit here: this InstructPix2Pix expression predicts noise, whereas the earlier flow-matching example predicted velocity.
请明确 epsilon 的含义：这里的 InstructPix2Pix 公式预测噪声，而前面的流匹配例子预测速度。

951
01:24:11,496 --> 01:24:13,540
Read the pseudocode in two groups.
把伪代码分成两组来读。

952
01:24:14,211 --> 01:24:22,345
First we prepare an example: source, instruction, target, conditions, and target latent.
首先准备样本：原图、指令、目标、条件，以及目标潜变量。

953
01:24:23,149 --> 01:24:32,334
Then we sample noise and time, construct the intermediate latent, predict velocity, and compare that prediction with the known training target.
然后采样噪声和时间，构造中间潜变量，预测速度，并将预测与已知训练目标比较。

954
01:24:32,917 --> 01:24:36,495
The final update changes trainable weights based on the loss.
最后根据损失更新可训练权重。

955
01:24:37,079 --> 01:24:41,342
At inference, there is no known edited target to supply in this way.
推理时，并没有已知编辑目标可以这样提供。

956
01:24:41,342 --> 01:24:45,358
We instead follow the learned generative predictions from a starting state.
我们从初始状态出发，遵循学习得到的生成预测。

957
01:24:46,161 --> 01:24:54,279
This pseudocode explains the conceptual loop; it omits practical engineering such as batches, precision, and scheduling.
这段伪代码解释概念循环，省略了批处理、精度和调度等实际工程细节。

958
01:24:55,083 --> 01:25:00,617
Its most important distinction is what information is available during learning versus use.
最重要的区别，是学习与使用时分别有哪些信息可用。

959
01:25:01,217 --> 01:25:02,984
Look for the line that disappeared.
请找出消失的那一行。

960
01:25:03,349 --> 01:25:07,963
The inference loop has no known edited target and no loss that updates weights.
推理循环没有已知编辑目标，也没有用于更新权重的损失。

961
01:25:08,474 --> 01:25:18,083
Instead, we encode the source and instruction, begin with noise, repeatedly predict a velocity and move the latent, then decode the result.
我们编码原图和指令，从噪声开始，反复预测速度并移动潜变量，最后解码结果。

962
01:25:18,666 --> 01:25:21,776
Training uses examples to adjust the learned model.
训练用样本调整学习到的模型。

963
01:25:21,776 --> 01:25:27,602
In this simplified inference loop, we use that model to construct an output we do not yet possess.
在这个简化推理循环中，我们用该模型构造尚未拥有的输出。

964
01:25:28,026 --> 01:25:34,670
The code is conceptual; practical systems use their own representations and sampling schedules.
代码是概念性的；实际系统使用各自的表示和采样调度。

965
01:25:35,254 --> 01:25:44,964
But you can now explain why providing another reference changes the information available for a request without automatically becoming a model-training operation.
但你现在可以解释：增加参考图会改变请求可用的信息，却不会自动变成模型训练。

966
01:25:44,964 --> 01:25:49,696
It is the same distinction we used when choosing between references and LoRA.
这与选择参考图还是 LoRA 时使用的区别相同。

967
01:25:50,295 --> 01:25:55,625
Before we turn to the readings' discussion questions, use this table as a compact decision aid.
转向阅读讨论问题之前，请把这张表当作简明决策辅助。

968
01:25:56,209 --> 01:26:00,385
Exact untouched pixels suggest a role for masks and compositing.
如果未编辑像素必须精确不变，可以考虑蒙版与合成。

969
01:26:00,969 --> 01:26:04,387
A new visual language may benefit from reference conditioning.
新的视觉语言，可能受益于参考图条件控制。

970
01:26:04,970 --> 01:26:08,314
A reusable learned concept may motivate adaptation.
可复用的学习概念，可能需要模型适配。

971
01:26:08,825 --> 01:26:14,009
A transformation that must survive motion needs a video workflow and temporal evidence.
必须经受运动的变换，需要视频工作流程和时间证据。

972
01:26:14,681 --> 01:26:17,864
These approaches can cooperate in the same artwork.
这些方法可以在同一件作品中协作。

973
01:26:17,864 --> 01:26:22,565
Choose according to the contract, then inspect the failure most likely to undermine it.
根据约定选择，再检查最可能破坏约定的失败。

974
01:26:23,150 --> 01:26:33,984
That leaves us with a larger question for the readings: once production becomes easier, how do we choose a worthwhile direction and retain meaningful judgment over the result?
这给阅读留下更大的问题：制作更容易之后，如何选择值得追求的方向，并对结果保留有意义的判断？

975
01:26:34,584 --> 01:26:41,242
An Alien Mind distinguishes achieving an assigned goal from generalizing human values in unfamiliar circumstances.
《An Alien Mind》区分了完成指定目标，与在陌生情境中泛化人类价值观。

976
01:26:41,914 --> 01:26:46,456
Pachocki also raises questions about monitoring increasingly capable systems.
Pachocki 也提出了如何监测日益强大系统的问题。

977
01:26:47,039 --> 01:26:51,508
Treat the essay as an argument containing claims and forecasts that we can examine.
请把文章视为包含主张与预测的论证，我们可以对其检视。

978
01:26:52,018 --> 01:27:00,517
For our discussion, imagine an exhibition team delegating production while retaining responsibility for what the work communicates.
讨论时，想象一个展览团队委托制作，但仍对作品传达的内容负责。

979
01:27:00,517 --> 01:27:11,132
Discuss the three questions on screen: explain the goal-and-values distinction, identify a forecast and the evidence it would need, and defend a boundary for human control.
讨论屏幕上的三个问题：解释目标与价值观的区别，找出一个预测及所需证据，并为人类控制的边界辩护。

980
01:27:12,389 --> 01:27:13,761
Pause the video here.
请在这里暂停视频。

981
01:27:14,346 --> 01:27:23,354
Our art-direction example is an analogy; it does not give an image editor the same agency or risk profile as an autonomous research system.
艺术指导只是类比；它并不赋予图像编辑器与自主研究系统相同的能动性或风险特征。

982
01:27:27,354 --> 01:27:32,100
Koe's title makes a dramatic promise: How to fix your entire life in one day.
Koe 的标题作出了强烈承诺：《如何在一天内修复你的整个生活》。

983
01:27:32,771 --> 01:27:40,277
The essay proposes examining identity and goals, interrupting habitual behavior, and turning reflection into action.
文章建议审视身份与目标、打断习惯行为，并把反思转化为行动。

984
01:27:41,080 --> 01:27:45,942
We can discuss the usefulness of that proposal without accepting the title as a guarantee.
我们可以讨论这些建议是否有用，而不必把标题当作保证。

985
01:27:46,746 --> 01:27:52,295
Imagine an artist who can generate a hundred attractive images but cannot choose what to make.
想象一个艺术家，能生成一百张好看的图像，却无法决定要创作什么。

986
01:27:52,295 --> 01:27:57,771
Does easy production help clarify a direction, or make avoiding the decision easier?
轻松制作会帮助明确方向，还是让逃避决定更容易？

987
01:27:58,354 --> 01:28:00,603
Discuss the three questions on screen.
请讨论屏幕上的三个问题。

988
01:28:01,186 --> 01:28:07,656
Choose an idea worth using or challenging, explain the reason, and connect it to creative intention.
选一个值得采用或质疑的观点，解释理由，并把它与创作意图联系起来。

989
01:28:08,328 --> 01:28:09,744
Pause the video here.
请在这里暂停视频。

990
01:28:10,329 --> 01:28:16,125
You can use a fictional artist or public example; no personal disclosure is needed.
可以使用虚构艺术家或公开案例，不需要披露个人经历。

991
01:28:20,125 --> 01:28:25,542
Runway's reading proposes approving an edited image before applying its appearance through a video.
Runway 的阅读提出：先批准编辑图像，再把它的外观应用到整段视频。

992
01:28:26,054 --> 01:28:34,654
It offers a concrete answer to a communication problem: an art director can point to the desired look, rather than describe every feature in words.
它具体回答了一个沟通问题：艺术总监可以指出目标外观，而不必用文字描述每个特征。

993
01:28:35,239 --> 01:28:36,684
Now bring back the column.
现在，让柱子重新出现。

994
01:28:37,267 --> 01:28:42,261
A convincing keyframe does not tell us what the pear will look like after it reappears.
可信关键帧并不能告诉我们，梨再次显露后会是什么样子。

995
01:28:42,261 --> 01:28:52,337
Discuss the three questions on screen: what the frame establishes, which claim deserves a harder test, and how the workflow serves AFTER RAIN's intention.
讨论屏幕上的三个问题：这一帧确定了什么，哪个主张需要更严格测试，以及流程如何服务于《雨后》的意图。

996
01:28:52,920 --> 01:28:54,148
Pause the video.
请暂停视频。

997
01:28:54,570 --> 01:29:03,755
The source is a vendor announcement; our job is to distinguish a useful demonstrated workflow from a reliability claim still needing evaluation.
来源是厂商公告；我们的任务是区分有用的已演示流程，与仍需评价的可靠性主张。

998
01:29:07,755 --> 01:29:10,413
Return to your first judgment of the glass pear.
回到你最初对玻璃梨的判断。

999
01:29:11,217 --> 01:29:20,124
We began with a small request and discovered that it touched the scene's physics, the artwork's intention, and the audience's experience over time.
我们从一个小要求出发，发现它触及场景物理、作品意图，以及观众随时间展开的体验。

1000
01:29:20,926 --> 01:29:26,928
You now have more precise ways to say what should change, what should survive, and how to judge the result.
现在，你有更精确的方法说明什么应改变、什么应保留，以及如何判断结果。

1001
01:29:27,599 --> 01:29:30,009
Finish with the three questions on screen.
最后，请讨论屏幕上的三个问题。

1002
01:29:30,009 --> 01:29:33,250
Where would you place the boundary between control and surprise?
你会把控制与惊喜之间的边界放在哪里？

1003
01:29:33,835 --> 01:29:37,310
How would you balance technical success and artistic purpose?
你会如何平衡技术成功与艺术目的？

1004
01:29:37,893 --> 01:29:40,990
What would you delegate, and what evidence would you require?
你会委托什么，又要求什么证据？

1005
01:29:41,573 --> 01:29:44,026
Pause the video for the final discussion.
请暂停视频，进行最终讨论。

1006
01:29:44,697 --> 01:29:48,450
Connect one mechanism or reading to a concrete artistic decision.
把一种机制或一篇阅读，与具体艺术决定联系起来。

1007
01:29:49,034 --> 01:29:52,144
Listen for an answer that makes you revise your own.
留意一个能让你修正自己看法的回答。

1008
01:29:52,145 --> 01:29:56,000
That revision is a fitting last act for a class about editing.
对一堂关于编辑的课来说，这样的修正，是恰当的最后一笔。
