SLIDE 001 · 00:00:00 · Opening
Generative Editing
00:00:00Imagine you are the art director.
想象一下,你是艺术总监。
00:00:02Your exhibition opens tomorrow.
你的展览明天开幕。
00:00:04You love this glass pear, but you want to see it in ivory ceramic.
你很喜欢这只玻璃梨,但想看看它变成象牙白陶瓷的样子。
00:00:08You give the editor a tiny request: change the material, keep everything else.
你给编辑器一个小小的要求:改变材质,其他一切保持不变。
00:00:13Then you notice the reflection.
然后,你注意到了倒影。
00:00:15Should it still look like glass?
它还应该像玻璃一样吗?
00:00:18Suddenly, a three-word edit contains a whole argument about the world.
突然,一个简单的编辑指令,包含了关于世界如何运作的一整套判断。
00:00:22That is our starting point.
这就是我们的起点。
00:00:24We will follow this fictional artwork through image editing, exhibition design, and a moving shot.
我们将跟随这件虚构作品,经历图像编辑、展览设计,再进入一段运动镜头。
00:00:31The exhibition is called AFTER RAIN.
展览名叫《雨后》,英文是 AFTER RAIN。
00:00:34The pear is our recurring character.
这只梨会反复登场。
00:00:37By the end, you should be able to explain a model's mechanism, defend a visual choice, and catch a failure that a beautiful preview can hide.
到课程结束,你应该能够解释模型机制、为视觉选择提供理由,并发现精美预览可能掩盖的失败。
00:00:46First, look carefully at what you think must survive.
首先,仔细看看你认为必须保留的东西。
SLIDE 002 · 00:00:50 · Opening
One object, three creative decisions
00:00:50We are going to give the same object three different jobs.
我们要让同一个物体承担三种不同的任务。
00:00:53First it is a technical puzzle: can its material change while its identity survives?
首先,它是一个技术难题:能否改变材质,同时保留它的身份特征?
00:00:59Then it becomes an artwork: can its placement and lighting make us feel that it is fragile?
接着,它成为艺术作品:能否通过摆放和灯光,让我们感到它的脆弱?
00:01:06Finally it becomes a character in time: can it disappear behind something and return as the same object?
最后,它成为时间中的角色:能否在被遮挡后重新出现,仍然是同一个物体?
00:01:12Those jobs demand different judgments.
这些任务需要不同的判断。
00:01:15A technically clean image may say very little.
技术上干净的图像,可能表达得很少。
00:01:18An expressive poster may contain an intentional impossibility.
富有表现力的海报,可能故意包含不可能的景象。
00:01:23A wonderful still may belong to a broken video.
一张精彩的静帧,可能来自一段存在问题的视频。
00:01:26As its job changes, you may find yourself approving a change that you would have rejected five minutes earlier.
随着任务改变,你可能会接受一个五分钟前还会拒绝的变化。
SLIDE 003 · 00:01:33 · Opening
The reflection is part of the edit
00:01:33For a few seconds, ignore the pear.
先暂时别看梨。
00:01:36Look at the water beneath it.
看看它下面的水。
00:01:38Now look back at the body.
现在再看回梨的主体。
00:01:40If the body becomes opaque ceramic, what happens to the light that used to pass through the glass?
如果主体变成不透明的陶瓷,原先穿过玻璃的光会怎样?
00:01:46The answer cannot live entirely inside the object's outline.
答案不可能完全局限在物体轮廓之内。
00:01:49This comparison is a prepared classroom illustration.
这组对比是事先准备的课堂示意图。
00:01:53Use it to locate the consequences of the request: transmission, highlights, and the reflection below.
用它来找出指令引发的影响:透光、高光,以及下方的倒影。
00:02:00We still want the blue stem and recognizable silhouette.
我们仍然希望保留蓝色的柄和可辨认的轮廓。
00:02:04But preserving every surrounding pixel could preserve the wrong optics. a successful local edit may require a carefully justified change somewhere else.
但保留周围的每一个像素,可能会保留错误的光学效果。成功的局部编辑,有时需要在别处做出有充分理由的改变。
SLIDE 004 · 00:02:14 · Opening
Necessary change or accidental drift?
00:02:14Let us make the opening comparison more disciplined.
让我们更严谨地审视开场的对比。
00:02:17The disappearance of transmitted amber light is consistent with changing transparent glass into opaque ceramic.
从透明玻璃变为不透明陶瓷后,透过的琥珀色光消失,是合理的。
00:02:25A changed blue stem, by contrast, would need a separate justification because the instruction did not request that transformation.
相比之下,如果蓝色的柄也变了,就需要另外解释,因为指令并没有要求这种变化。
00:02:33The reflection is more interesting.
倒影则更有意思。
00:02:35A material change can require its appearance to change, while its placement should still agree with the object and the water.
材质改变可能要求倒影外观随之变化,但它的位置仍应与物体和水面一致。
00:02:44We cannot classify every difference using a simple rule that change is bad.
我们不能用“变化就是坏事”这一简单规则来判断每个差异。
00:02:49We need a model of the intended scene.
我们需要理解目标场景应当如何运作。
00:02:51That is why the edit contract includes allowed consequences as well as invariants.
因此,编辑约定既要包含不变量,也要包含允许发生的连带变化。
SLIDE 005 · 00:02:57 · Opening
Tonight’s route
00:02:57Here is the route through that puzzle.
下面是我们破解这个难题的路线。
00:03:00We begin inside an image editor: its representations, sampling process, and controls.
我们先进入图像编辑器内部,研究它的表示方式、采样过程和控制手段。
00:03:09Then we take the art director's seat and decide what those controls should accomplish.
然后坐到艺术总监的位置,决定这些控制手段应该实现什么。
00:03:14In the final technical section, the image starts moving and the preservation problem becomes harder.
在最后一个技术部分,图像开始运动,保留原有信息的问题也变得更难。
00:03:21The spoken script is planned for about ninety minutes.
讲解稿按约九十分钟规划。
00:03:24The marked discussions, the break, and student presentations need additional class time.
标出的讨论、课间休息和学生展示,需要额外的课堂时间。
00:03:30You will see the same pear return in different situations.
你会看到同一只梨出现在不同情境中。
00:03:34Each return asks us to revise our judgment, rather than learn a completely unrelated example.
每次重逢,都要求我们修正判断,而不是学习一个毫不相关的新例子。
00:03:41At the end, three readings let us question who sets the goal and who decides whether it was achieved.
最后三篇阅读,将让我们追问:谁来设定目标,又由谁判断目标是否达成?
SLIDE 006 · 00:03:47 · Opening
Three decisions you can defend
00:03:47Here are three decisions I want you to be able to defend by the end.
到课程结束,我希望你能为这三类决定提供理由。
00:03:51When an edit fails, which control would you change, and why?
编辑失败时,你会调整哪一种控制手段,为什么?
00:03:56When two posters both look good, which belongs in this exhibition?
两张海报都很好看时,哪一张更适合这个展览?
00:04:01When a video looks convincing, which moment would you inspect before accepting it?
视频看起来可信时,你会在接受它之前检查哪个时刻?
00:04:07You do need a link between mechanism and consequence.
你确实需要把机制与结果联系起来。
00:04:10A mask tells us where; a reference can show us what; an artistic brief tells us why the result matters.
蒙版告诉我们改哪里;参考图可以展示改成什么;艺术创作说明告诉我们结果为何重要。
00:04:17Today we will practice making those links aloud.
今天,我们会练习把这些联系说清楚。
00:04:21The aim is to leave with reasons you can use when the next tool has a different interface.
目标是让你带走一套理由,即使下一款工具换了界面,也能继续使用。
SLIDE 007 · 00:04:26 · Part 1 / Image editing
How image editors work
00:04:26The pear gives us a concrete technical problem: change its material while retaining the information that makes it recognizable.
这只梨给我们一个具体的技术问题:改变材质,同时保留让人认出它的信息。
00:04:34We can describe this as conditional generation under preservation constraints.
我们可以把它描述为:在保留约束下进行条件生成。
00:04:39The inputs specify the change; the contract specifies what should survive.
输入指定要改变什么;编辑约定指定要保留什么。
00:04:44Throughout this section, connect each mechanism to one part of that problem.
在这一部分,请把每种机制对应到这个问题的某一环节。
00:04:49We begin by making our expectations explicit.
我们先把预期说清楚。
SLIDE 008 · 00:04:53 · Part 1 / Image editing
Before Part 1: image editing
00:04:53Before we open the machinery, decide what success should look like.
在打开技术黑箱之前,先确定什么才算成功。
00:04:57When glass becomes ceramic, what should stay the same?
玻璃变成陶瓷时,哪些东西应该保持不变?
00:05:00Which parts of the scene must change with the material?
场景的哪些部分必须随材质一起改变?
00:05:03What could a model misunderstand in that request?
模型可能会如何误解这个要求?
00:05:07Discuss these three questions and point to visible evidence for your choices.
请讨论这三个问题,并指出支持你们选择的可见证据。
00:05:11If your partner wants to preserve a detail that you want to change, identify the artistic intention behind each answer.
如果同伴想保留的细节恰好是你想改变的,请找出两种答案背后的艺术意图。
00:05:21We will return to these expectations when we write the worked ceramic-edit contract.
等到我们为陶瓷编辑写出完整约定时,再回头检视这些预期。
SLIDE 009 · 00:05:30 · Part 1 / Image editing
The edit contract
00:05:30The gallery has become a forest.
画廊变成了森林。
00:05:32Let us write the edit contract.
让我们写下编辑约定。
00:05:34The intended change is the environment.
我们想改变的是环境。
00:05:37The pear's identity, its pedestal, and the framing should remain recognizable.
梨的身份特征、底座和构图,都应该仍然可辨认。
00:05:43The ambient light and reflections are allowed to adapt to the forest.
环境光和反射则允许适应森林。
00:05:48Why include that last category?
为什么要加入最后这一类?
00:05:50Because preserving every visible relationship would contradict the new setting.
因为保留所有可见关系,会与新环境矛盾。
00:05:55Imagine retaining a bright gallery-window reflection inside a dark forest.
想象一下,在昏暗森林里还保留着明亮画廊窗户的反光。
00:05:59The model might preserve the source faithfully and still produce an implausible scene.
模型可能忠实保留了原图,却仍然生成了不合理的场景。
00:06:05Before generating, separate the requested change, the invariants, and the consequences that should adapt.
生成之前,要区分所要求的变化、不变量,以及应当随之调整的连带效果。
00:06:12Afterwards, inspect those categories separately.
生成之后,再分别检查这些类别。
SLIDE 010 · 00:06:16 · Part 1 / Image editing
An invariant can be semantic or exact
00:06:16Suppose a curator says, keep the pear exactly the same.
假设策展人说:让这只梨完全保持原样。
00:06:20There are at least two meanings hiding in that sentence.
这句话至少隐藏着两种含义。
00:06:24One means that a visitor should recognize the same sculpture in a new room.
一种是:观众在新房间里,仍然能认出这是同一件雕塑。
00:06:29Its highlights may change with the room.
它的高光可以随着房间而变化。
00:06:31The other means that selected pixel values must remain identical.
另一种是:选定像素的数值必须完全一致。
00:06:36Those are different requirements.
这是两种不同的要求。
00:06:38Generative conditioning can encourage recognizable identity.
生成式条件控制可以帮助保留可辨认的身份特征。
00:06:41Copying source pixels or compositing can enforce exact preservation in a specified region.
复制原图像素或进行合成,可以强制指定区域完全不变。
00:06:47Neither requirement automatically makes the whole image coherent: a copied region may now have the wrong lighting.
但这两种要求都不会自动保证整幅图像协调:复制过来的区域,光照可能已经不合适。
00:06:55Before choosing a tool, settle the meaning of same.
选择工具之前,先确定“相同”到底是什么意思。
SLIDE 011 · 00:06:59 · Part 1 / Image editing
Editing as conditional generation
00:06:59Read the expression as a distribution of possible edited results, given the source, instruction, and optional references.
把这个表达式理解为:给定原图、指令和可选参考信息后,可能产生的编辑结果的分布。
00:07:07The symbol theta represents the model's learned parameters.
符号 theta 表示模型学习到的参数。
00:07:11The vertical bar means 'given.' We are not asking for a unique answer in every case.
竖线表示“给定”。我们并不是在所有情况下都要求唯一答案。
00:07:17Two different forest scenes might both satisfy the same brief.
两个不同的森林场景,都可能满足同一份创作说明。
00:07:21However, a source image is a condition rather than a promise that every unmentioned pixel will be copied.
但原图只是条件,并不意味着每一个未提及的像素都会被复制。
00:07:27To judge success, we need a more precise contract than 'it looks plausible.'
要判断是否成功,我们需要比“看起来合理”更精确的约定。
SLIDE 012 · 00:07:33 · Part 1 / Image editing
A condition is not a hard constraint
00:07:33The probability expression says that the system produces a candidate given conditions.
概率表达式表示:系统在给定条件下生成一个候选结果。
00:07:38It does not contain a certificate that the candidate passes our checks.
它并不附带证明,保证这个结果通过我们的检查。
00:07:43The acceptance step is a separate decision that we impose on the result.
是否接受,是我们对结果做出的另一个独立判断。
00:07:47Imagine two forest outputs.
想象两个森林版本。
00:07:50Both look plausible, but one changes the pear's stem and one retains it.
两个都看起来合理,但一个改变了梨柄,另一个保留了它。
00:07:54The model can assign probability to both; our contract can reject one.
模型可以给两者都分配概率,而我们的约定可以拒绝其中一个。
00:08:00In practice, acceptance may involve human inspection, image comparisons, or explicit production constraints.
实际上,验收可能依赖人工检查、图像对比,或明确的制作约束。
SLIDE 013 · 00:08:07 · Part 1 / Image editing
A smaller workspace: image latents
00:08:07The RGB image has 1024 by 1024 pixels and three channels: just over three million scalar values.
RGB 图像有 1024 乘 1024 个像素,每个像素有三个通道,总共略多于三百万个标量值。
00:08:15With spatial reduction by sixteen and sixty-four latent channels, the representation becomes 64 by 64 by 64: 262,144 values.
空间尺寸缩小十六倍、潜在通道数为六十四时,表示变成 64 乘 64 乘 64,即 262,144 个数值。
00:08:31Sixteen times smaller along each spatial axis does not mean sixteen times fewer values overall, because the number of channels changes too.
每个空间轴缩小十六倍,并不意味着总数值量减少十六倍,因为通道数也变了。
00:08:39Here the scalar count falls by a factor of twelve.
在这个例子中,标量数量减少了十二倍。
00:08:43That is specific to this configuration.
这个比例只适用于这一具体配置。
00:08:45The model works in a learned representation, and its decoder must recover visible detail.
模型在学习得到的表示中工作,解码器必须从中恢复可见细节。
00:08:52Our next question is what that compression already loses before we request any edit.
接下来要问的是:在我们提出任何编辑要求之前,压缩已经丢失了什么?
SLIDE 014 · 00:08:57 · Part 1 / Image editing
Is the detail lost before the edit?
00:08:57Suppose the small exhibition caption is already blurry after an edit.
假设编辑后,展览的小字号说明文字变模糊了。
00:09:02You might spend an hour rewriting the prompt to say, preserve every letter.
你可能花一个小时重写提示词,强调要保留每一个字母。
00:09:08There is a quicker diagnostic: encode the original image and decode it again without requesting a change.
有一个更快的诊断方法:不要求任何改变,先把原图编码,再解码回来。
00:09:15If the letters are already damaged in that round trip, part of the problem lies in representation and reconstruction.
如果这次往返已经损坏了文字,那么问题的一部分就在表示和重建环节。
00:09:23A more emphatic instruction cannot restore a detail the editing path never retained faithfully.
再强硬的指令,也无法恢复编辑流程从未忠实保留的细节。
00:09:29Look at fine text, thin edges, and small textures first.
首先检查小字、细线边缘和微小纹理。
00:09:33If reconstruction is sound but the edit damages them, investigate the editing stage.
如果重建没有问题,而编辑损坏了它们,就调查编辑阶段。
00:09:39It also explains why exact exhibition typography deserves its own editable layer.
这也说明,要求精确的展览文字排版,应当保留在独立的可编辑图层中。
SLIDE 015 · 00:09:45 · Part 1 / Image editing
Why token count matters
00:09:45Now consider the number of tokens a transformer processes.
现在考虑 Transformer 需要处理多少个 token。
00:09:49In our toy example, one token per position on a 64 by 64 grid gives 4,096 tokens.
在这个简化例子里,64 乘 64 的网格上,每个位置对应一个 token,共有 4,096 个。
00:09:58Grouping each two-by-two region gives 1,024 tokens, a reduction by four.
把每个二乘二区域组合起来,就得到 1,024 个 token,数量减少四倍。
00:10:06Dense self-attention compares token pairs.
稠密自注意力会比较 token 对。
00:10:09Squaring the two counts gives 16,777,216 and 1,048,576 pair scores.
将两个数量分别平方,得到 16,777,216 和 1,048,576 个配对分数。
00:10:21That is a reduction by sixteen.
数量减少了十六倍。
00:10:24This is why representation choices matter so much for cost.
这就是为什么表示方式会如此显著地影响成本。
00:10:28It is not a claim that the whole system runs sixteen times faster.
但这不等于整个系统的运行速度会提高十六倍。
00:10:32Text tokens, reference tokens, other layers, and implementation details also matter.
文本 token、参考图 token、其他网络层和实现细节,也会影响结果。
SLIDE 016 · 00:10:38 · Part 1 / Image editing
Double the resolution: count the cost
00:10:38Before reading the second number, make a prediction.
看第二个数字之前,先预测一下。
00:10:42We double the image width and double its height, keeping the patch scheme fixed.
我们把图像宽度和高度都加倍,同时保持分块方式不变。
00:10:47How many image tokens do we get?
图像 token 会变成多少?
00:10:51In dense self-attention, each token can compare with every other token.
在稠密自注意力中,每个 token 都可以与其他所有 token 比较。
00:10:56Four times four gives sixteen times as many pair scores.
四乘四,配对分数的数量就变成十六倍。
00:10:59This calculation describes the image-token attention component, not a promise that the entire system becomes sixteen times slower.
这个计算只描述图像 token 的注意力部分,并不保证整个系统会慢十六倍。
00:11:08Other operations and implementations matter.
其他操作和实现方式同样重要。
00:11:12The useful habit is to ask what grew: pixels, tokens, comparisons, or measured runtime.
一个有用的习惯是追问:增加的究竟是像素、token、比较次数,还是实测运行时间?
00:11:19They are related quantities, but they are not interchangeable.
这些量相互关联,但不能互换。
SLIDE 017 · 00:11:23 · Part 1 / Image editing
Inside a current image editor
00:11:23Let us walk through this architecture from the input side.
让我们从输入端走一遍这个架构。
00:11:27The source image contributes semantic information through a vision-language component and visual information through a VAE.
原图通过视觉语言组件提供语义信息,并通过 VAE 提供视觉信息。
00:11:35The target stream begins with an evolving noisy representation.
目标分支从一个不断演化的带噪表示开始。
00:11:39The transformer processes that stream together with the conditions.
Transformer 将这个分支与条件信息一起处理。
00:11:44During training, we possess an example of the desired edited target, so we can measure what the model should learn.
训练时,我们拥有所需编辑结果的样本,因此可以衡量模型应该学习什么。
00:11:51During inference, we have the source and request but not the desired target.
推理时,我们有原图和要求,却没有所需的目标图像。
00:11:55The model must generate it.
模型必须把它生成出来。
00:11:57That difference is essential.
这个区别至关重要。
00:12:00The architecture does not quietly receive the answer when you ask it to edit your photograph.
当你让模型编辑照片时,架构并没有偷偷收到答案。
SLIDE 018 · 00:12:06 · Part 1 / Image editing
The missing target at inference
00:12:06There is a piece of information on the left that we do not possess on the right: the desired edited image.
左边有一项信息,是右边没有的:期望得到的编辑图像。
00:12:12During supervised training, a source, an instruction, and a target show the model what a successful change looks like.
在监督训练中,原图、指令和目标图一起,向模型展示成功的变化是什么样子。
00:12:20During use, we provide the request because that target does not yet exist.
使用时,我们之所以提出要求,正是因为目标图像还不存在。
00:12:26This distinction explains a surprisingly common confusion.
这个区别解释了一个相当常见的混淆。
00:12:29Showing a reference to an editor is not automatically teaching new weights.
给编辑器看一张参考图,并不自动意味着学习新的权重。
00:12:34It may simply be conditioning this particular generation.
它可能只是在为这一次生成提供条件。
00:12:38Adaptation, which we will reach with LoRA, changes trainable parameters.
适配则会改变可训练参数,讲到 LoRA 时我们会展开。
00:12:43It is available when we construct the learning signal, and absent when we ask the trained model to make our ceramic pear.
构造学习信号时,目标图像是已知的;要求训练好的模型生成陶瓷梨时,它则是未知的。
SLIDE 019 · 00:12:50 · Part 1 / Image editing
Two representations of the source
00:12:50Think of these two representations using our pear.
用这只梨来理解这两种表示。
00:12:54Semantic features help with questions such as: which object is the pear, what does ceramic mean, and what does 'it' refer to in the instruction?
语义特征帮助回答:哪个物体是梨,陶瓷是什么意思,以及指令中的“它”指什么?
00:13:05Visual latents provide appearance information, including shapes and textures that help keep the particular source recognizable.
视觉潜变量提供外观信息,包括形状和纹理,帮助保留原图的独特特征。
00:13:13This is a conceptual distinction, not a claim that the branches have perfectly isolated responsibilities.
这是一种概念上的区分,并不是说两个分支的职责完全隔离。
00:13:20Learned representations can overlap in what they encode.
学习到的表示所编码的信息可能重叠。
00:13:24If an edit fails, ask whether the model misunderstood the request or lost important visual information.
编辑失败时,要问:模型是误解了要求,还是丢失了重要的视觉信息?
SLIDE 020 · 00:13:31 · Part 1 / Image editing
OpenSubject: identity across new scenes
00:13:31How do you teach a system the difference between a pear and this particular pear?
如何教会系统区分“一只梨”和“这只特定的梨”?
00:13:36That question connects to OpenSubject, research I coauthored with Yexin Liu and our collaborators.
这个问题与 OpenSubject 有关,这是我与Yexin Liu及其他合作者共同完成的研究。
00:13:43A video offers something valuable: repeated observations of a subject as the view changes.
视频提供了一种宝贵资源:随着视角变化,对同一主体进行反复观察。
00:13:49The shared identity is a learning opportunity.
共同的身份特征提供了学习机会。
00:13:53Follow the main path in the figure.
请沿着图中的主路径看。
00:13:55We curate clips, verify subjects across frames, and select diverse pairs.
我们筛选视频片段,跨帧验证主体,并挑选具有多样性的图像对。
00:14:01Inpainting or outpainting helps synthesize reference inputs, followed by verification.
通过图像修补或扩图来合成参考输入,再进行验证。
00:14:08The corpus contains 2.5 million samples; the contribution is training data and a benchmark.
数据集包含二百五十万个样本;这项工作的贡献是训练数据和评测基准。
00:14:15For our exhibition, imagine the same sculpture photographed in several rooms.
对于我们的展览,可以想象在不同房间拍摄同一件雕塑。
00:14:20We want freedom to change the setting without losing its distinctive features.
我们希望自由改变环境,同时不丢失它的独特特征。
00:14:24That is a classroom application of the identity problem, rather than a claim that our fictional pear was tested in the paper.
这是身份保持问题的课堂应用,并不意味着论文测试过这只虚构的梨。
SLIDE 021 · 00:14:32 · Part 1 / Image editing
Flow matching: the training target
00:14:32The editor needs a way to learn how to move from noise toward an image.
编辑器需要学习如何从噪声走向图像。
00:14:36In this simple flow-matching setup, we construct an intermediate latent by mixing noise, epsilon, with a known target latent, z one.
在这个简单的流匹配设定中,我们将噪声 epsilon 与已知目标潜变量 z₁ 混合,构造中间潜变量。
00:14:47The mixing variable t tells us where we are along that training path.
混合变量 t 告诉我们处于这条训练路径的什么位置。
00:14:52For this straight interpolation, the target velocity is the target latent minus the noise.
对于这种直线插值,目标速度等于目标潜变量减去噪声。
00:14:57The model sees an intermediate state and learns to predict that direction, given the source and instruction.
模型看到中间状态,并在原图和指令的条件下,学习预测这个方向。
00:15:04There is also an important clock distinction: t here is the generative process's time.
还要区分两种时钟:这里的 t 是生成过程中的时间。
00:15:09Later, our video will have another time axis—the seconds that pass in the scene.
稍后的视频还有另一条时间轴:场景中实际流逝的秒数。
SLIDE 022 · 00:15:15 · Part 1 / Image editing
The interpolation endpoints
00:15:15Before trusting an equation, check its easiest cases.
相信一个方程之前,先检查最简单的情况。
00:15:19At t equals zero, the coefficient on the target becomes zero, so we recover the noise.
当 t 等于零时,目标的系数为零,所以我们得到噪声。
00:15:26At t equals one, the coefficient on noise vanishes, so we recover the target latent.
当 t 等于一时,噪声的系数消失,所以我们得到目标潜变量。
00:15:34Along this simple straight path, the difference between target and noise gives the direction we train the model to predict.
沿着这条简单直线路径,目标减去噪声,就是我们训练模型预测的方向。
00:15:41You do not need to imagine a recognizable half-finished image at every intermediate point.
不必想象每个中间点都对应一张能辨认的半成品图像。
00:15:47The calculation takes place in a learned representation, and this path is a teaching construction rather than the only possible design.
计算发生在学习得到的表示中;这条路径用于教学,并不是唯一可能的设计。
SLIDE 023 · 00:15:57 · Part 1 / Image editing
Predict the next sampling step
00:15:57Here is the smallest possible sampling calculation.
这是一个最小的采样计算例子。
00:16:01One coordinate currently equals 0.20.
当前某个坐标值是 0.20。
00:16:04Its predicted velocity is 0.60, and the next step has size 0.10.
预测速度为 0.60,下一步的步长为 0.10。
00:16:12Before looking at the result, say what operation should happen.
看结果之前,先说说应该做什么运算。
00:16:17We add a small amount of motion to the current state: 0.20 plus 0.10 times 0.60, giving 0.26.
在当前状态上加一小步:0.20 加上 0.10 乘 0.60,得到 0.26。
00:16:27In a real latent, many coordinates update together.
在真实的潜变量中,很多坐标会同时更新。
00:16:31These numbers are invented to show an Euler step, and practical samplers can use more elaborate solvers.
这些数字是为了演示欧拉步而编造的,实际采样器可以采用更复杂的求解器。
00:16:37But the sequence is now less mysterious: predict a direction, take a step, repeat, then decode.
但这个过程现在不再那么神秘:预测方向,迈出一步,重复,然后解码。
SLIDE 024 · 00:16:45 · Part 1 / Image editing
A better solver cannot clarify the brief
00:16:45Imagine following a beautifully accurate set of directions to the wrong gallery.
想象你沿着极其精确的路线指引,却走向了错误的画廊。
00:16:50Taking smaller steps will not repair the destination.
把步子迈小一点,并不能修正目的地。
00:16:54The same distinction helps with generative sampling.
同样的区分也有助于理解生成采样。
00:16:57A better numerical solver can follow a learned field more accurately.
更好的数值求解器,可以更准确地沿着学到的向量场前进。
00:17:02That alone does not settle an ambiguous editing request.
但仅凭这一点,无法解决含糊的编辑要求。
00:17:06If the pear's boundary is unstable across sampling settings, numerical or model behavior may matter.
如果不同采样设置下梨的边界不稳定,数值计算或模型行为可能有影响。
00:17:13If we never specified whether the glass reflection should become ceramic, we also have a brief problem.
如果我们从未说明玻璃倒影是否应变成陶瓷倒影,那么创作说明本身也有问题。
00:17:20Ask whether the system struggled to realize a clear instruction, or whether we have not agreed on what success means.
要问:系统难以实现一个清楚的指令,还是我们尚未就成功的含义达成一致?
00:17:27Those situations call for different next actions.
这两种情况需要不同的下一步行动。
SLIDE 025 · 00:17:31 · Part 1 / Image editing
Where edit supervision comes from
00:17:31Where does the ability to follow an edit instruction come from?
遵循编辑指令的能力从何而来?
00:17:35A training example can contain a source image, an instruction, and the corresponding edited target.
一个训练样本可以包含原图、指令,以及对应的编辑目标图像。
00:17:42InstructPix2Pix is an early example that used synthetic editing data to teach this relationship.
InstructPix2Pix 是一个较早的例子,它用合成编辑数据来学习这种关系。
00:17:48Suppose a training pair says 'make the object ceramic,' but its target also moves the camera and replaces the background.
假设某个训练对的指令是“把物体变成陶瓷”,但目标图还移动了相机并替换了背景。
00:17:56The learning signal no longer cleanly identifies the requested transformation.
这时,学习信号就不再能清晰地指向所要求的变换。
00:18:01The model may learn unwanted associations.
模型可能学到不希望出现的关联。
00:18:05This gives us an important data-quality question: did the target accomplish the requested change while preserving what should remain?
因此,一个重要的数据质量问题是:目标图是否完成了要求的变化,同时保留了应该保留的内容?
00:18:14Attractive targets alone are not enough to teach dependable editing.
仅有好看的目标图,还不足以教会可靠的编辑。
SLIDE 026 · 00:18:18 · Part 1 / Image editing
What a mismatched training pair teaches
00:18:18A training pair is a lesson, and lessons can accidentally teach the wrong thing.
一个训练对就是一堂课,而课程可能无意中教错东西。
00:18:23Imagine a source glass pear and a target ceramic pear that also has a different background.
想象原图是一只玻璃梨,目标图是一只陶瓷梨,但背景也变了。
00:18:29The written request says only to change the material.
文字要求却只说改变材质。
00:18:32Which visible changes should the model associate with that request?
模型应当把哪些可见变化与这个要求联系起来?
00:18:36If this mismatch recurs in the training data, the model can learn an unwanted association between a material change and a scene change.
如果训练数据中反复出现这种不匹配,模型可能把材质变化与场景变化错误地联系起来。
00:18:44Compare the instruction with the actual difference between source and target.
请将指令与原图、目标图之间的实际差异进行比较。
00:18:49Ask what supervision is rewarding, including changes nobody meant to label.
要问监督信号在奖励什么,包括那些没人打算标注的变化。
00:18:54The preservation contract begins in the examples used to teach the editor.
保留约定,应当从用来教会编辑器的样本开始。
SLIDE 027 · 00:18:59 · Part 1 / Image editing
Classifier-free guidance (CFG)
00:18:59Sometimes a conditioned prediction moves in the right direction but too weakly.
有时,条件预测的方向是对的,但力度不足。
00:19:04Classifier-free guidance uses the difference between a conditioned prediction and a baseline prediction to steer the result.
无分类器引导利用条件预测与基线预测之间的差异,来引导结果。
00:19:12Read the formula as baseline, plus a scaled change in direction.
可以把公式读成:基线,加上经过缩放的方向变化。
00:19:16When s equals one, the baseline terms cancel and we recover the conditioned prediction.
当 s 等于一时,基线项抵消,我们得到条件预测。
00:19:22Above one, we extrapolate beyond it.
大于一时,我们会外推到条件预测之外。
00:19:25That can strengthen the requested change, but it can strengthen errors as well.
这可以强化所要求的变化,但也可能强化错误。
00:19:30In an editing system, the baseline may still retain image information; baseline does not always mean no information at all.
在编辑系统中,基线可能仍保留图像信息;基线并不总意味着完全没有信息。
00:19:39Think of guidance as a specific operation on predictions, rather than a universal quality dial.
把引导理解为对预测值的一种具体运算,而不是万能的质量旋钮。
SLIDE 028 · 00:19:46 · Part 1 / Image editing
Guidance goes beyond the prediction
00:19:46The baseline predicts 0.2 and the conditioned version predicts 0.5.
基线预测是 0.2,条件预测是 0.5。
00:19:52With guidance scale two, do we get a value somewhere between them?
引导强度设为二,结果会落在两者之间吗?
00:19:58We take 0.2 plus twice the difference of 0.3, which gives 0.8.
我们用 0.2 加上两倍的差值 0.3,得到 0.8。
00:20:05We have moved beyond the conditioned prediction.
我们已经越过了条件预测。
00:20:08If the useful direction includes a slight mistake, we can amplify both.
如果有效方向中包含一点错误,那么两者都可能被放大。
00:20:13For the pear, the new material might become clearer while the stem or boundary becomes less faithful.
对于这只梨,新材质可能更明显,但柄或边界可能变得不那么忠实。
00:20:19Compare the requested change and preservation separately.
请分别比较所要求的变化和保留效果。
00:20:24One slider can move those judgments in opposite directions; a single overall impression can hide the tradeoff.
同一个滑块可能让两项评价朝相反方向变化;单一的总体印象会掩盖这种取舍。
SLIDE 029 · 00:20:31 · Part 1 / Image editing
Different controls carry different information
00:20:31Different controls communicate different kinds of information.
不同控制手段传达不同类型的信息。
00:20:35Text describes a requested change.
文本描述所要求的变化。
00:20:38A mask specifies a region.
蒙版指定区域。
00:20:41A spatial condition can describe pose, depth, or edges.
空间条件可以描述姿态、深度或边缘。
00:20:46An image reference can supply appearance information that would be difficult to express precisely in words.
参考图可以提供难以用语言精确表达的外观信息。
00:20:52These are not interchangeable knobs, and every product does not expose all of them.
这些控制并不能互相替代,也不是每款产品都提供全部选项。
00:20:57ControlNet and IP-Adapter are examples of distinct architectural approaches.
ControlNet 和 IP-Adapter 就是两种不同架构思路的例子。
00:21:01If exact untouched pixels are essential, a generation mask alone may be insufficient; explicit copying or compositing can enforce that requirement.
如果必须精确保留未编辑像素,仅有生成蒙版可能不够;明确复制像素或合成,可以强制实现这一要求。
00:21:12Choose a control by asking what information is missing, then check whether the resulting image actually respected it.
选择控制手段时,先问缺少什么信息,再检查结果是否真正遵循了它。
00:21:18A mask indicates a permitted region; it does not by itself solve seams or the reflection of a changed material.
蒙版指出允许编辑的区域,但它本身不能解决接缝,也不能处理材质改变后的倒影。
00:21:25The next slide names an architecture that learns to use a spatial condition.
下一页介绍一种学习使用空间条件的架构。
SLIDE 030 · 00:21:30 · Part 1 / Image editing
ControlNet: add spatial conditioning
00:21:30Suppose your sentence is understood, but the pear keeps changing shape.
假设模型理解了你的话,但梨的形状总在变化。
00:21:34You can describe its outline with more adjectives, or supply spatial evidence.
你可以用更多形容词描述轮廓,也可以直接提供空间证据。
00:21:39ControlNet gives an edge map, depth map, or pose a learned route into a compatible diffusion model.
ControlNet 为边缘图、深度图或姿态信息,提供了一条进入兼容扩散模型的可学习路径。
00:21:46Trace the two paths in this original architecture.
请追踪原始架构中的两条路径。
00:21:49The pretrained path stays frozen.
预训练主路径保持冻结。
00:21:51A trainable copy processes the added condition, and zero-initialized one-by-one convolutions connect its features to the main path.
一个可训练副本处理新增条件,零初始化的一乘一卷积把它的特征连接到主路径。
00:22:00Initially those connections contribute zero; training learns their contribution.
初始时,这些连接的贡献为零;训练会学出它们应有的贡献。
00:22:06The copied branch itself starts from pretrained weights, not all zeros.
副本分支本身从预训练权重开始,而不是全部从零开始。
00:22:11Now return to our request.
现在回到我们的要求。
00:22:13Edges can help specify the outline.
边缘可以帮助指定轮廓。
00:22:16They cannot, by themselves, tell us whether the reflected ceramic looks convincing or the blue stem retains its identity.
但仅凭边缘,无法判断陶瓷倒影是否可信,也无法保证蓝色梨柄保留身份特征。
SLIDE 031 · 00:22:24 · Part 1 / Image editing
Multiple references need distinct identities
00:22:24With several references, the system must know which reference contributes which information.
当有多张参考图时,系统必须知道每张图分别提供什么信息。
00:22:30Imagine asking for the sculpture from image one and the atmosphere from image two.
想象你要求使用图一的雕塑,以及图二的氛围。
00:22:37If the references become confused, you may get the wrong object with the right lighting.
如果参考图的角色混淆,可能会得到光照正确、物体却错误的结果。
00:22:43The illustrated research approach uses separators and image-index information to distinguish inputs.
图示研究方法使用分隔符和图像索引信息来区分输入。
00:22:48In an art brief, name the role of each image rather than presenting a pile of vaguely related inspiration.
在艺术创作说明中,要说明每张图的角色,而不是堆放一批关系模糊的灵感图。
00:22:55Then inspect for leakage: did a composition reference accidentally replace the subject?
然后检查信息是否串用:构图参考是否意外替换了主体?
00:23:02This figure describes one proposed mechanism, not a universal design used by every editor.
这张图描述的是一种提出的机制,并不是所有编辑器通用的设计。
SLIDE 032 · 00:23:08 · Part 1 / Image editing
A reference-role failure is visible in the output
00:23:08Look at these published multi-image examples with reference roles in mind.
请带着“参考图角色”的概念,观察这些已发表的多图示例。
00:23:13Before judging the output, identify what each input was supposed to contribute.
评价输出之前,先确定每个输入原本应该贡献什么。
00:23:18Then trace the subject and the requested transformation into the result.
然后在结果中追踪主体和所要求的变换。
00:23:23A reference-role failure can look superficially attractive.
参考图角色混淆的结果,表面上也可能很好看。
00:23:27The system may borrow the wrong object's appearance or import a background that was never intended.
系统可能借用了错误物体的外观,或引入了原本不需要的背景。
00:23:33We are examining qualitative examples from a particular paper, not conducting a broad comparison between products.
我们正在查看某篇论文的定性示例,并不是对产品做全面比较。
00:23:40The figure helps us practice an inspection method: follow the intended contribution of each input and look for unintended transfers between them.
这张图帮助我们练习一种检查方法:追踪每个输入的预期贡献,寻找意外的信息转移。
SLIDE 033 · 00:23:49 · Part 1 / Image editing
Attention as weighted information gathering
00:23:49Attention lets a token gather information from other tokens with different weights.
注意力让一个 token 以不同权重,从其他 token 收集信息。
00:23:55Our simplified scalar example gives weight 0.8 to a value of 0.9, and weight 0.2 to a value of 0.1.
在这个简化的标量例子中,数值 0.9 的权重为 0.8,数值 0.1 的权重为 0.2。
00:24:05The weighted result is 0.74.
加权结果是 0.74。
00:24:08The real mechanism operates on learned vectors, so this is arithmetic intuition rather than a literal description of artistic decision-making.
真实机制处理的是学习到的向量,因此这里只是在建立算术直觉,并非字面描述艺术决策。
00:24:17The important structure is selective combination: every available source need not contribute equally.
关键结构是选择性组合:不是每个可用来源都必须贡献同样多的信息。
00:24:23In a multi-reference edit, the system also needs to distinguish where information came from.
在多参考图编辑中,系统还需要区分信息来自哪里。
00:24:29Otherwise, gathering information successfully can still produce the wrong mixture of subject identity and visual style.
否则,即使成功收集了信息,也可能错误地混合主体身份与视觉风格。
SLIDE 034 · 00:24:36 · Part 1 / Image editing
Softmax weights sum to one
00:24:36The weights in this simplified attention example sum to one.
在这个简化注意力例子中,权重之和为一。
00:24:41That makes the result a weighted combination of the values.
因此,结果是各个数值的加权组合。
00:24:45A larger weight makes the associated value contribute more to this particular calculation.
权重越大,对应数值在这次计算中的贡献就越大。
00:24:51Now be cautious about jumping from the calculation to an explanation of the final image.
但要谨慎,不能直接从这个计算跳到对最终图像的解释。
00:24:57Real networks contain many layers, heads, and transformations.
真实网络包含许多层、注意力头和变换。
00:25:01A high weight at one point may not tell us why a visible feature ultimately appeared.
某处权重很高,未必能解释某个可见特征最终为何出现。
00:25:06For practical reference editing, the useful question remains whether the intended information survives in the output, regardless of how compelling an attention visualization looks.
在实际参考图编辑中,关键仍是预期信息是否出现在输出中,无论注意力可视化看起来多么有说服力。
SLIDE 035 · 00:25:18 · Part 1 / Image editing
LoRA: low-rank adaptation
00:25:18What if a useful concept must recur across many requests?
如果一个有用的概念需要在许多请求中反复出现,怎么办?
00:25:22A reference can condition a generation, while LoRA adapts selected learned weights using a low-rank update.
参考图可以为一次生成提供条件;LoRA 则通过低秩更新,调整选定的模型权重。
00:25:29The update is a product of two smaller matrices.
更新量是两个较小矩阵的乘积。
00:25:33Instead of learning every entry of a large update matrix independently, we express that update as a product of two smaller matrices.
我们不独立学习大型更新矩阵的每个元素,而是用两个较小矩阵的乘积来表示更新。
00:25:41For a 4096 by 4096 weight matrix and rank sixteen, the full matrix has 16,777,216 entries.
对于一个 4096 乘 4096、秩设为十六的例子,完整矩阵有 16,777,216 个元素。
00:25:52The two factors together have 131,072, a factor of 128 fewer for this update.
两个因子合计只有 131,072 个元素,这次更新的参数量减少了一百二十八倍。
00:26:01That is not a claim that the entire model shrinks by 128.
这并不意味着整个模型缩小了一百二十八倍。
00:26:05The base weights still exist.
基础权重仍然存在。
00:26:08Low rank makes adaptation economical, but does not certify that the learned subject survives a new pose.
低秩让适配更经济,但不能证明学到的主体在新姿态下仍能保持一致。
00:26:14That question requires examples the adaptation did not already see.
这个问题需要用适配时未见过的样例来检验。
SLIDE 036 · 00:26:19 · Part 1 / Image editing
Did it learn the subject or the setting?
00:26:19An adapted model reproduces your favorite training portrait perfectly.
适配后的模型完美复现了你最喜欢的训练肖像。
00:26:24Is that enough to use the character in a new film?
这就足以让这个角色出演新电影了吗?
00:26:27Consider what else it may have learned: the familiar camera angle, the background, even the lighting that always accompanied the subject.
想一想它还可能学到了什么:熟悉的机位、背景,甚至总与主体一起出现的光照。
00:26:35A held-out check deliberately changes those circumstances.
留出测试会有意改变这些条件。
00:26:39Ask for a new view or a different setting and inspect the distinctive features.
要求一个新视角或不同环境,并检查独特特征。
00:26:45We want a reusable concept, rather than a narrow ability to reproduce familiar combinations.
我们需要的是可复用的概念,而不只是复现熟悉组合的狭窄能力。
00:26:51For the pear, keep the blue stem recognizable while moving the exhibition outdoors.
对于这只梨,要在把展览移到户外时,仍保留可辨认的蓝色柄。
00:26:57If identity collapses there, more praise for the training examples will not solve the problem.
如果身份特征在那里崩溃,再称赞训练样例也解决不了问题。
SLIDE 037 · 00:27:03 · Part 1 / Image editing
What does the reward favor?
00:27:03A reward tells a learning process what kinds of outputs to favor.
奖励告诉学习过程应当偏好哪些输出。
00:27:07The figure shows a recent approach that separates task-specific reward training and then distills what is learned.
图中是一种近期方法:分别训练任务特定的奖励,再蒸馏学到的能力。
00:27:14Look at the categories: editing quality involves more than a single judgment of attractiveness.
看这些类别:编辑质量不只是单一的美观判断。
00:27:21Now imagine an artwork whose purpose is to feel awkward or disturbing.
现在想象一件艺术作品,其目的就是让人感到别扭或不安。
00:27:25A generic preference for polished images could work against that purpose.
对精致图像的普遍偏好,可能反而违背这个目的。
00:27:29Human preference signals are useful, but they do not define artistic merit for every project.
人类偏好信号很有用,但不能为所有项目定义艺术价值。
00:27:35This is the bridge to the next section: technical optimization can help produce candidates, while an artist still needs to decide which properties serve the work.
这正是通向下一部分的桥梁:技术优化可以帮助生成候选作品,而艺术家仍要决定哪些属性服务于作品。
SLIDE 038 · 00:27:46 · Part 1 / Image editing
One reward cannot stand in for every intention
00:27:46These edit examples invite a question about evaluation.
这些编辑示例引出了一个评价问题。
00:27:50Which properties would you reward separately?
你会分别奖励哪些属性?
00:27:53You might ask whether the instruction was carried out, whether identity survived, and whether the image is visually convincing.
可以问:指令是否完成、身份是否保留,以及图像在视觉上是否可信。
00:28:01Those judgments can disagree.
这些判断可能彼此冲突。
00:28:04A highly polished output might erase an awkward feature that is central to the artwork.
高度精致的输出,可能抹掉作品中至关重要的别扭特征。
00:28:09An unusual composition might serve the brief while attracting a lower generic preference score.
不寻常的构图可能符合创作要求,却得到较低的通用偏好分数。
00:28:15We should understand what a reward encourages before treating a high score as artistic approval.
在把高分视为艺术认可之前,我们应当理解奖励究竟鼓励什么。
00:28:20The published examples illustrate the research setting; our classroom task is to articulate the intention against which we would judge a particular result.
已发表示例展示的是研究情境;课堂任务则是明确创作意图,并据此判断具体结果。
SLIDE 039 · 00:28:30 · Part 1 / Image editing
Attractiveness and fidelity are separate tests
00:28:30Which failure would you notice first?
你会先注意到哪种失败?
00:28:33On one side, the output is beautiful, but it has quietly replaced our particular sculpture with a generic decorative pear.
一边的结果很美,却悄悄把我们的特定雕塑换成了普通的装饰梨。
00:28:40On the other, the right sculpture is present, but its reflection still behaves like the old material.
另一边保留了正确雕塑,但倒影仍像原来的材质。
00:28:46The first may win a quick aesthetic vote.
第一种可能赢得快速的审美投票。
00:28:49The second may pass a checklist that only asks whether the requested object changed.
第二种可能通过只检查“目标物体是否改变”的清单。
00:28:53Neither satisfies the full brief.
两者都没有满足完整要求。
00:28:56Has the editor lost identity, broken the scene's optical logic, or missed the work's intention?
编辑器是丢失了身份、破坏了场景光学逻辑,还是偏离了作品意图?
00:29:02Naming the mismatch gives us a next action.
说清不匹配之处,才能决定下一步。
00:29:05Saying only that an output looks a little wrong leaves the diagnosis unfinished.
只说输出“看起来有点不对”,诊断还没有完成。
SLIDE 040 · 00:29:11 · Part 1 / Image editing
A useful diagnosis changes the next action
00:29:11A diagnosis becomes useful when it changes the next action.
只有能改变下一步行动的诊断,才有用。
00:29:15If the wrong object changes, first suspect an ambiguous reference or region specification.
如果改错了物体,首先怀疑参考对象或区域描述有歧义。
00:29:22If the correct object becomes another instance, appearance grounding may be the issue.
如果改对了物体,却变成另一个个体,问题可能在外观信息的约束。
00:29:28If the material looks disconnected from its reflection, the dependent change may be missing.
如果材质与倒影脱节,可能缺少必要的连带变化。
00:29:34These are candidate explanations, not automatic conclusions.
这些只是候选解释,不是自动成立的结论。
00:29:38Choose a small revision that tests one of them.
选择一个小调整,检验其中一种解释。
00:29:40If it does not help, reconsider the hypothesis.
如果没有帮助,就重新考虑假设。
00:29:44This is more informative than changing the prompt, seed, reference, and model all at once, because then a successful result would leave us unsure which intervention mattered.
这比同时改提示词、随机种子、参考图和模型更有信息量;否则,即使成功,也不知道哪个改动起了作用。
SLIDE 041 · 00:29:54 · Part 1 / Image editing
Discuss: the ceramic edit
00:29:54We now have several ways to communicate an edit.
现在,我们有了几种表达编辑要求的方法。
00:29:58Would you preserve the reflection or let it change, and why?
你会保留倒影,还是允许它变化?为什么?
00:30:02When would a mask help more than a longer prompt?
什么时候蒙版比更长的提示词更有帮助?
00:30:06What would edges or depth control still leave uncertain?
边缘或深度控制仍会留下哪些不确定性?
00:30:10Discuss these three questions using the pear already on screen.
请用屏幕上的这只梨,讨论这三个问题。
00:30:14For each control you favor, name the failure it addresses and something it cannot settle.
对于你支持的每种控制手段,指出它解决什么失败,以及它不能决定什么。
00:30:21The next slide offers one possible contract; compare it with your reasoning rather than treating it as the only artistic answer.
下一页给出一种可能的约定;请与你的推理比较,不要把它当成唯一的艺术答案。
SLIDE 042 · 00:30:33 · Part 1 / Image editing
A worked ceramic-edit contract
00:30:33Here is one defensible answer to the ceramic puzzle.
对于陶瓷难题,这是一种有理有据的答案。
00:30:37Preserve the blue stem and recognizable silhouette.
保留蓝色柄和可辨认的轮廓。
00:30:40Allow transmission, highlights, and the reflection to change with the new material.
允许透光、高光和倒影随着新材质改变。
00:30:47Then inspect the boundary and water, because that is where the request reaches beyond the object.
然后检查边界和水面,因为要求的影响会在那里超出物体本身。
00:30:53Notice the wording: allow a justified consequence, rather than permit arbitrary drift.
注意措辞:允许有理由的连带变化,而不是允许任意漂移。
00:30:58We have not given the model permission to redesign the gallery.
我们没有授权模型重新设计画廊。
00:31:02We have named the changes necessary to make this material transformation coherent.
我们明确指出了,让这次材质变换协调一致所必需的变化。
00:31:07A mask could help localize work, while compositing might protect a region that truly must remain exact.
蒙版可以帮助限定工作区域;合成则可以保护真正必须完全不变的区域。
00:31:15The contract tells us how to use those tools.
编辑约定告诉我们如何使用这些工具。
SLIDE 043 · 00:31:18 · Part 1 / Image editing
The brief, the model, and the review
00:31:18We can now explain our opening puzzle at three levels.
现在,我们可以在三个层面解释开场难题。
00:31:22The brief decides what should change and what should survive.
创作说明决定什么应改变、什么应保留。
00:31:26The model uses representations, conditioning, and sampling to propose a result.
模型通过表示、条件控制和采样,提出候选结果。
00:31:31Our review checks whether that proposal actually satisfies the brief.
我们的审查检验这个候选结果是否真正满足要求。
00:31:36Keep those levels separate when you diagnose a failure.
诊断失败时,要区分这三个层面。
00:31:39Stronger guidance will not choose the exhibition's purpose.
更强的引导不会替你选择展览的目的。
00:31:43A more eloquent artistic statement will not enforce identical pixels.
更动人的艺术陈述不会强制像素完全一致。
00:31:48A beautiful preview will not prove preservation.
精美预览不会证明保留成功。
00:31:51The useful skill is connecting the right intervention to the observed problem.
有用的能力,是把正确的干预对应到观察到的问题。
00:31:56If your partner can tell when to use it and what it leaves uncertain, you have understood more than its name.
如果同伴能说清何时使用它、它还留下什么不确定性,你就不只是记住了它的名字。
SLIDE 044 · 00:32:03 · Part 1 / Image editing
After Part 1: explain the control
00:32:03We can now explain our controls rather than just name them.
现在,我们能够解释控制手段,而不只是说出名称。
00:32:06How does ControlNet differ from a mask or a text prompt?
ControlNet 与蒙版或文本提示词有什么区别?
00:32:10Why can stronger guidance make an edit worse?
为什么更强的引导可能让编辑变差?
00:32:14What evidence would show that an edit preserved identity?
什么证据能够表明编辑保留了身份特征?
00:32:18Discuss these three questions.
请讨论这三个问题。
00:32:20Choose a concrete requirement—a silhouette, an untouched region, or a distinctive stem—and connect your explanation to it.
选一个具体要求——轮廓、未编辑区域,或独特的梨柄——并把解释与它联系起来。
00:32:27Then consider whether the control guarantees that requirement or only helps express it.
然后考虑:这种控制是保证实现要求,还是仅仅帮助表达要求?
00:32:32Pause the video here before the break.
休息之前,请在这里暂停视频。
SLIDE 045 · 00:32:39 · Part 1 / Image editing
Break
00:32:39We will take ten minutes and resume when the class is ready.
我们休息十分钟,等大家准备好再继续。
00:32:43When you come back, think about a series of images you would recognize as belonging to one artwork.
回来时,请想一组你能认出属于同一件艺术作品的图像。
00:32:48What makes them belong together: the subject, the palette, the treatment of space, or something else?
是什么让它们属于一体:主体、色彩、空间处理,还是别的因素?
00:32:56Leave that question open for now.
暂时先保留这个问题。
00:32:59We will use it to move from individual editing operations to a coherent visual language.
我们将用它,从单次编辑操作转向连贯的视觉语言。
SLIDE 046 · 00:33:08 · Part 1 / Image editing
One exhibition, different visual worlds
00:33:08Look again at the artwork we have been using as a technical test.
再看看这件一直被我们用来做技术测试的作品。
00:33:12After the break, it has a different job.
休息之后,它有了不同的任务。
00:33:15We are no longer asking only whether the edit obeys a request.
我们不再只问编辑是否遵循了要求。
00:33:19We are asking what a visitor might feel, and which visual decisions create that feeling.
我们要问:观众可能感受到什么,哪些视觉决定创造了这种感受?
00:33:26Two images can have different surfaces and still belong to one exhibition.
两张图像的表面可以不同,却仍属于同一个展览。
00:33:30Two images can share a palette and feel unrelated.
两张图像可以使用相同色彩,却感觉毫不相关。
00:33:33The difference is worth arguing about.
这种差别值得讨论。
00:33:36In the next section, I will show prepared alternatives for AFTER RAIN.
下一部分,我会展示为《雨后》准备的不同方案。
00:33:41Choose a direction in your mind, and be ready to explain a visible reason.
请在心里选择一个方向,并准备说出可见的理由。
00:33:46Your neighbor may choose the other one.
你的邻座可能会选另一个。
SLIDE 047 · 00:33:49 · Part 2 / Art direction
Art direction
00:33:49Now take the art director’s seat.
现在,请坐到艺术总监的位置。
00:33:51The model can produce many plausible outputs; we decide which differences matter for AFTER RAIN.
模型能生成许多合理的结果;我们来决定哪些差异对《雨后》重要。
00:33:57As we compare the prepared images, name the visible relationship that carries the idea.
比较这些准备好的图像时,请指出承载创意的可见关系。
00:34:03That could be scale, light, or the treatment of space.
它可能是尺度、光线,或空间处理。
00:34:09Your explanation will be more useful than simply calling an image impressive.
这样的解释,比简单地说图像很惊艳更有用。
SLIDE 048 · 00:34:13 · Part 2 / Art direction
Before Part 2: artistic intention
00:34:13Before seeing the poster alternatives, consider what you want a visitor to experience.
看海报方案之前,先想想你希望观众经历什么。
00:34:18What makes an image feel fragile rather than merely attractive?
什么让图像显得脆弱,而不仅仅是好看?
00:34:22Can very different images belong to the same artwork, and why?
差别很大的图像,能否属于同一件作品?为什么?
00:34:27Which artistic decisions would you keep for yourself?
哪些艺术决定,你会留给自己?
00:34:31Discuss these three questions.
请讨论这三个问题。
00:34:33You might disagree about the effect of empty space or the meaning of a material.
你们可能对留白的效果或材质的含义有不同看法。
00:34:37Locate the visual evidence behind that disagreement.
请找出分歧背后的视觉证据。
00:34:43Keep your starting position in mind when we compare the prepared directions.
比较预先准备的方向时,请记住你的初始立场。
SLIDE 049 · 00:34:51 · Part 2 / Art direction
The exhibition brief
00:34:51Here is our commission: AFTER RAIN, a fictional exhibition about fragile objects in changing environments.
这是我们的委托:《雨后》,一个关于变化环境中脆弱物体的虚构展览。
00:35:00Imagine a visitor seeing its poster across a corridor before encountering a five-second moving image inside.
想象观众先在走廊另一端看到海报,再进入展厅看到一段五秒的动态图像。
00:35:07What should that visitor expect to feel?
你希望这位观众期待怎样的感受?
00:35:10The prepared examples keep the amber pear and blue stem recognizable.
准备好的示例都保留了可辨认的琥珀色梨和蓝色柄。
00:35:14We will examine photographic and collage directions, then consider how a moving reveal changes the experience.
我们将研究摄影和拼贴两个方向,再考虑运动中的揭示如何改变体验。
00:35:22Nobody needs to generate an image during this lecture.
这节课不需要任何人现场生成图像。
00:35:25Fragile is a useful beginning, but it is not yet an art direction.
“脆弱”是一个有用的起点,但还不算艺术方向。
00:35:30We need to translate that word into something visible enough to compare and precise enough to revise.
我们需要把这个词转化为足够可见、能够比较,又足够精确、能够修改的东西。
SLIDE 050 · 00:35:36 · Part 2 / Art direction
Making fragility visible
00:35:36Try replacing fragile with expensive in the brief.
试着把创作说明中的“脆弱”换成“昂贵”。
00:35:40You might still choose glass, but would you use the same composition?
你可能仍然选择玻璃,但还会用同样的构图吗?
00:35:44A large centered object and assertive lighting might suggest a luxury product.
居中放大的物体配合强势光照,可能让人联想到奢侈品。
00:35:49A small object surrounded by quiet space might instead seem exposed.
安静空间中一个小小的物体,则可能显得无所庇护。
00:35:54Scale can make the pear feel vulnerable.
尺度可以让梨显得易受伤害。
00:35:57Restrained light can make us look closely.
克制的光线可以引导我们细看。
00:35:59A reflection can suggest a world less stable than the object itself.
倒影可以暗示一个比物体本身更不稳定的世界。
00:36:04These are artistic hypotheses to test against an audience's reading, not a formula that makes every image fragile.
这些是需要通过观众解读来检验的艺术假设,不是让所有图像显得脆弱的公式。
00:36:12Point to the feature doing the work.
请指出真正起作用的特征。
00:36:15If nobody can locate it in the image, the intention may still be living only in the prompt.
如果没人能在图像中找到它,意图可能仍然只存在于提示词里。
SLIDE 051 · 00:36:21 · Part 2 / Art direction
Hokusai: scale, rhythm, and tension
00:36:21In Hokusai's Great Wave, look first for Mount Fuji.
看葛饰北斋的《神奈川冲浪里》,先找富士山。
00:36:24It is small and distant.
它很小,也很远。
00:36:26Now follow the curves of the wave and the boats beneath it.
现在沿着巨浪的曲线,以及下方的小船看。
00:36:30Scale and rhythm create a tension we can discuss without turning the artwork into a style label.
尺度和节奏创造了张力;我们可以讨论这种张力,而不是把作品变成一个风格标签。
00:36:36For AFTER RAIN, we might borrow a relationship: a small, vulnerable form facing a much larger environment.
对《雨后》来说,我们可以借鉴一种关系:一个小而脆弱的形体,面对远大于自身的环境。
00:36:43We do not need to reproduce the wave or ask for a generic imitation of the artist.
不需要复制巨浪,也不需要要求泛泛地模仿这位艺术家。
00:36:48The reference becomes useful when we can say which visual decision we are studying.
当我们能够说明正在研究哪个视觉决定时,参考作品才真正有用。
00:36:53Then we must test its translation.
然后,还必须检验这种转化。
00:36:55Our quiet flooded gallery has a different subject and emotional register.
我们安静而积水的画廊,有着不同的主体和情感基调。
SLIDE 052 · 00:37:00 · Part 2 / Art direction
Reference as a relationship, not a label
00:37:00Instead of using a reference as a label, describe a relationship you can observe.
不要把参考作品当作标签,而要描述你能观察到的关系。
00:37:06You might want a small stable form beneath a large dynamic curve, or a rhythm that moves the eye through the composition.
你可能想要一个巨大动态曲线下的小型稳定形体,或引导视线穿过构图的节奏。
00:37:14Then translate that relationship into the new subject and brief.
然后,把这种关系转化到新的主体和创作要求中。
00:37:18The result should be judged as its own artwork, not as a contest to resemble the reference.
应把结果作为独立作品评价,而不是比赛谁更像参考图。
00:37:23This approach makes references more useful to collaborators because they can understand what you are borrowing.
这种方法让参考资料对合作者更有用,因为他们能理解你借鉴的是什么。
00:37:29It also makes iteration more focused: if the intended tension is missing, you can revise scale or rhythm rather than vaguely asking for more influence from the source.
它也让迭代更聚焦:如果缺少预期张力,可以调整尺度或节奏,而不是含糊地要求“更受原作影响”。
SLIDE 053 · 00:37:41 · Part 2 / Art direction
Where does your eye arrive?
00:37:41Before admiring the surface detail, decide where your eye arrives.
欣赏表面细节之前,先判断你的视线最先落在哪里。
00:37:45Is it the pear, the reflection, or the space waiting above them?
是梨、倒影,还是它们上方留出的空间?
00:37:50A poster has to organize that first encounter.
海报必须组织这第一次相遇。
00:37:54If every region is equally busy, the audience has to invent a hierarchy we have not provided.
如果每个区域都同样繁忙,观众就得自行建立我们没有提供的层次。
00:38:00Here, the open area can make the object feel small and give the title somewhere to live.
这里的开放区域能让物体显得小,也给标题留下位置。
00:38:05We can ruin the sense of quiet by filling every available gap with decorative detail, even if each addition is attractive on its own.
如果用装饰细节填满每一处空隙,就可能破坏安静感,即使每个新增细节单独看都很漂亮。
00:38:13When revising this image, I would first ask whether the composition expresses the intended scale and silence.
修改这张图时,我会先问:构图是否表达了预期的尺度感与寂静?
00:38:20More texture comes later, if the work needs it.
如果作品需要,之后再增加纹理。
SLIDE 054 · 00:38:24 · Part 2 / Art direction
Negative space has a job
00:38:24Negative space does two jobs in this poster.
留白在这张海报中承担两项任务。
00:38:27It changes how large or isolated the object feels, and it creates a place for typography.
它改变物体显得多大、多孤立,也为文字排版创造位置。
00:38:34These are connected design decisions rather than separate finishing steps.
这是相互关联的设计决定,而不是彼此独立的收尾步骤。
00:38:39If you fill the pale area with dramatic detail, the image may become more visually active but leave no calm route for reading the title.
如果在浅色区域填满戏剧性细节,画面可能更活跃,却没有平静的路径让人阅读标题。
00:38:47If you reserve too much space, the subject might lose the presence you wanted.
如果留白过多,主体又可能失去你想要的存在感。
00:38:51Evaluate the space in relation to the finished use.
请结合最终用途来评价空间。
00:38:55A generative background is part of a larger composition when it must carry text, branding, or other exact elements.
当背景需要承载文字、品牌或其他精确元素时,生成背景只是更大构图的一部分。
SLIDE 055 · 00:39:03 · Part 2 / Art direction
A visual language has rules
00:39:03The collage direction changes the rules of the world.
拼贴方向改变了这个世界的规则。
00:39:06Torn edges replace smooth contours.
撕裂边缘取代了平滑轮廓。
00:39:09Flat layers replace optical depth.
平面层次取代了光学纵深。
00:39:12Amber and indigo keep a connection to our recurring subject.
琥珀色和靛蓝色,让它仍与反复出现的主体保持联系。
00:39:17Look at how those decisions affect the pear's apparent weight and vulnerability.
看看这些决定如何影响梨看起来的重量与脆弱感。
00:39:22If the body looks like paper but the water behaves like a photograph, do we accept that collision?
如果主体像纸,水却像照片,我们是否接受这种碰撞?
00:39:28We might, if it is deliberate and supports the work.
如果它是有意的,并服务于作品,我们可能接受。
00:39:31We might reject it as an unresolved mixture.
我们也可能认为它是尚未解决的混杂,因而拒绝。
00:39:35The answer depends on the intended visual language.
答案取决于预期的视觉语言。
00:39:38The next comparison asks you to locate precisely where that language holds together or breaks.
下一组比较,请你精确指出:这种语言在哪里成立,又在哪里断裂。
SLIDE 056 · 00:39:44 · Part 2 / Art direction
When a reflection breaks the collage
00:39:44A photographic reflection inside a paper collage can be a mistake, or the most interesting decision in the image.
纸质拼贴中的摄影式倒影,可能是错误,也可能是整张图最有意思的决定。
00:39:51We need to know whether the mismatch is an intentional disruption and whether it produces the intended effect.
我们需要知道:这种不匹配是否是有意的打破,以及是否产生了预期效果。
00:39:59Suppose the exhibition explores unstable memories.
假设展览探索的是不稳定的记忆。
00:40:02An impossibly photographic reflection could make sense.
一个不可能的摄影式倒影,就可能说得通。
00:40:06Suppose the brief calls for a coherent world assembled from torn paper.
假设要求是构建一个由撕纸组成的统一世界。
00:40:10The same reflection might weaken it.
同样的倒影可能削弱它。
00:40:13This does not mean every accident deserves an explanation after the fact.
这并不意味着每个意外都值得事后找理由。
00:40:18Make the intention specific, examine what the viewer can actually see, and compare alternatives.
请明确意图,检查观众实际能看见什么,并比较不同方案。
00:40:25A strong critique can distinguish a productive contradiction from an excuse for an unresolved result.
有力的评论能区分富有成效的矛盾,与为未完成结果寻找的借口。
SLIDE 057 · 00:40:32 · Part 2 / Art direction
Reference roles
00:40:32A reference becomes easier to use when we assign it a role.
为参考图分配角色后,它就更容易使用。
00:40:36One image might define the subject, another the composition, and another the material treatment.
一张图可以定义主体,另一张定义构图,再一张定义材质处理。
00:40:43Name the property you want from each.
请说清你想从每张图中得到什么属性。
00:40:45The model may not isolate those roles perfectly.
模型未必能完美隔离这些角色。
00:40:49That is why you should look for unwanted transfers, such as importing a reference's background when you only wanted its palette.
因此,要检查意外的信息转移,例如只想借用配色,却把参考图的背景也引入了。
00:40:56Try removing one reference and observing what disappears.
试着移除一张参考图,观察什么随之消失。
00:41:00Start with the smallest useful set.
从最小但有用的参考集合开始。
00:41:03If five references contradict one another, adding a sixth may make the problem harder to diagnose rather than solve it.
如果五张参考图互相矛盾,再加第六张可能让问题更难诊断,而不是解决问题。
SLIDE 058 · 00:41:11 · Part 2 / Art direction
What did the reference contribute?
00:41:11We have added a reference because we think it helps.
我们添加参考图,是因为认为它有帮助。
00:41:14How would we discover whether it actually does?
怎样才能知道它是否真的有帮助?
00:41:19First state its intended role: perhaps the pear's silhouette, perhaps the composition's sense of scale.
先说明它的预期作用:可能是梨的轮廓,也可能是构图的尺度感。
00:41:26Then look for that quality in outputs with and without the reference.
然后,在有无这张参考图的输出中寻找这种品质。
00:41:30Also inspect what came along uninvited.
也要检查哪些信息不请自来。
00:41:33A useful subject reference can carry a background or lighting scheme we did not want.
有用的主体参考,也可能携带我们不想要的背景或光照。
00:41:38This is an ablation: change one component to understand its contribution.
这就是消融实验:改变一个组件,以理解它的贡献。
00:41:44Random generation makes one lucky comparison weak evidence, so several samples can help.
随机生成会让一次幸运的对比缺乏说服力,因此多个样本会有帮助。
00:41:51If we cannot explain what a reference contributes, we may be making the request more complicated without making the direction clearer.
如果说不清一张参考图的贡献,我们可能只是把要求变复杂,却没有让方向更清楚。
SLIDE 059 · 00:41:59 · Part 2 / Art direction
An art-direction prompt
00:41:59Listen to how this prompt distributes responsibility.
听听这个提示词如何分配职责。
00:42:03The pear reference supplies silhouette and the blue stem.
梨的参考图提供轮廓和蓝色柄。
00:42:06Placement low on the right organizes the composition.
将主体放在右下方,用来组织构图。
00:42:10A pale upper-left area reserves space for the title.
左上方的浅色区域,为标题留出空间。
00:42:14Soft daylight and restrained reflections support the mood.
柔和日光与克制的倒影,支撑整体情绪。
00:42:17Every phrase points toward something we could inspect in an output.
每个短语,都指向输出中可以检查的东西。
00:42:22Compare that with asking for a stunning masterpiece with beautiful lighting.
与“用漂亮光照创作一幅惊艳杰作”这样的要求比较一下。
00:42:26This prompt is still a proposal, not a binding contract enforced by the model.
这个提示词仍然只是提议,不是模型强制执行的约束协议。
00:42:31If an output fails, we can now name the failed relationship: the title area is crowded, the light is too assertive, or the reference's background has leaked into the scene.
输出失败时,我们现在可以指出哪种关系失败了:标题区太拥挤、光线太强势,或参考图背景渗入了场景。
SLIDE 060 · 00:42:43 · Part 2 / Art direction
Group Editing: one decision, many views
00:42:43A single poster can look coherent by itself.
单张海报本身可能看起来很统一。
00:42:47A series exposes a harder problem: does the same material decision survive across viewpoints?
系列作品暴露出更难的问题:同样的材质决定,能否在不同视角下保持?
00:42:53Our Group Editing collaboration studies related images that should be edited consistently.
我们合作开展的 Group Editing 研究,关注应当保持一致编辑的关联图像。
00:42:58Treating each image independently can produce slightly different costumes or materials.
独立处理每张图,可能生成略有不同的服装或材质。
00:43:03The method arranges related images as pseudo-video frames to use a video model's consistency prior.
该方法把相关图像排列成伪视频帧,利用视频模型的一致性先验。
00:43:10VGGT provides geometric correspondences.
VGGT 提供几何对应关系。
00:43:14Geometry-enhanced rotary positional embeddings connect geometry features with image latents, while Identity-RoPE supports identity preservation.
几何增强的旋转位置编码,将几何特征与图像潜变量联系起来;Identity-RoPE 则支持身份保持。
00:43:24Follow the penguin examples through the figure before reading every label.
阅读每个标签之前,先沿着图中的企鹅示例看。
00:43:28Reliable correspondence is part of the technical problem: when views cannot be matched well, repeating the same instruction alone does not establish a coherent series.
可靠对应关系是技术问题的一部分:视角无法准确匹配时,仅重复相同指令,并不能保证系列统一。
SLIDE 061 · 00:43:40 · Part 2 / Art direction
Iteration as a controlled comparison
00:43:40Iteration becomes informative when we know what changed between attempts.
当我们知道两次尝试之间改了什么,迭代才有信息量。
00:43:44Suppose we compare two lighting treatments.
假设我们比较两种灯光处理。
00:43:47Keep the subject and composition instructions stable, and record the references and model version.
保持主体和构图指令不变,并记录参考图和模型版本。
00:43:54If a seed control exists, holding it fixed can help the first comparison.
如果可以控制随机种子,固定种子有助于初步比较。
00:43:59A shared seed still does not guarantee identical composition after a prompt change.
但即使种子相同,改变提示词后,也不能保证构图完全一致。
00:44:04If results vary widely, examine several outputs for each condition.
如果结果差异很大,就检查每种条件下的多个输出。
00:44:09Otherwise, one lucky sample may decide the whole direction.
否则,一个幸运样本就可能决定整个方向。
00:44:13The table is a proposed experiment, not measured results.
这张表是拟议实验,不是实测结果。
00:44:17Its purpose is to make each generation answer a question about the artwork.
它的目的,是让每次生成都回答一个关于作品的问题。
SLIDE 062 · 00:44:21 · Part 2 / Art direction
A lucky image or a reliable direction?
00:44:21Imagine you compare two prompts once, and the second gives a wonderful poster.
想象你只比较了两条提示词各一次,第二条生成了精彩海报。
00:44:28Did its wording cause the improvement, or did it receive a favorable random sample?
改进是措辞造成的,还是它碰巧抽到了有利的随机样本?
00:44:33From one output, it can be difficult to tell.
只看一个输出,往往很难判断。
00:44:36Several outputs per condition reveal whether the direction is stable or merely fortunate.
每种条件生成多个结果,才能看出这个方向是稳定有效,还是仅仅幸运。
00:44:42For art direction, we can ask two different questions.
在艺术指导中,我们可以问两个不同的问题。
00:44:45Would I exhibit this particular image?
我愿意展出这张具体图像吗?
00:44:48Could I reliably develop a series using this direction?
我能用这个方向,可靠地发展出一个系列吗?
00:44:52One exceptional output may answer the first while leaving the second unresolved.
一张出色的输出可能回答了第一个问题,却没有解决第二个。
00:44:57Repeating generations is helpful when it reveals a pattern relevant to the brief; it becomes a distraction when we keep browsing alternatives to avoid deciding what we value.
重复生成若能揭示与创作要求相关的规律,就有帮助;若只是为了逃避价值判断而不断浏览方案,就成了干扰。
SLIDE 063 · 00:45:08 · Part 2 / Art direction
Typography is part of the image
00:45:08Now the image becomes a poster.
现在,图像变成了海报。
00:45:11The title is an editable typographic layer, which lets us choose its wording, line breaks, and placement exactly.
标题是可编辑的文字图层,让我们精确选择措辞、换行和位置。
00:45:20That matters when a work has to carry an exhibition name rather than merely resemble a poster in a generated preview.
作品需要准确呈现展览名称时,这很重要,而不是只在生成预览中“看起来像海报”。
00:45:28Look at how the title occupies the area we deliberately left open.
看看标题如何占据我们有意留出的区域。
00:45:32The image and type were designed to cooperate.
图像与文字是协同设计的。
00:45:35We can adjust hierarchy without asking the model to regenerate the sculpture and risk changing it.
我们可以调整层次,而不必让模型重新生成雕塑、冒着改变它的风险。
00:45:41This is a useful division of labor: generation develops the visual material, and direct layout controls exact communication.
这是一种有用的分工:生成负责发展视觉素材,直接排版负责精确传达信息。
00:45:49The audience sees one composition.
观众看到的是一个整体构图。
00:45:52It does not need to know which parts came from which tool.
他们不需要知道哪个部分来自哪种工具。
SLIDE 064 · 00:45:56 · Part 2 / Art direction
Typography needs a reading order
00:45:56Typography establishes a sequence of attention.
文字排版建立了注意力的顺序。
00:45:59Ask what the viewer should read first, what comes next, and where their eye returns to the artwork.
要问:观众应先读什么、接着读什么,以及视线在哪里回到作品?
00:46:05In our poster, reserved space allows the title to be clear without covering the sculpture.
在我们的海报中,预留空间让标题清楚可读,又不会遮住雕塑。
00:46:10This is also a production decision.
这也是一个制作决定。
00:46:13Exact text can remain editable, while the image carries the material and atmosphere.
精确文字可以保持可编辑,而图像承载材质和氛围。
00:46:18If the title feels too dominant, change size, placement, or contrast and inspect the whole composition again.
如果标题太抢眼,就调整大小、位置或对比度,再检查整个构图。
00:46:25Do not judge the text in isolation from the image.
不要脱离图像,单独评价文字。
00:46:29The poster is the relationship between them, including the empty space that lets each element do its job.
海报是两者之间的关系,也包括让各元素发挥作用的空白。
SLIDE 065 · 00:46:35 · Part 2 / Art direction
Two directions for the same brief
00:46:35Take a few seconds with both directions before I describe them.
在我描述之前,先花几秒看看两个方向。
00:46:40Which would you put outside AFTER RAIN?
你会把哪一个放在《雨后》展厅外?
00:46:42Choose privately first, so your answer is not just a response to mine.
先自己选择,避免答案只是对我的回应。
00:46:47Then identify the visible feature that made you choose.
然后指出促使你选择的可见特征。
00:46:51The photographic direction can invite attention to material and atmosphere.
摄影方向可以引导观众关注材质与氛围。
00:46:56The collage direction can make fragility feel constructed through paper edges and layers.
拼贴方向可以通过纸边和层次,让脆弱感显得被建构出来。
00:47:01Both can serve the brief, but they make different promises to a visitor.
两者都可能符合要求,但向观众许下的承诺不同。
00:47:06A useful defense goes beyond realism versus abstraction.
有用的辩护,不应止于写实与抽象的区别。
00:47:10Tell us what the audience is likely to notice or feel, and how the composition produces that reading.
请告诉我们,观众可能注意或感受到什么,以及构图如何产生这种解读。
00:47:17We will use the disagreement to decide what to revise, rather than vote for a universally better image.
我们将用分歧来决定修改方向,而不是投票选出普遍更好的图像。
SLIDE 066 · 00:47:24 · Part 2 / Art direction
Why would you exhibit this one?
00:47:24Selection is an artistic decision.
选择本身就是艺术决定。
00:47:27Once generation gives us many plausible alternatives, choosing one determines what the work becomes.
当生成提供许多合理方案时,选中哪一个,就决定作品最终成为什么。
00:47:34The reason should connect to the brief: perhaps the small object feels more exposed, or the torn edge makes fragility physically legible.
理由应与创作要求相连:可能是小物体更显得无所庇护,或撕裂边缘让脆弱感变得可触可见。
00:47:43The rejected direction is useful evidence too.
被放弃的方向也是有用证据。
00:47:47Explain what it does well and why you are choosing something else for this exhibition.
请解释它哪里做得好,以及为什么为这个展览选择了别的方案。
00:47:52An unexpected model output can also change the direction, if you choose to develop it deliberately.
如果你决定有意识地发展一个意外输出,它也可以改变创作方向。
00:47:58The key question is whether you can now articulate the intention and carry it through subsequent decisions.
关键是:你现在能否说清意图,并在后续决定中贯彻它?
00:48:05Surprise can begin a work; it does not finish the artist's judgment.
惊喜可以开启作品,却不能替艺术家完成判断。
SLIDE 067 · 00:48:11 · Part 2 / Art direction
A critique with observable evidence
00:48:11A useful critique names a visible feature, explains its effect, and proposes a revision.
有用的评论会指出可见特征、解释效果,并提出修改建议。
00:48:16For example: the crisp reflection makes the collage feel photographic, so I would simplify its shape to match the flatter layers.
例如:清晰倒影让拼贴显得像摄影,所以我会简化它的形状,以配合更平面的层次。
00:48:24Compare that with 'I do not like the reflection.' The first comment gives the artist a relationship to inspect and a possible next move.
与“我不喜欢倒影”相比,前一种评论给艺术家提供了可检查的关系,以及可能的下一步。
00:48:33You can disagree about the intended effect, but make the disagreement specific.
你们可以对预期效果有不同看法,但要把分歧说具体。
00:48:39This is a classroom critique framework rather than an official grading rubric.
这是课堂评论框架,不是正式评分标准。
00:48:44Use it to connect evidence in the image to the idea the work is trying to communicate.
用它把图像中的证据,与作品试图传达的观念联系起来。
SLIDE 068 · 00:48:49 · Part 2 / Art direction
LightCtrl: lighting as an artistic choice
00:48:49A curator asks, can the same sculpture feel more vulnerable without changing its shape?
策展人问:能否不改变形状,却让同一件雕塑显得更脆弱?
00:48:55Lighting is one way to answer.
灯光是一种回答方式。
00:48:58Our LightCtrl collaboration studies controllable relighting from a single image, including light direction, intensity, and color temperature.
我们合作开展的 LightCtrl 研究,探索从单张图像进行可控重光照,包括光线方向、强度和色温。
00:49:07Follow the chair through the figure.
请沿着图中的椅子示例看。
00:49:09The latent proxy encoder extracts compact physical cues.
潜在代理编码器提取紧凑的物理线索。
00:49:13A lighting-aware mask guides the denoiser toward regions affected by the change, and preference optimization in the proxy branch supports physical consistency.
光照感知蒙版引导去噪器关注受变化影响的区域,代理分支中的偏好优化则支持物理一致性。
00:49:24The method connects an interpretable lighting request with image generation.
这种方法把可解释的光照要求与图像生成连接起来。
00:49:29For our pear, look for consequences rather than the word dramatic: where do highlights move, what happens to shadow, and how does the material read?
对于梨,不要只看“戏剧性”这个词,而要看后果:高光移向哪里、阴影如何变化、材质呈现如何?
00:49:40Inferring geometry and material from one photograph is ambiguous, so the result still needs inspection.
从单张照片推断几何和材质存在歧义,因此仍需检查结果。
00:49:47The exhibition example is our application of the idea, not an additional result reported by the paper.
展览示例是我们对这一思路的应用,不是论文另外报告的实验结果。
SLIDE 069 · 00:49:54 · Part 2 / Art direction
Discuss: two artistic directions
00:49:54Choose between the two prepared directions for AFTER RAIN.
请在为《雨后》准备的两个方向中做出选择。
00:49:58Which poster communicates fragility more clearly, and why?
哪张海报更清楚地传达脆弱感?为什么?
00:50:02What does the rejected direction reveal about your choice?
被放弃的方向,揭示了你选择中的什么?
00:50:06Which single revision would most change the audience’s reading?
哪一个单独的修改,会最大程度改变观众的解读?
00:50:10Discuss these three questions with a specific visual feature in view.
请看着一个具体视觉特征,讨论这三个问题。
00:50:15You can prefer quiet space or the instability of torn paper, but explain how that choice serves the exhibition.
你可以偏爱安静空间,也可以偏爱撕纸的不稳定感,但要解释这种选择如何服务于展览。
00:50:23Listen for a persuasive reason to choose the direction you initially rejected.
试着听取一个有说服力的理由,支持你最初拒绝的方向。
SLIDE 070 · 00:50:32 · Part 2 / Art direction
A process record makes critique more precise
00:50:32Consider what a process record lets us discuss.
想一想,过程记录能让我们讨论什么。
00:50:35An artist can record the intention and invariants, the prompt and reference roles, and one observation about a rejected result.
艺术家可以记录意图与不变量、提示词与参考图角色,以及对一个被拒绝结果的观察。
00:50:44That record explains what an attempt was testing.
这些记录说明了每次尝试在检验什么。
00:50:48Without that context, a folder of attractive images can be difficult to interpret.
缺少背景信息,一文件夹好看的图像也可能很难解读。
00:50:53We may not know which input changed or why a version was rejected.
我们可能不知道改了哪个输入,或为什么拒绝某个版本。
00:50:57A few precise notes make a comparison more informative.
几条精确笔记,就能让比较更有信息量。
00:51:01In a critique, this lets us ask about the artist’s decisions and the evidence behind them rather than guessing only from the final image.
在评论中,这让我们能够追问艺术家的决定及其证据,而不只是从最终图像猜测。
00:51:09It also helps distinguish an intentional departure from an accidental change.
它也有助于区分有意偏离与意外变化。
SLIDE 071 · 00:51:14 · Part 2 / Art direction
A gallery conversation
00:51:14Imagine we are standing between the two finished posters in a gallery review.
想象我们站在画廊评审现场,面前是两张完成的海报。
00:51:19I would begin with a visible decision that carries the idea, then locate an unintended change, then propose one revision worth discussing.
我会先指出一个承载观念的可见决定,再找出一个意外变化,最后提出一个值得讨论的修改。
00:51:29The order matters: the critique begins by understanding the work before prescribing a repair.
顺序很重要:评论应先理解作品,再给出修正方案。
00:51:35We can hear two contrasting readings of the same poster.
对同一张海报,我们可能听到两种相反的解读。
00:51:38One person may find the empty space quiet; another may find it emotionally distant.
一个人觉得留白安静,另一个人却觉得它情感疏离。
00:51:44Ask which details support each reading.
请问,哪些细节支持各自的解读?
00:51:46There is no image-making task here.
这里没有图像制作任务。
00:51:49We are practicing how to make feedback useful to the next decision.
我们是在练习如何让反馈有助于下一次决定。
00:51:54A good proposal names what would change and what effect we expect, so that a later version could confirm or challenge the reasoning.
好建议会说明改什么、预期产生什么效果,让后续版本能够证实或挑战这套推理。
00:52:02Discuss these three questions, then pause the video for the gallery conversation.
请讨论这三个问题,然后暂停视频,进行画廊对话。
SLIDE 072 · 00:52:11 · Part 2 / Art direction
After Part 2: defend a decision
00:52:11Before the artwork begins to move, defend a decision about the images.
在作品开始运动之前,请为一个图像决定提供理由。
00:52:17How can a technically correct image fail as an artwork?
技术上正确的图像,为什么可能在艺术上失败?
00:52:20Which visual rule should remain across a series?
一个系列中,哪条视觉规则应该保持?
00:52:24What evidence makes a critique useful for the next revision?
什么证据能让评论对下一轮修改有用?
00:52:28Discuss these three questions through one of the prepared posters.
请结合其中一张准备好的海报,讨论这三个问题。
00:52:31Make words such as coherent or expressive concrete: locate the feature, describe its effect, and explain what changing it would do.
把“统一”“有表现力”等词说具体:指出特征、描述效果,并解释改变它会怎样。
00:52:43We will carry those artistic rules into the video section.
我们会把这些艺术规则带入视频部分。
SLIDE 073 · 00:52:50 · Topic 3 / Video editing
Generative video editing
00:52:50Our image now has to survive time.
现在,图像必须经受时间的考验。
00:52:53The subject can move, disappear, and return.
主体会运动、消失,然后返回。
00:52:57We will examine selected frames from published research, follow the technical representations, and use a prepared AFTER RAIN scenario to decide what a convincing edit requires.
我们将查看已发表研究的精选帧、理解技术表示,并用准备好的《雨后》情境判断可信编辑需要什么。
00:53:09Keep the preservation contract, but add an event that could break it.
保留编辑约定,再加入一个可能让它失效的事件。
SLIDE 074 · 00:53:14 · Topic 3 / Video editing
Before Topic 3: images in motion
00:53:14A moving image creates new ways to break our contract.
动态图像会以新的方式破坏约定。
00:53:17What new failures become possible when an image moves?
图像运动起来后,会出现哪些新的失败?
00:53:21What should happen when the pear disappears behind a column?
梨消失在柱子后面时,应该发生什么?
00:53:25Can four convincing frames prove that a video works?
四张可信的帧,能证明整段视频成立吗?
00:53:28Discuss these three questions.
请讨论这三个问题。
00:53:31Separate a correct change in visibility from an unwanted change in identity.
区分正确的可见性变化,与不希望发生的身份变化。
00:53:36Name a moment you would need to see between the selected frames before accepting the shot.
接受这个镜头之前,请指出你需要在这些精选帧之间看到的某个时刻。
00:53:43Your prediction will give us a concrete test for the methods that follow.
你的预测会为后续方法提供一个具体测试。
SLIDE 075 · 00:53:51 · Topic 3 / Video editing
A video edit must survive the next frame
00:53:51A still image lets us choose a flattering instant.
静态图像让我们挑选一个最好看的瞬间。
00:53:54A video makes the object keep its promises in the next frame.
视频则要求物体在下一帧继续信守承诺。
00:53:58Look at the source sequence.
看原始序列。
00:54:00As viewpoint and visibility change, we continue to recognize the subject.
随着视角和可见性变化,我们仍能认出主体。
00:54:06An edit has to preserve that relationship while introducing its requested transformation.
编辑必须在引入所要求变换的同时,保留这种关系。
00:54:14The sampling steps describe how a model generates a result.
采样步描述模型如何生成结果。
00:54:17The frames here describe time passing in the depicted scene.
这里的帧描述被描绘场景中时间的流逝。
00:54:22A system may use many sampling steps to produce a short clip, but those steps are not extra seconds of action.
系统可能用很多采样步生成短片,但这些步数不是额外的剧情秒数。
00:54:29Now the preservation contract must hold across an event, not just inside a frame.
现在,保留约定必须跨越一个事件成立,而不只是局限于单帧内部。
SLIDE 076 · 00:54:35 · Topic 3 / Video editing
Consistency is not stillness
00:54:35Would a video be perfectly consistent if every frame were identical?
如果每帧完全相同,视频就算完美一致吗?
00:54:39Only if the intended scene were perfectly still.
只有目标场景本来就完全静止时才是。
00:54:43In our moving shot, pose, viewpoint, illumination, and visibility should change.
在运动镜头中,姿态、视角、光照和可见性都应变化。
00:54:52Consistency means that those changes remain coherent with the scene.
一致性意味着这些变化与场景保持协调。
00:54:56The blue stem may become hidden during a turn.
蓝色柄可能在转动时被遮住。
00:55:00That is different from its color changing without a lighting explanation.
这与没有光照原因却突然变色,是不同的。
00:55:04A texture should move with its surface, rather than crawl independently across it.
纹理应该随表面一起移动,而不是独自在表面爬动。
00:55:09If there is no convincing explanation, you may have found instability.
如果找不到可信解释,你可能发现了不稳定性。
00:55:14This is why the goal cannot simply be to minimize change from one frame to the next.
因此,目标不能只是尽量减少相邻帧之间的变化。
SLIDE 077 · 00:55:20 · Topic 3 / Video editing
A moving image can reveal an idea
00:55:20Here is a prepared storyboard, not a generated video result.
这是准备好的分镜,不是生成视频的结果。
00:55:24In the first view, the reflection suggests an object we have not fully seen.
第一个视角中,倒影暗示了一个我们尚未看全的物体。
00:55:29A partial view then gives us enough evidence to make a guess.
接着,局部视角提供足够线索,让我们猜测。
00:55:33The wide view finally changes our understanding of the setting.
最后的全景改变了我们对环境的理解。
00:55:36The audience is doing something during those five seconds: forming an expectation, testing it, and revising it.
在这五秒中,观众一直在行动:建立预期、检验预期,然后修正它。
00:55:44Compare this sequence with revealing everything in the first frame.
把这个序列与第一帧就揭示一切的版本比较一下。
00:55:48The same sculpture could be present, but the experience would differ.
同一件雕塑可以都在场,但体验会不同。
00:55:52For AFTER RAIN, the timing of information can carry fragility as strongly as material or lighting.
对于《雨后》,信息揭示的时机,可以像材质或灯光一样有力地承载脆弱感。
00:55:58We should decide that structure before asking a tool to fill in motion.
要求工具补充运动之前,我们应该先决定这种结构。
SLIDE 078 · 00:56:03 · Topic 3 / Video editing
Five seconds can have a structure
00:56:03Five seconds is short, but it can still have a structure.
五秒虽然短,却仍然可以有结构。
00:56:06Our proposed first interval shows a reflection, giving the viewer clues.
我们设想的第一段展示倒影,给观众线索。
00:56:11The second offers a partial view.
第二段提供局部视角。
00:56:13The final interval reveals the wider setting and changes how the object is understood.
最后一段揭示更广阔的环境,改变人们对物体的理解。
00:56:18Try another allocation and predict the effect.
试着重新分配时间,并预测效果。
00:56:21A longer reflection might create uncertainty; a quick reveal might make the piece feel like a product shot.
更长的倒影段落可能制造不确定感;快速揭示则可能让它像产品广告镜头。
00:56:28These timings are artistic proposals, not measured outputs from a video model.
这些时间安排是艺术提案,不是视频模型的实测输出。
00:56:34By deciding the experience first, we can evaluate whether the generated motion and cuts support it.
先决定体验,才能评价生成的运动和剪辑是否支持它。
00:56:40A technically smooth sequence can still have the wrong rhythm.
技术上流畅的序列,节奏仍可能不对。
SLIDE 079 · 00:56:44 · Topic 3 / Video editing
An image editor can read a contact sheet
00:56:44The research begins with a surprisingly simple experiment: arrange video frames into a contact sheet and give that image to an image editor.
这项研究始于一个出人意料的简单实验:把视频帧排成接触印样,再交给图像编辑器。
00:56:53This asks whether an existing image-editing capability can transfer across multiple views presented together.
它想检验:现有图像编辑能力,能否迁移到一起呈现的多个视角上?
00:57:01Treat the result as evidence that motivates a research direction.
请把结果当作启发研究方向的证据。
00:57:05A grid can make frames available in a shared context, but a successful-looking selection does not prove reliable video editing.
网格让多帧共享上下文,但一组看似成功的精选帧,并不能证明视频编辑可靠。
00:57:12There may be flicker between the sampled frames.
采样帧之间可能仍有闪烁。
00:57:16The next steps investigate how to represent video more effectively and adapt the model, rather than assuming a contact sheet alone solves time.
后续研究探索更有效的视频表示和模型适配,而不是假设接触印样本身就解决了时间问题。
SLIDE 080 · 00:57:24 · Topic 3 / Video editing
Four good frames can hide a bad video
00:57:24But the missing intervals are where some of the most revealing failures live.
恰恰是在缺失的间隔中,可能藏着最有揭示性的失败。
00:57:29A feature can jump, disappear, and return between the frames on this page.
某个特征可能在这一页的两帧之间跳动、消失,再返回。
00:57:33Use the contact sheet for what it does well: compare appearance across selected moments and locate regions to inspect.
发挥接触印样的长处:比较选定时刻的外观,并定位需要检查的区域。
00:57:41Then review full playback, with closer attention around turns and occlusion.
然后完整播放,特别关注转向和遮挡附近。
00:57:45Think of a film review based only on publicity stills.
想象只凭宣传剧照来评论电影。
00:57:49You might judge the costume and lighting, but you have not yet seen the performance.
你或许能评价服装和灯光,却还没有看过表演。
00:57:55In generative video, that missing performance includes whether the same object continues to exist convincingly from moment to moment.
对生成视频而言,缺失的“表演”包括:同一个物体是否在每个时刻都可信地持续存在。
SLIDE 081 · 00:58:04 · Topic 3 / Video editing
The grid can live in latent space
00:58:04The grid can also be constructed in latent space.
网格也可以在潜在空间中构造。
00:58:07Instead of only combining visible pixel images, the method encodes frames and arranges their tokens into a virtual image grid.
这种方法不只是组合可见像素图像,而是先编码各帧,再把 token 排列成虚拟图像网格。
00:58:16The published example changes a bus into a graphics card.
论文中的示例把一辆公交车变成显卡。
00:58:20The key idea is shared treatment of positions within that constructed representation.
核心思路是,在这种构造的表示中统一处理各位置。
00:58:25Do not confuse this step with using a video VAE; they are different design decisions.
不要把这一步与使用视频 VAE 混为一谈;它们是不同的设计决定。
00:58:31Putting frames near one another in a representation can help the model exchange information, but it does not impose a guarantee of physical continuity.
在表示中把帧放在一起,有助于模型交换信息,但不会强制保证物理连续性。
00:58:40We still need to inspect what survives across changing views.
我们仍需检查视角变化时,哪些信息得以保留。
SLIDE 082 · 00:58:45 · Topic 3 / Video editing
A virtual grid is a representation choice
00:58:45A virtual grid provides a way to arrange information for a model trained around image-like structure.
虚拟网格为围绕图像结构训练的模型,提供了一种信息排列方式。
00:58:51It creates shared context and positional relationships among frame representations.
它在各帧表示之间建立共享上下文和位置关系。
00:58:56That can be a useful bridge when reusing an image editor.
复用图像编辑器时,这可以成为有用的桥梁。
00:59:00But arrangement alone does not enforce the laws of motion or the persistence of a hidden object.
但排列方式本身,不能强制执行运动规律或被遮挡物体的持续存在。
00:59:05Those abilities depend on the model, adaptation, data, and other parts of the workflow.
这些能力取决于模型、适配、数据,以及工作流程中的其他部分。
00:59:12Separate the representation choice from the capability claim.
请区分表示选择与能力主张。
00:59:16A diagram can explain how frames enter a system without proving that the system handles every difficult event.
图表可以解释帧如何进入系统,却不能证明系统处理得了所有困难事件。
00:59:23That proof would require appropriate outputs and evaluation.
证明这一点,需要适当的输出和评价。
SLIDE 083 · 00:59:28 · Topic 3 / Video editing
Repurposing an image model for video
00:59:28This architecture asks whether an image editor's abilities can be reused for video.
这个架构检验图像编辑器的能力能否复用于视频。
00:59:33The video VAE encodes a clip into a compressed latent representation.
视频 VAE 将片段编码成压缩的潜在表示。
00:59:39Learned projections connect those video latents with the image-editing transformer, so the model can operate on a compatible arrangement of information.
学习得到的投影把视频潜变量与图像编辑 Transformer 连接起来,让模型处理兼容的信息排列。
00:59:48Trace that main route before reading the smaller branches.
阅读较小分支之前,先沿着这条主路径看。
00:59:51Adaptation connects them; the paper explores LoRA and full training choices, with an optional enhancement stage.
适配把它们连接起来;论文探索了 LoRA 和全量训练,并包含可选的增强阶段。
00:59:59The attractive idea is reuse of editing knowledge.
吸引人的想法是复用编辑知识。
01:00:03The test is whether that reuse preserves the temporal relationships our shot needs.
真正的检验是:这种复用是否保留了镜头需要的时间关系?
01:00:08Architecture explains where information flows; the edited clip shows whether the intended event survives.
架构解释信息流向哪里;编辑后的片段展示预期事件是否得以保留。
SLIDE 084 · 01:00:15 · Topic 3 / Video editing
Adaptation connects incompatible representations
01:00:15The image model and video VAE do not necessarily speak the same representational language.
图像模型与视频 VAE 未必使用相同的表示语言。
01:00:22Learned projections help connect them.
学习得到的投影帮助连接二者。
01:00:25Adaptation then allows the editing transformer to operate usefully with the video representation.
之后,适配让编辑 Transformer 能有效使用视频表示。
01:00:31This is a common engineering pattern: reuse a capable component while learning the interface and behavior needed for another task.
这是一种常见工程模式:复用有能力的组件,同时学习另一项任务所需的接口和行为。
01:00:39The benefit is not automatic.
收益并不会自动出现。
01:00:41A projection must preserve useful information, and the adapted model must learn how the new structure relates to edits.
投影必须保留有用信息,适配后的模型也必须学会新结构如何与编辑相关联。
01:00:49In the published pipeline, decoding returns the edited representation to video.
在已发表的流程中,解码将编辑后的表示还原为视频。
01:00:56Follow the information through every stage rather than treating the name of the reused model as an explanation by itself.
请追踪信息经过每个阶段,而不要把被复用模型的名字本身当作解释。
SLIDE 085 · 01:01:04 · Topic 3 / Video editing
45 frames become 12: why?
01:01:04Here is a counting puzzle with a small trap.
这里有一个带小陷阱的计数题。
01:01:06In this example, forty-five pixel frames are compressed with a special first frame and a factor of four for the remaining temporal groups.
在这个例子中,四十五个像素帧采用特殊首帧处理,其余时间组按四倍压缩。
01:01:15Forty-five is four times eleven, plus one.
四十五等于四乘十一,再加一。
01:01:19The latent sequence therefore has eleven plus one, or twelve frames.
因此,潜在序列有十一加一,即十二帧。
01:01:24Those twelve latent frames can be arranged in the illustrated three-by-four virtual grid.
这十二个潜在帧可以排成图中的三乘四虚拟网格。
01:01:29Simply dividing forty-five by four would miss the first-frame convention.
简单地用四十五除以四,会忽略首帧约定。
01:01:34This arithmetic belongs to the representation used in this example; it is not a rule for every video model.
这个算法属于本例采用的表示,并不是所有视频模型的通用规则。
01:01:41A small detail in temporal compression changes what the editing model actually receives.
时间压缩中的一个小细节,会改变编辑模型实际收到什么。
SLIDE 086 · 01:01:47 · Topic 3 / Video editing
Temporal compression preserves a special first frame
01:01:47Solve the frame-count relationship step by step.
一步步解出帧数关系。
01:01:50Start with four k plus one equals forty-five.
从 4k 加一等于四十五开始。
01:01:54Subtract one to obtain forty-four, then divide by four to get k equals eleven.
减一得到四十四,再除以四,得到 k 等于十一。
01:02:00The latent sequence contains k plus one frames, so its length is twelve.
潜在序列包含 k 加一帧,因此长度为十二。
01:02:06This arithmetic encodes the special handling of the first frame in the stated convention.
这个计算反映了该约定对首帧的特殊处理。
01:02:11It is different from simply dividing the total number of pixel frames by four.
它不同于直接把像素帧总数除以四。
01:02:17Understanding them can prevent mistakes when arranging latent frames into a virtual grid or comparing the representation with the original clip.
理解这些关系,可以避免在排列潜在帧网格、或与原始片段比较时出错。
SLIDE 087 · 01:02:26 · Topic 3 / Video editing
Guidance in a published video ablation
01:02:26The requested edit turns the scene into a cyberpunk workshop with holographic documents.
编辑要求是把场景变成赛博朋克工作室,并加入全息文档。
01:02:32Compare the results before focusing on the setting labels.
先比较结果,再关注设置标签。
01:02:36Which version carries out the requested transformation more completely, and what visible evidence supports your answer?
哪个版本更完整地实现了要求的变换?有什么可见证据?
01:02:43The paper presents this case as an example where classifier-free guidance improves edit completeness.
论文用这个案例说明,无分类器引导可以改善编辑完整性。
01:02:50Its baseline retains source latents, and its implementation includes rescaling.
其基线保留了原始潜变量,实现中还包含重缩放。
01:02:55This does not establish one universally best guidance value.
这并不意味着存在一个普遍最优的引导值。
01:02:59Guidance also has a computation cost when it requires another model pass.
如果引导需要额外一次模型前向计算,也会带来计算成本。
01:03:05Relate the example back to our equation: changing a prediction combination affects how strongly the result follows the condition.
联系之前的方程:改变预测组合,会影响结果遵循条件的力度。
SLIDE 088 · 01:03:13 · Topic 3 / Video editing
What this comparison actually shows
01:03:13An ablation asks what changes when one component is removed.
消融实验追问:移除某个组件后,会发生什么变化?
01:03:17In the displayed example, guidance makes more of the requested transformation visible.
在展示的例子中,引导让更多要求的变换可见。
01:03:22That is useful evidence about this comparison.
这是关于这次比较的有用证据。
01:03:26It is not yet a measurement of reliability across the kinds of shots we might produce for an exhibition.
但它还不是对我们可能用于展览的各类镜头的可靠性测量。
01:03:32As a reader, separate the observation from the next question.
作为读者,请区分观察结果与下一步问题。
01:03:35We can observe a more complete edit here.
我们可以观察到,这里的编辑更加完整。
01:03:38We would still want to know what happens across other clips, how often preservation suffers, and what the computation costs.
但仍想知道:其他片段会怎样、保留效果多常受损,以及计算成本是多少。
01:03:46You can learn a mechanism from a selected example while designing a stronger test for the decision you actually need to make.
你可以从精选示例中学习机制,同时为真正需要做出的决定设计更强的测试。
SLIDE 089 · 01:03:55 · Topic 3 / Video editing
Local editing across frames
01:03:55This is a local edit: the sheep's face changes and receives a white star-shaped patch.
这是一次局部编辑:绵羊的脸发生变化,增加了一块白色星形斑纹。
01:04:00Look at the patch across poses.
观察不同姿态下的这块斑纹。
01:04:03Does it remain attached to the same facial region, with a plausible change in apparent shape as the head turns?
它是否始终附着在同一面部区域,并在头部转动时合理改变表观形状?
01:04:10The most attractive frame is not enough.
仅看最漂亮的一帧是不够的。
01:04:12An identity marker can slide, disappear, or change shape at another moment.
身份标记可能在另一个时刻滑动、消失或变形。
01:04:17Selected frames let us ask the right questions, but the full clip is needed to check continuity between them.
精选帧让我们提出正确问题,但要检查帧间连续性,仍需完整视频。
01:04:25For our discussion, identify a small recognizable feature that will make identity drift easier to notice.
请为讨论找一个可辨认的小特征,让身份漂移更容易被发现。
SLIDE 090 · 01:04:33 · Topic 3 / Video editing
Follow the identity marker
01:04:33Choose a distinctive feature and follow it through the shot: its attachment to the subject, its shape, and its reappearance after partial visibility.
选一个独特特征,并在镜头中追踪它:与主体的附着关系、形状,以及部分遮挡后的再次出现。
01:04:43For our pear, the blue stem is a useful witness.
对这只梨而言,蓝色柄是有用的见证者。
01:04:47It may become hidden as the camera moves; we should allow that.
随着相机移动,它可能被遮住;这是应当允许的。
01:04:51When it returns, it should still belong to the same object.
当它返回时,仍应属于同一个物体。
01:04:54A marker that slides across the surface or reappears in a new shape tells us something a general impression of smooth motion might miss.
如果标记在表面滑动,或以新形状返回,就揭示了整体流畅感可能掩盖的问题。
01:05:02We are using the marker as a diagnostic aid, while still judging the whole object's identity and the shot's intended motion.
我们用标记辅助诊断,同时仍评价整个物体的身份和镜头预期运动。
SLIDE 091 · 01:05:11 · Topic 3 / Video editing
Global stylization across frames
01:05:11Here the instruction changes the whole sequence into a minimal monochrome sketch.
这里的指令把整个序列变成极简单色素描。
01:05:16Local object replacement and global stylization allow different degrees of visual freedom.
局部物体替换与整体风格化,允许的视觉自由程度不同。
01:05:22Even with a large style change, motion and scene structure should remain readable.
即使风格变化很大,运动和场景结构也应保持可读。
01:05:28Look for rules across frames: line density, silhouette treatment, and the handling of depth.
寻找跨帧规则:线条密度、轮廓处理,以及深度的表现方式。
01:05:35Do they feel like one visual language?
它们是否像同一种视觉语言?
01:05:38This connects directly to our collage discussion.
这与我们的拼贴讨论直接相连。
01:05:42A style is more useful to an art director when described through operations that can persist through time.
当风格被描述为能够持续存在于时间中的操作时,它对艺术总监更有用。
01:05:48If every frame reinvents those operations, the result may feel unstable even when each still is appealing.
如果每一帧都重新发明这些操作,即使每张静帧好看,整体也可能不稳定。
SLIDE 092 · 01:05:56 · Topic 3 / Video editing
A style can flicker while motion stays correct
01:05:56A video can preserve object motion while its style flickers.
视频可以保留物体运动,风格却仍然闪烁。
01:06:00Line density may jump, a paper texture may crawl, or shading may switch between flat and volumetric treatments without a scene explanation.
线条密度可能跳变,纸张纹理可能爬动,明暗处理也可能无故在平面与立体之间切换。
01:06:09Review style as a set of temporal rules, just as we reviewed it across surfaces in the collage.
像审查拼贴不同表面的风格那样,把视频风格作为一组时间规则来审查。
01:06:15Some variation is appropriate when the camera or light changes.
相机或光线变化时,某些变化是合理的。
01:06:19The question is whether the material language remains coherent.
问题是:材质语言是否仍然连贯?
01:06:23If your artwork depends on a particular drawing or collage treatment, this check is as important as whether the subject stays in the same place.
如果作品依赖特定绘画或拼贴处理,这项检查与主体位置是否稳定同样重要。
01:06:32Motion correctness alone does not establish visual consistency.
运动正确,本身并不能证明视觉一致。
SLIDE 093 · 01:06:37 · Topic 3 / Video editing
Why raw frame difference is misleading
01:06:37If we subtract one frame from the next, a perfectly correct camera movement can create a large difference.
如果将相邻帧相减,完全正确的相机运动也会产生很大差异。
01:06:44The same surface has simply moved to another pixel location.
同一个表面,只是移动到了另一个像素位置。
01:06:49A more meaningful comparison first aligns corresponding visible content, then measures the remaining appearance difference.
更有意义的比较,是先对齐对应的可见内容,再测量剩余外观差异。
01:06:57The teaching equation uses a motion warp for that alignment and a visibility mask to exclude occluded regions.
这个教学公式使用运动变换来对齐,并用可见性蒙版排除被遮挡区域。
01:07:04We should not demand agreement for content that is hidden or newly revealed.
对于隐藏或刚被揭示的内容,不应要求直接一致。
01:07:09This is an illustrative metric, not a claim about the exact evaluation used by the paper.
这是示意性指标,不是在声称论文使用了这一精确评测。
01:07:15A low difference can also reward a video that barely moves.
低差异也可能奖励几乎不动的视频。
01:07:19A metric must be checked against the task it is supposed to represent.
必须对照指标所要代表的任务,检查指标本身。
SLIDE 094 · 01:07:24 · Topic 3 / Video editing
The frozen video wins the wrong test
01:07:24Imagine two candidates for our five-second shot.
想象五秒镜头的两个候选版本。
01:07:27One shows a coherent camera move around the pear with modest frame differences.
一个展示相机连贯地绕梨移动,帧间差异适中。
01:07:32The other repeats a single beautiful frame.
另一个反复显示同一张漂亮图像。
01:07:35A naive difference score might prefer the frozen version.
简单的差异分数,可能偏爱冻结版本。
01:07:39Our audience would immediately notice that the intended reveal never happened.
观众却会立刻发现,预期的揭示根本没有发生。
01:07:43That is a useful example of optimizing the measurement while missing the purpose.
这是优化测量指标却错失目的的一个例子。
01:07:48We need evidence for both coherent appearance and requested motion.
我们需要外观连贯和所要求运动两方面的证据。
01:07:52Neither can substitute for the other.
两者不能相互替代。
01:07:55Each time, one easy-to-observe quality tempted us to stand in for the whole brief.
每一次,我们都容易让某个易于观察的品质,代表整份创作要求。
01:08:00Good evaluation keeps the intended result in view, especially when a convenient score seems reassuring.
良好的评测始终关注预期结果,尤其当方便的分数令人安心时。
SLIDE 095 · 01:08:08 · Topic 3 / Video editing
A keyframe can anchor the edit
01:08:08An edited keyframe can make an art direction easier to approve.
编辑后的关键帧,可以让艺术方向更容易获得认可。
01:08:12Instead of describing every material and lighting choice in words, we can point to an image and say, this is the appearance we want.
不必用文字描述每个材质和灯光决定,我们可以指着一张图说:这就是想要的外观。
01:08:21The workflow then carries that approved look through the source clip.
工作流程再把获批的外观延续到原始片段中。
01:08:24Follow the four stages: select a representative frame, edit and inspect its appearance, propagate the look, and review the resulting motion.
请跟随四个阶段:选代表帧,编辑并检查外观,传播外观,再审查运动结果。
01:08:36Will the material persist during a turn?
转向时,材质能保持吗?
01:08:39Will the identity survive a column passing in front of it?
柱子从前方经过时,身份能保留吗?
01:08:44This distinction is central to the Runway reading later: an appealing interaction pattern still needs a shot review suited to the artwork.
这个区别是后面 Runway 阅读的核心:吸引人的交互方式,仍需要适合这件作品的镜头审查。
SLIDE 096 · 01:08:52 · Topic 3 / Video editing
AlignVid: when the image overrules the text
01:08:52What if the approved image is so influential that the requested event never happens?
如果获批图像影响力太强,以至于要求的事件根本没发生,怎么办?
01:08:57Our AlignVid collaboration studies that tension in image-to-video generation.
我们合作开展的 AlignVid 研究,探索图生视频中的这种张力。
01:09:02In the published examples, a baseline omits a sunflower or leaves a person standing; the corresponding AlignVid results implement more of the requested event.
论文示例中,基线漏掉了向日葵,或让人物一直站着;对应的 AlignVid 结果实现了更多要求的事件。
01:09:13The intervention scales queries or keys in selected attention blocks and denoising steps without retraining.
这种干预无需重新训练,只在选定注意力模块和去噪步骤中,缩放查询或键。
01:09:19In a scalar form, scaling Q by gamma changes the weights to softmax of gamma times Q K transpose over square root d.
在标量缩放形式下,用 gamma 缩放 Q 后,权重变成 softmax(gamma × QKᵀ / √d)。
01:09:28It changes attention concentration.
它改变了注意力的集中程度。
01:09:31Classifier-free guidance instead combines predictions with different conditioning; these are distinct operations.
无分类器引导则组合不同条件下的预测;两者是不同操作。
01:09:38For an artist, the interesting failure is a faithful-looking image that refuses to do what the scene requires.
对艺术家来说,有意思的失败是:图像看起来很忠实,却拒绝做场景要求的事。
01:09:44The published frames illustrate that tension; a finished shot still needs temporal review.
已发表的帧展示了这种张力;完成的镜头仍需进行时间连续性审查。
SLIDE 097 · 01:09:50 · Topic 3 / Video editing
A five-second video edit brief
01:09:50Our five-second brief changes the pear's body to ivory ceramic while retaining the blue stem, camera movement, and position.
我们的五秒要求,是把梨的主体变为象牙白陶瓷,同时保留蓝色柄、相机运动和位置。
01:09:58Reflections may adapt to the new material.
倒影可以适应新材质。
01:10:00The pear must remain the same object after passing behind a column.
梨经过柱子后面再出现时,必须仍是同一个物体。
01:10:05Which clause is hardest to verify?
哪一条最难验证?
01:10:08The reappearance is a strong candidate, because the system must maintain identity through a period of invisibility.
再次出现很可能最难,因为系统必须在一段不可见期间保持身份。
01:10:14This is more demanding than transferring a visible color from frame to frame.
这比逐帧传递可见颜色更具挑战。
01:10:19Begin with one clear transformation in a short clip.
先在短片中做一个明确的变换。
01:10:23Then make the critical visibility event part of the review, rather than discovering it only after choosing a favorite result.
然后把关键可见性事件纳入审查,而不是选好最喜欢的结果后才发现它。
SLIDE 098 · 01:10:31 · Topic 3 / Video editing
The moment the pear returns
01:10:31The column is the moment of truth.
柱子就是关键考验。
01:10:34Before the pear disappears, we can inspect its silhouette, material, and stem.
梨消失前,我们能检查它的轮廓、材质和柄。
01:10:39During occlusion, there may be nothing visible to compare.
遮挡期间,可能没有可见内容可供比较。
01:10:43When it returns, the model has to make the same object convincing again.
返回时,模型必须再次让同一个物体显得可信。
01:10:48Do not reject correct invisibility as a failure.
不要把正确的不可见状态当作失败。
01:10:51Instead, compare the identity before and after the event, including how the object emerges at the boundary.
而应比较事件前后的身份,包括物体如何从边界处显露。
01:10:58A smooth-looking clip can still reveal a subtly redesigned pear on the other side.
看起来流畅的视频,也可能在另一侧露出一只被悄悄重新设计的梨。
01:11:03We are not merely watching for anything strange; we are testing the preservation requirement at the point where the shot makes it hardest to satisfy.
我们不只是寻找怪异之处,而是在镜头最难满足要求的地方,检验保留约束。
SLIDE 099 · 01:11:12 · Topic 3 / Video editing
A small addition creates new relationships
01:11:12Adding a small object creates several new relationships.
添加一个小物体,会创造多种新关系。
01:11:16In this published example, the edit adds a drone.
在这个已发表示例中,编辑添加了一架无人机。
01:11:20Even if the drone looks convincing by itself, its scale, placement, and movement must fit the scene.
即使无人机本身很可信,它的尺度、位置和运动也必须适合场景。
01:11:27Imagine a camera moving forward while the added object changes size in the wrong direction.
想象相机向前移动,新增物体的大小却朝错误方向变化。
01:11:32The object might be beautifully rendered, but its relationship with the camera would expose the edit.
物体可能渲染得很漂亮,但它与相机的关系会暴露编辑。
01:11:38Or it might hover at an unintended location relative to other objects.
或者,它可能相对其他物体悬停在错误位置。
01:11:42When you review an addition, evaluate those relationships across time.
审查新增物体时,要跨时间评价这些关系。
01:11:47A close-up of the inserted object cannot answer every question about whether it belongs.
仅看新增物体的特写,无法回答它是否融入场景的所有问题。
SLIDE 100 · 01:11:53 · Topic 3 / Video editing
An addition must belong to camera motion
01:11:53An added object must fit the camera as well as the scene.
新增物体既要符合场景,也要符合相机。
01:11:57A convincing texture and shape in one still do not establish a coherent trajectory.
单张静帧中可信的纹理和形状,不能证明运动轨迹连贯。
01:12:02If the camera approaches, the apparent size and position of the addition should evolve in a compatible way.
如果相机靠近,新增物体的表观大小和位置应相应变化。
01:12:09The exact expectation depends on whether the object is stationary or moving independently.
具体预期取决于物体是静止,还是独立运动。
01:12:15That is why the brief should specify the intended relationship.
因此,创作说明应该明确预期关系。
01:12:19Inspect the addition relative to nearby objects and the background, not only in a crop.
要相对附近物体和背景检查新增内容,而不只看裁剪图。
01:12:26A production-quality edit is a collection of relationships that survive over time, rather than a new object that looks impressive in isolation.
制作级编辑是一组经得起时间考验的关系,而不是一个单独看很惊艳的新物体。
SLIDE 101 · 01:12:35 · Topic 3 / Video editing
The shot review
01:12:35Review the whole shot at normal speed, then inspect moments where failure is most likely.
先以正常速度看完整个镜头,再检查最可能失败的时刻。
01:12:41Pay attention to identity, intended motion, occlusion, boundaries, and the ending.
关注身份、预期运动、遮挡、边界和结尾。
01:12:46High-motion content is a limitation the research authors specifically identify.
高运动量内容,是研究作者明确指出的局限。
01:12:51If a clip fails, choose a revision based on the failure.
片段失败时,要根据具体失败选择修改。
01:12:54You might reduce the transformation, shorten the segment, add a reference frame where supported, or repair part of the result through conventional compositing.
可以减小变换幅度、缩短片段、在支持时添加参考帧,或用传统合成修复部分结果。
01:13:05Another generation is useful when it tests a reasoned hypothesis.
当新一轮生成用于检验有理由的假设时,它才有用。
01:13:09Do not let repeated attempts distract you from whether the piece still communicates the artistic intention you began with.
不要让反复尝试分散注意力,忘记作品是否仍在传达最初的艺术意图。
SLIDE 102 · 01:13:16 · Topic 3 / Video editing
A failed shot suggests a different workflow
01:13:16A failed shot is information about the workflow.
失败镜头提供了关于工作流程的信息。
01:13:20If the appearance is wrong from the beginning, revising the keyframe or the material direction may help.
如果一开始外观就不对,修改关键帧或材质方向可能有帮助。
01:13:26If the first frame is convincing but identity breaks after occlusion, the problem calls for temporal review and a more suitable propagation or editing approach.
如果首帧可信,但遮挡后身份破坏,就需要时间连续性审查,以及更合适的传播或编辑方法。
01:13:36If a region must remain exact, direct compositing may be part of the solution.
如果某个区域必须完全不变,直接合成可以是解决方案的一部分。
01:13:41We can also split a complicated transformation into stages, provided the transitions remain coherent.
也可以把复杂变换拆成多个阶段,但必须保证过渡连贯。
01:13:48The important move is to connect the observed failure to the next intervention.
关键是把观察到的失败,与下一步干预联系起来。
01:13:53Repeating a vague request with more enthusiasm does not use what the failed shot has taught us.
更热情地重复模糊要求,并没有利用失败镜头教给我们的东西。
SLIDE 103 · 01:13:59 · Topic 3 / Video editing
Discuss: the difficult moment
01:13:59Return to the proposed five-second ceramic-pear shot.
回到拟议的五秒陶瓷梨镜头。
01:14:03Which moment would you inspect most closely?
你会最仔细检查哪个时刻?
01:14:05What should remain recognizable after the column?
经过柱子后,什么应该仍然可辨认?
01:14:08What failure would make you reject an attractive clip?
什么失败会让你拒绝一段好看的视频?
01:14:12Discuss these three questions by describing an event and the evidence you would seek.
请通过描述事件和所需证据,讨论这三个问题。
01:14:17If your answer is temporal consistency, make it visible: what changes, when, and why would that violate the intended shot?
如果你的答案是“时间一致性”,请让它可见:什么在何时改变,为什么违反了镜头意图?
01:14:29Compare your acceptance criteria with another person’s before we examine the compact review plan.
在查看简要审查方案之前,先与另一位同学比较验收标准。
SLIDE 104 · 01:14:38 · Topic 3 / Video editing
A five-second test plan
01:14:38Here is a compact test plan.
这是一个简要测试方案。
01:14:41The request is an ivory ceramic body with the blue stem retained.
要求是象牙白陶瓷主体,并保留蓝色柄。
01:14:45Camera motion and object motion should follow the source.
相机运动和物体运动应遵循原片。
01:14:49The clip includes a column so that we can inspect a difficult reappearance.
片段中包含柱子,让我们检查困难的再次出现过程。
01:14:53The evidence includes normal-speed playback and a closer look before and after that event.
证据包括正常速度播放,以及事件前后的仔细观察。
01:14:59We also inspect reflections because a material change has optical consequences.
还要检查倒影,因为材质变化会产生光学后果。
01:15:04This plan does not require a long benchmark suite.
这个方案不需要庞大的基准测试集。
01:15:08It simply makes the request, the likely failure, and the acceptance evidence explicit.
它只是明确要求、可能的失败,以及验收证据。
01:15:14You can use the same structure to test a different subject or visual transformation.
你可以用同样的结构测试其他主体或视觉变换。
SLIDE 105 · 01:15:19 · Topic 3 / Video editing
Three decisions without a model
01:15:19Let us check whether the mechanisms are now useful without a model in front of us.
现在没有模型在面前,让我们检验这些机制是否真的有用了。
01:15:24Consider the three questions on screen.
请考虑屏幕上的三个问题。
01:15:27A material change reaches the reflection: explain why.
材质变化会影响倒影:请解释原因。
01:15:32A subject reference and LoRA both help represent a concept: explain the difference.
主体参考图和 LoRA 都有助于表达概念:请解释两者区别。
01:15:39Four frames look excellent: explain why the clip can still fail.
四帧看起来很出色:请解释整段视频为什么仍可能失败。
01:15:44Take a short silent beat before answering.
回答前,先安静思考片刻。
01:15:47The next slide names the decisions hiding inside the questions.
下一页会指出这些问题中隐藏的决定。
01:15:51If you get stuck, locate the type of problem first: the physical scene, the model's information, or evidence across time.
如果卡住了,先定位问题类型:物理场景、模型获得的信息,还是跨时间的证据。
01:16:00That is often enough to begin a clear answer.
通常,做到这一步就足以开始给出清晰答案。
SLIDE 106 · 01:16:03 · Topic 3 / Video editing
Match the problem to the intervention
01:16:03Here are the decisions those questions conceal.
这些问题隐藏着以下决定。
01:16:06If the reflection must change with the new material, we need to define the permitted region and consequences of the edit.
如果倒影必须随新材质变化,就需要定义编辑允许的区域和连带后果。
01:16:14If we want a concept for this request, a reference can condition generation; if we want learned adaptation, LoRA changes selected weights.
如果只为这次要求表达概念,参考图可提供生成条件;如果需要学习式适配,LoRA 会改变选定权重。
01:16:25If continuity matters, the evidence must include the intervals between our chosen stills.
如果重视连续性,证据就必须包含精选静帧之间的间隔。
01:16:30Each decision rules out a tempting shortcut.
每个决定,都排除了一条诱人的捷径。
01:16:34Freezing every surrounding pixel may freeze the wrong reflection.
冻结周围每一个像素,可能冻结错误的倒影。
01:16:38Calling every supplied image training confuses conditioning with adaptation.
把所有输入图像都称为训练,会混淆条件控制与适配。
01:16:43Calling four stills a successful video omits time.
把四张静帧称为成功视频,就遗漏了时间。
01:16:47Use this slide to repair your explanation, rather than memorize a preferred sentence.
请用这一页修正你的解释,而不是背诵某句标准答案。
SLIDE 107 · 01:16:53 · Topic 3 / Video editing
The mechanisms behind your decisions
01:16:53The material changes how light interacts with the pear, so its visible consequences can extend into the reflection.
材质改变光与梨的相互作用,因此可见后果会延伸到倒影。
01:17:01A reference conditions an output; LoRA learns a low-rank update to selected model weights.
参考图为输出提供条件;LoRA 为选定模型权重学习低秩更新。
01:17:07Four good frames leave unobserved transitions, including events such as occlusion where identity can fail.
四张好帧仍留下未观察的过渡,包括可能导致身份失败的遮挡事件。
01:17:14Those are compact answers, but each has an application.
这些答案很简短,但每个都有实际用途。
01:17:18They help us write a better preservation contract, choose between two kinds of intervention, and design a more revealing review.
它们帮助我们写出更好的保留约定、在两类干预中选择,并设计更有揭示性的审查。
01:17:26If your answer used different words and preserved those distinctions, it works.
如果你用了不同措辞,但保留了这些区别,答案就成立。
01:17:31Tomorrow's interface may rename its controls, but the difference between specifying a request, adapting a model, and checking its result will still matter.
明天的界面可能给控件改名,但提出要求、适配模型与检查结果之间的区别,仍然重要。
SLIDE 108 · 01:17:41 · Topic 3 / Video editing
Three levels of checking an edit
01:17:41We can check an edit at three levels.
我们可以在三个层面检查编辑。
01:17:43First, did the requested transformation occur?
第一,所要求的变换发生了吗?
01:17:47Second, did the necessary invariants survive?
第二,必要的不变量保留了吗?
01:17:50Third, does the result serve the artwork's intention?
第三,结果是否服务于作品意图?
01:17:54A result can pass the first level and fail the second, as when the correct material appears on a different object.
结果可能通过第一层却不通过第二层,例如正确材质出现在了另一个物体上。
01:18:01It can pass both technical levels and still fail the artistic one, as when the chosen effect undermines the intended mood.
也可能通过两个技术层面,却在艺术层面失败,例如所选效果削弱了预期情绪。
01:18:10Keeping the levels separate makes critique more precise.
区分这些层面,会让评论更精确。
01:18:14It also explains why neither a generic aesthetic score nor a strict pixel comparison can answer every question we care about.
这也解释了为什么通用审美分数和严格像素比较,都无法回答我们关心的所有问题。
SLIDE 109 · 01:18:22 · Topic 3 / Video editing
Who decides?
01:18:22We have spent the lecture making the editing request more precise.
整堂课,我们都在让编辑要求更精确。
01:18:27Now consider who gets to choose the request in the first place.
现在想想:一开始由谁来选择提出什么要求?
01:18:31A system can offer convincing alternatives, but somebody still decides which intention matters and which evidence counts as success.
系统可以提供可信方案,但仍需要有人决定什么意图重要,以及什么证据才算成功。
01:18:41The three readings let us examine that responsibility from different positions.
三篇阅读让我们从不同位置审视这种责任。
01:18:46Pachocki raises questions about capable AI systems and human values.
Pachocki 提出了能力强大的 AI 系统与人类价值观之间的问题。
01:18:51Koe asks readers to reconsider their own goals and habits.
Koe 请读者重新审视自己的目标与习惯。
01:18:55Runway presents a workflow for turning an approved image into a video edit.
Runway 展示把获批图像转化为视频编辑的工作流程。
01:18:59Keep our pear in mind as we read them: we can delegate parts of its production while still arguing about what the artwork should become.
阅读时请记住这只梨:我们可以委托部分制作,同时仍讨论作品应该成为什么。
SLIDE 110 · 01:19:08 · Topic 3 / Video editing
After Topic 3: judge the whole shot
01:19:08We have an approved appearance and a moving shot to judge.
现在,我们有了获批外观,以及需要评价的运动镜头。
01:19:11What can an approved keyframe establish, and what can it not?
获批关键帧能确定什么,又不能确定什么?
01:19:16Why might a low frame-difference score reward a bad video?
为什么低帧差分数可能奖励一段坏视频?
01:19:20Which evidence would convince you that identity survived motion?
什么证据能说服你,身份经受住了运动?
01:19:24Discuss these three questions.
请讨论这三个问题。
01:19:26Your answer should account for the requested event as well as the subject’s appearance.
答案既要考虑主体外观,也要考虑所要求的事件。
01:19:30Use the column, a turn, or the sketch treatment as a concrete example.
可以用柱子、转向或素描处理作为具体例子。
01:19:36We will next ask how the readings change our view of responsibility for those decisions.
接下来,我们将追问阅读如何改变我们对这些决定所涉及责任的看法。
SLIDE 111 · 01:19:45 · Readings and technical extensions
AI goals and human direction
01:19:45These first two readings ask about direction from different sides.
前两篇阅读从不同方面追问方向。
01:19:49An Alien Mind considers the relationship between an AI system achieving a goal and respecting human values.
《An Alien Mind》思考 AI 系统实现目标与尊重人类价值观之间的关系。
01:19:57Dan Koe's essay invites people to examine the goals and habits directing their own lives.
Dan Koe 的文章邀请人们审视引导自己生活的目标和习惯。
01:20:02We will use its proposal as material for critical discussion of creative practice.
我们把他的主张作为批判性讨论创作实践的材料。
01:20:07Find a specific claim, explain your interpretation, and connect it to a decision from the lecture.
找出一个具体主张,解释你的理解,并联系课堂中的一个决定。
01:20:14A fictional artist is fine for discussing personal direction.
讨论个人方向时,可以使用虚构艺术家的例子。
01:20:18The fifth quiz question asks you to reflect on one reading of your choice.
测验第五题要求你从自选的一篇阅读出发进行反思。
01:20:23Our class discussion can compare all three perspectives without requiring everyone to disclose personal experiences.
课堂可以比较三种视角,而不要求每个人披露个人经历。
SLIDE 112 · 01:20:30 · Readings and technical extensions
Reading 3: image-guided video editing
01:20:30The third reading moves from questions about direction to a concrete editing workflow.
第三篇阅读从方向问题,转向具体编辑流程。
01:20:36Runway's announcement of Aleph 2.0 and Edit Studio describes establishing an appearance in an edited image and applying that change through video.
Runway 关于 Aleph 2.0 和 Edit Studio 的公告,描述了先在编辑图像中确定外观,再把变化应用到视频的方法。
01:20:45It connects directly to our keyframe discussion.
它与关键帧讨论直接相连。
01:20:48Read it with the ceramic pear and column in mind.
阅读时,请想着陶瓷梨和柱子。
01:20:51Which decision does the interface make easier to express?
这个界面让哪项决定更容易表达?
01:20:55Which result would you still need to watch before approval?
批准之前,哪些结果仍需要亲眼播放检查?
01:20:59The page is a vendor's account, so its demonstrations illustrate claimed capabilities rather than supply an independent comparative evaluation.
这是一篇厂商介绍,因此演示说明的是其声称的能力,而不是独立的比较评测。
01:21:08We can learn from the interaction design and still propose an occlusion test that asks whether the workflow meets this exhibition's particular needs.
我们可以借鉴交互设计,同时提出遮挡测试,检验流程是否满足这个展览的具体需要。
SLIDE 113 · 01:21:17 · Readings and technical extensions
Optional technical reading
01:21:17These technical papers are optional companions to the discussion readings.
这些技术论文,是讨论阅读之外的可选配套材料。
01:21:21Choose according to the question you want to investigate.
请根据你想研究的问题选择。
01:21:24InstructPix2Pix helps explain edit supervision; Flow Matching develops the generative training framework.
InstructPix2Pix 帮助解释编辑监督;Flow Matching 发展了生成训练框架。
01:21:32Qwen-Image-2.0 and Qwen-Video-Edit provide recent architecture examples.
Qwen-Image-2.0 和 Qwen-Video-Edit 提供了近期架构示例。
01:21:38You do not need to read all of them before trying the class discussion.
参加课堂讨论之前,不必读完全部论文。
01:21:42Start with a question, locate the part of the paper that addresses it, and distinguish the proposed method from the authors' evidence.
从一个问题开始,找到论文中回应它的部分,并区分提出的方法与作者的证据。
01:21:50The slide notes retain the other references on multi-image inputs, rewards, composition, and video editing.
幻灯片备注保留了多图输入、奖励、构图和视频编辑方面的其他参考资料。
SLIDE 114 · 01:21:58 · Readings and technical extensions
Three source types, three kinds of evidence
01:21:58The three readings provide different kinds of material.
三篇阅读提供不同类型的材料。
01:22:02A research leader's perspective develops an argument about alignment and future development.
研究领导者的观点文章,提出关于对齐和未来发展的论证。
01:22:07A reflective essay offers a way to examine personal direction.
反思性文章提供一种审视个人方向的方法。
01:22:12A vendor announcement presents a workflow and promotes its capabilities.
厂商公告展示工作流程,并推广其能力。
01:22:16Read each according to what it can support.
阅读每种材料时,要考虑它能够支持什么。
01:22:19Identify forecasts, personal claims, and demonstrations rather than treating every sentence as the same kind of evidence.
识别预测、个人主张和演示,不要把每句话都当作同类证据。
01:22:27You may disagree with an author and still find a useful question.
你可以不同意作者,却仍从中找到有用的问题。
01:22:31Our synthesis is practical: what do you want to make, what will you delegate, and what evidence will you use to decide whether the collaboration served that intention?
我们的综合问题很实际:你想创作什么、委托什么,以及用什么证据判断合作是否服务于这一意图?
SLIDE 115 · 01:22:42 · Readings and technical extensions
Separate image and text guidance
01:22:42This technical extension separates image guidance and text guidance in InstructPix2Pix.
这个技术延伸,区分 InstructPix2Pix 中的图像引导和文本引导。
01:22:49Begin with the prediction using neither condition.
从不使用任何条件的预测开始。
01:22:52Add a scaled difference for including the image, then another scaled difference for adding the text alongside that image.
加上引入图像后的缩放差值,再加上在图像之外引入文本后的另一个缩放差值。
01:22:59Set both scales to one and follow the cancellation: the intermediate terms disappear, leaving the fully conditioned noise prediction.
把两个强度都设为一,观察抵消:中间项消失,只留下完全条件化的噪声预测。
01:23:07This is a useful check that you understand the expression.
这是检查你是否理解表达式的有用方法。
01:23:11Epsilon here denotes a noise prediction, unlike the velocity in our flow-matching slides.
这里的 epsilon 表示噪声预测,与流匹配页面中的速度不同。
01:23:18These controls belong to this formulation and should not be assumed to map directly onto every current editor's interface.
这些控制属于这一公式,不应假设它们直接对应所有当前编辑器的界面。
SLIDE 116 · 01:23:26 · Readings and technical extensions
What remains when the terms cancel?
01:23:26Now set both guidance scales to one and follow the cancellation.
现在把两个引导强度都设为一,观察抵消。
01:23:30The initial no-condition prediction cancels its negative copy in the first difference.
初始无条件预测,与第一个差值中的负项抵消。
01:23:36The image-only prediction then cancels its negative copy in the second difference.
随后,仅图像条件的预测,与第二个差值中的负项抵消。
01:23:41The result is the prediction with both image and text conditions.
最终得到同时使用图像和文本条件的预测。
01:23:45That gives us a reference point for understanding the two controls.
这为理解两个控制提供了参考点。
01:23:49If the text scale were zero while the image scale stayed one, we would instead recover the image-only prediction.
如果文本强度为零、图像强度仍为一,就会得到仅使用图像条件的预测。
01:23:57It shows what information each difference adds.
这说明了每个差值添加什么信息。
01:24:00Keep epsilon's meaning explicit here: this InstructPix2Pix expression predicts noise, whereas the earlier flow-matching example predicted velocity.
请明确 epsilon 的含义:这里的 InstructPix2Pix 公式预测噪声,而前面的流匹配例子预测速度。
SLIDE 117 · 01:24:11 · Readings and technical extensions
Training: the target supplies the lesson
01:24:11Read the pseudocode in two groups.
把伪代码分成两组来读。
01:24:14First we prepare an example: source, instruction, target, conditions, and target latent.
首先准备样本:原图、指令、目标、条件,以及目标潜变量。
01:24:23Then we sample noise and time, construct the intermediate latent, predict velocity, and compare that prediction with the known training target.
然后采样噪声和时间,构造中间潜变量,预测速度,并将预测与已知训练目标比较。
01:24:32The final update changes trainable weights based on the loss.
最后根据损失更新可训练权重。
01:24:37At inference, there is no known edited target to supply in this way.
推理时,并没有已知编辑目标可以这样提供。
01:24:41We instead follow the learned generative predictions from a starting state.
我们从初始状态出发,遵循学习得到的生成预测。
01:24:46This pseudocode explains the conceptual loop; it omits practical engineering such as batches, precision, and scheduling.
这段伪代码解释概念循环,省略了批处理、精度和调度等实际工程细节。
01:24:55Its most important distinction is what information is available during learning versus use.
最重要的区别,是学习与使用时分别有哪些信息可用。
SLIDE 118 · 01:25:01 · Readings and technical extensions
Inference: where did the target go?
01:25:01Look for the line that disappeared.
请找出消失的那一行。
01:25:03The inference loop has no known edited target and no loss that updates weights.
推理循环没有已知编辑目标,也没有用于更新权重的损失。
01:25:08Instead, we encode the source and instruction, begin with noise, repeatedly predict a velocity and move the latent, then decode the result.
我们编码原图和指令,从噪声开始,反复预测速度并移动潜变量,最后解码结果。
01:25:18Training uses examples to adjust the learned model.
训练用样本调整学习到的模型。
01:25:21In this simplified inference loop, we use that model to construct an output we do not yet possess.
在这个简化推理循环中,我们用该模型构造尚未拥有的输出。
01:25:28The code is conceptual; practical systems use their own representations and sampling schedules.
代码是概念性的;实际系统使用各自的表示和采样调度。
01:25:35But you can now explain why providing another reference changes the information available for a request without automatically becoming a model-training operation.
但你现在可以解释:增加参考图会改变请求可用的信息,却不会自动变成模型训练。
01:25:44It is the same distinction we used when choosing between references and LoRA.
这与选择参考图还是 LoRA 时使用的区别相同。
SLIDE 119 · 01:25:50 · Readings and technical extensions
Choose the approach by the requirement
01:25:50Before we turn to the readings' discussion questions, use this table as a compact decision aid.
转向阅读讨论问题之前,请把这张表当作简明决策辅助。
01:25:56Exact untouched pixels suggest a role for masks and compositing.
如果未编辑像素必须精确不变,可以考虑蒙版与合成。
01:26:00A new visual language may benefit from reference conditioning.
新的视觉语言,可能受益于参考图条件控制。
01:26:04A reusable learned concept may motivate adaptation.
可复用的学习概念,可能需要模型适配。
01:26:08A transformation that must survive motion needs a video workflow and temporal evidence.
必须经受运动的变换,需要视频工作流程和时间证据。
01:26:14These approaches can cooperate in the same artwork.
这些方法可以在同一件作品中协作。
01:26:17Choose according to the contract, then inspect the failure most likely to undermine it.
根据约定选择,再检查最可能破坏约定的失败。
01:26:23That leaves us with a larger question for the readings: once production becomes easier, how do we choose a worthwhile direction and retain meaningful judgment over the result?
这给阅读留下更大的问题:制作更容易之后,如何选择值得追求的方向,并对结果保留有意义的判断?
SLIDE 120 · 01:26:34 · Closing reading discussions
An Alien Mind
01:26:34An Alien Mind distinguishes achieving an assigned goal from generalizing human values in unfamiliar circumstances.
《An Alien Mind》区分了完成指定目标,与在陌生情境中泛化人类价值观。
01:26:41Pachocki also raises questions about monitoring increasingly capable systems.
Pachocki 也提出了如何监测日益强大系统的问题。
01:26:47Treat the essay as an argument containing claims and forecasts that we can examine.
请把文章视为包含主张与预测的论证,我们可以对其检视。
01:26:52For our discussion, imagine an exhibition team delegating production while retaining responsibility for what the work communicates.
讨论时,想象一个展览团队委托制作,但仍对作品传达的内容负责。
01:27:00Discuss the three questions on screen: explain the goal-and-values distinction, identify a forecast and the evidence it would need, and defend a boundary for human control.
讨论屏幕上的三个问题:解释目标与价值观的区别,找出一个预测及所需证据,并为人类控制的边界辩护。
01:27:14Our art-direction example is an analogy; it does not give an image editor the same agency or risk profile as an autonomous research system.
艺术指导只是类比;它并不赋予图像编辑器与自主研究系统相同的能动性或风险特征。
SLIDE 121 · 01:27:27 · Closing reading discussions
How to fix your entire life in 1 day
01:27:27Koe's title makes a dramatic promise: How to fix your entire life in one day.
Koe 的标题作出了强烈承诺:《如何在一天内修复你的整个生活》。
01:27:32The essay proposes examining identity and goals, interrupting habitual behavior, and turning reflection into action.
文章建议审视身份与目标、打断习惯行为,并把反思转化为行动。
01:27:41We can discuss the usefulness of that proposal without accepting the title as a guarantee.
我们可以讨论这些建议是否有用,而不必把标题当作保证。
01:27:46Imagine an artist who can generate a hundred attractive images but cannot choose what to make.
想象一个艺术家,能生成一百张好看的图像,却无法决定要创作什么。
01:27:52Does easy production help clarify a direction, or make avoiding the decision easier?
轻松制作会帮助明确方向,还是让逃避决定更容易?
01:27:58Discuss the three questions on screen.
请讨论屏幕上的三个问题。
01:28:01Choose an idea worth using or challenging, explain the reason, and connect it to creative intention.
选一个值得采用或质疑的观点,解释理由,并把它与创作意图联系起来。
01:28:10You can use a fictional artist or public example; no personal disclosure is needed.
可以使用虚构艺术家或公开案例,不需要披露个人经历。
SLIDE 122 · 01:28:20 · Closing reading discussions
Introducing Aleph 2.0 and Edit Studio
01:28:20Runway's reading proposes approving an edited image before applying its appearance through a video.
Runway 的阅读提出:先批准编辑图像,再把它的外观应用到整段视频。
01:28:26It offers a concrete answer to a communication problem: an art director can point to the desired look, rather than describe every feature in words.
它具体回答了一个沟通问题:艺术总监可以指出目标外观,而不必用文字描述每个特征。
01:28:35Now bring back the column.
现在,让柱子重新出现。
01:28:37A convincing keyframe does not tell us what the pear will look like after it reappears.
可信关键帧并不能告诉我们,梨再次显露后会是什么样子。
01:28:42Discuss the three questions on screen: what the frame establishes, which claim deserves a harder test, and how the workflow serves AFTER RAIN's intention.
讨论屏幕上的三个问题:这一帧确定了什么,哪个主张需要更严格测试,以及流程如何服务于《雨后》的意图。
01:28:54The source is a vendor announcement; our job is to distinguish a useful demonstrated workflow from a reliability claim still needing evaluation.
来源是厂商公告;我们的任务是区分有用的已演示流程,与仍需评价的可靠性主张。
SLIDE 123 · 01:29:07 · Closing reading discussions
Your final judgment
01:29:07Return to your first judgment of the glass pear.
回到你最初对玻璃梨的判断。
01:29:11We began with a small request and discovered that it touched the scene's physics, the artwork's intention, and the audience's experience over time.
我们从一个小要求出发,发现它触及场景物理、作品意图,以及观众随时间展开的体验。
01:29:20You now have more precise ways to say what should change, what should survive, and how to judge the result.
现在,你有更精确的方法说明什么应改变、什么应保留,以及如何判断结果。
01:29:27Finish with the three questions on screen.
最后,请讨论屏幕上的三个问题。
01:29:30Where would you place the boundary between control and surprise?
你会把控制与惊喜之间的边界放在哪里?
01:29:33How would you balance technical success and artistic purpose?
你会如何平衡技术成功与艺术目的?
01:29:37What would you delegate, and what evidence would you require?
你会委托什么,又要求什么证据?
01:29:41Pause the video for the final discussion.
请暂停视频,进行最终讨论。
01:29:44Connect one mechanism or reading to a concrete artistic decision.
把一种机制或一篇阅读,与具体艺术决定联系起来。
01:29:49Listen for an answer that makes you revise your own.
留意一个能让你修正自己看法的回答。
01:29:52That revision is a fitting last act for a class about editing.
对一堂关于编辑的课来说,这样的修正,是恰当的最后一笔。