← Watch Lecture 02

AMCC 5160 · LECTURE 02 · 90 MINUTES

Generative editing
English / 中文

Read the lecture in English and Chinese. Select a timestamp to watch the corresponding moment in the video.

SLIDE 001 · 00:00:00 · Opening

Generative Editing

00:00:00

Imagine you are the art director.

想象一下,你是艺术总监。

00:00:02

Your exhibition opens tomorrow.

你的展览明天开幕。

00:00:04

You love this glass pear, but you want to see it in ivory ceramic.

你很喜欢这只玻璃梨,但想看看它变成象牙白陶瓷的样子。

00:00:08

You give the editor a tiny request: change the material, keep everything else.

你给编辑器一个小小的要求:改变材质,其他一切保持不变。

00:00:13

Then you notice the reflection.

然后,你注意到了倒影。

00:00:15

Should it still look like glass?

它还应该像玻璃一样吗?

00:00:18

Suddenly, a three-word edit contains a whole argument about the world.

突然,一个简单的编辑指令,包含了关于世界如何运作的一整套判断。

00:00:22

That is our starting point.

这就是我们的起点。

00:00:24

We will follow this fictional artwork through image editing, exhibition design, and a moving shot.

我们将跟随这件虚构作品,经历图像编辑、展览设计,再进入一段运动镜头。

00:00:31

The exhibition is called AFTER RAIN.

展览名叫《雨后》,英文是 AFTER RAIN。

00:00:34

The pear is our recurring character.

这只梨会反复登场。

00:00:37

By the end, you should be able to explain a model's mechanism, defend a visual choice, and catch a failure that a beautiful preview can hide.

到课程结束,你应该能够解释模型机制、为视觉选择提供理由,并发现精美预览可能掩盖的失败。

00:00:46

First, look carefully at what you think must survive.

首先,仔细看看你认为必须保留的东西。

SLIDE 002 · 00:00:50 · Opening

One object, three creative decisions

00:00:50

We are going to give the same object three different jobs.

我们要让同一个物体承担三种不同的任务。

00:00:53

First it is a technical puzzle: can its material change while its identity survives?

首先,它是一个技术难题:能否改变材质,同时保留它的身份特征?

00:00:59

Then it becomes an artwork: can its placement and lighting make us feel that it is fragile?

接着,它成为艺术作品:能否通过摆放和灯光,让我们感到它的脆弱?

00:01:06

Finally it becomes a character in time: can it disappear behind something and return as the same object?

最后,它成为时间中的角色:能否在被遮挡后重新出现,仍然是同一个物体?

00:01:12

Those jobs demand different judgments.

这些任务需要不同的判断。

00:01:15

A technically clean image may say very little.

技术上干净的图像,可能表达得很少。

00:01:18

An expressive poster may contain an intentional impossibility.

富有表现力的海报,可能故意包含不可能的景象。

00:01:23

A wonderful still may belong to a broken video.

一张精彩的静帧,可能来自一段存在问题的视频。

00:01:26

As its job changes, you may find yourself approving a change that you would have rejected five minutes earlier.

随着任务改变,你可能会接受一个五分钟前还会拒绝的变化。

SLIDE 003 · 00:01:33 · Opening

The reflection is part of the edit

00:01:33

For a few seconds, ignore the pear.

先暂时别看梨。

00:01:36

Look at the water beneath it.

看看它下面的水。

00:01:38

Now look back at the body.

现在再看回梨的主体。

00:01:40

If the body becomes opaque ceramic, what happens to the light that used to pass through the glass?

如果主体变成不透明的陶瓷,原先穿过玻璃的光会怎样?

00:01:46

The answer cannot live entirely inside the object's outline.

答案不可能完全局限在物体轮廓之内。

00:01:49

This comparison is a prepared classroom illustration.

这组对比是事先准备的课堂示意图。

00:01:53

Use it to locate the consequences of the request: transmission, highlights, and the reflection below.

用它来找出指令引发的影响:透光、高光,以及下方的倒影。

00:02:00

We still want the blue stem and recognizable silhouette.

我们仍然希望保留蓝色的柄和可辨认的轮廓。

00:02:04

But preserving every surrounding pixel could preserve the wrong optics. a successful local edit may require a carefully justified change somewhere else.

但保留周围的每一个像素,可能会保留错误的光学效果。成功的局部编辑,有时需要在别处做出有充分理由的改变。

SLIDE 004 · 00:02:14 · Opening

Necessary change or accidental drift?

00:02:14

Let us make the opening comparison more disciplined.

让我们更严谨地审视开场的对比。

00:02:17

The disappearance of transmitted amber light is consistent with changing transparent glass into opaque ceramic.

从透明玻璃变为不透明陶瓷后,透过的琥珀色光消失,是合理的。

00:02:25

A changed blue stem, by contrast, would need a separate justification because the instruction did not request that transformation.

相比之下,如果蓝色的柄也变了,就需要另外解释,因为指令并没有要求这种变化。

00:02:33

The reflection is more interesting.

倒影则更有意思。

00:02:35

A material change can require its appearance to change, while its placement should still agree with the object and the water.

材质改变可能要求倒影外观随之变化,但它的位置仍应与物体和水面一致。

00:02:44

We cannot classify every difference using a simple rule that change is bad.

我们不能用“变化就是坏事”这一简单规则来判断每个差异。

00:02:49

We need a model of the intended scene.

我们需要理解目标场景应当如何运作。

00:02:51

That is why the edit contract includes allowed consequences as well as invariants.

因此,编辑约定既要包含不变量,也要包含允许发生的连带变化。

SLIDE 005 · 00:02:57 · Opening

Tonight’s route

00:02:57

Here is the route through that puzzle.

下面是我们破解这个难题的路线。

00:03:00

We begin inside an image editor: its representations, sampling process, and controls.

我们先进入图像编辑器内部,研究它的表示方式、采样过程和控制手段。

00:03:09

Then we take the art director's seat and decide what those controls should accomplish.

然后坐到艺术总监的位置,决定这些控制手段应该实现什么。

00:03:14

In the final technical section, the image starts moving and the preservation problem becomes harder.

在最后一个技术部分,图像开始运动,保留原有信息的问题也变得更难。

00:03:21

The spoken script is planned for about ninety minutes.

讲解稿按约九十分钟规划。

00:03:24

The marked discussions, the break, and student presentations need additional class time.

标出的讨论、课间休息和学生展示,需要额外的课堂时间。

00:03:30

You will see the same pear return in different situations.

你会看到同一只梨出现在不同情境中。

00:03:34

Each return asks us to revise our judgment, rather than learn a completely unrelated example.

每次重逢,都要求我们修正判断,而不是学习一个毫不相关的新例子。

00:03:41

At the end, three readings let us question who sets the goal and who decides whether it was achieved.

最后三篇阅读,将让我们追问:谁来设定目标,又由谁判断目标是否达成?

SLIDE 006 · 00:03:47 · Opening

Three decisions you can defend

00:03:47

Here are three decisions I want you to be able to defend by the end.

到课程结束,我希望你能为这三类决定提供理由。

00:03:51

When an edit fails, which control would you change, and why?

编辑失败时,你会调整哪一种控制手段,为什么?

00:03:56

When two posters both look good, which belongs in this exhibition?

两张海报都很好看时,哪一张更适合这个展览?

00:04:01

When a video looks convincing, which moment would you inspect before accepting it?

视频看起来可信时,你会在接受它之前检查哪个时刻?

00:04:07

You do need a link between mechanism and consequence.

你确实需要把机制与结果联系起来。

00:04:10

A mask tells us where; a reference can show us what; an artistic brief tells us why the result matters.

蒙版告诉我们改哪里;参考图可以展示改成什么;艺术创作说明告诉我们结果为何重要。

00:04:17

Today we will practice making those links aloud.

今天,我们会练习把这些联系说清楚。

00:04:21

The aim is to leave with reasons you can use when the next tool has a different interface.

目标是让你带走一套理由,即使下一款工具换了界面,也能继续使用。

SLIDE 007 · 00:04:26 · Part 1 / Image editing

How image editors work

00:04:26

The pear gives us a concrete technical problem: change its material while retaining the information that makes it recognizable.

这只梨给我们一个具体的技术问题:改变材质,同时保留让人认出它的信息。

00:04:34

We can describe this as conditional generation under preservation constraints.

我们可以把它描述为:在保留约束下进行条件生成。

00:04:39

The inputs specify the change; the contract specifies what should survive.

输入指定要改变什么;编辑约定指定要保留什么。

00:04:44

Throughout this section, connect each mechanism to one part of that problem.

在这一部分,请把每种机制对应到这个问题的某一环节。

00:04:49

We begin by making our expectations explicit.

我们先把预期说清楚。

SLIDE 008 · 00:04:53 · Part 1 / Image editing

Before Part 1: image editing

00:04:53

Before we open the machinery, decide what success should look like.

在打开技术黑箱之前,先确定什么才算成功。

00:04:57

When glass becomes ceramic, what should stay the same?

玻璃变成陶瓷时,哪些东西应该保持不变?

00:05:00

Which parts of the scene must change with the material?

场景的哪些部分必须随材质一起改变?

00:05:03

What could a model misunderstand in that request?

模型可能会如何误解这个要求?

00:05:07

Discuss these three questions and point to visible evidence for your choices.

请讨论这三个问题,并指出支持你们选择的可见证据。

00:05:11

If your partner wants to preserve a detail that you want to change, identify the artistic intention behind each answer.

如果同伴想保留的细节恰好是你想改变的,请找出两种答案背后的艺术意图。

00:05:19

Pause the video here.

请在这里暂停视频。

00:05:21

We will return to these expectations when we write the worked ceramic-edit contract.

等到我们为陶瓷编辑写出完整约定时,再回头检视这些预期。

SLIDE 009 · 00:05:30 · Part 1 / Image editing

The edit contract

00:05:30

The gallery has become a forest.

画廊变成了森林。

00:05:32

Let us write the edit contract.

让我们写下编辑约定。

00:05:34

The intended change is the environment.

我们想改变的是环境。

00:05:37

The pear's identity, its pedestal, and the framing should remain recognizable.

梨的身份特征、底座和构图,都应该仍然可辨认。

00:05:43

The ambient light and reflections are allowed to adapt to the forest.

环境光和反射则允许适应森林。

00:05:48

Why include that last category?

为什么要加入最后这一类?

00:05:50

Because preserving every visible relationship would contradict the new setting.

因为保留所有可见关系,会与新环境矛盾。

00:05:55

Imagine retaining a bright gallery-window reflection inside a dark forest.

想象一下,在昏暗森林里还保留着明亮画廊窗户的反光。

00:05:59

The model might preserve the source faithfully and still produce an implausible scene.

模型可能忠实保留了原图,却仍然生成了不合理的场景。

00:06:05

Before generating, separate the requested change, the invariants, and the consequences that should adapt.

生成之前,要区分所要求的变化、不变量,以及应当随之调整的连带效果。

00:06:12

Afterwards, inspect those categories separately.

生成之后,再分别检查这些类别。

SLIDE 010 · 00:06:16 · Part 1 / Image editing

An invariant can be semantic or exact

00:06:16

Suppose a curator says, keep the pear exactly the same.

假设策展人说:让这只梨完全保持原样。

00:06:20

There are at least two meanings hiding in that sentence.

这句话至少隐藏着两种含义。

00:06:24

One means that a visitor should recognize the same sculpture in a new room.

一种是:观众在新房间里,仍然能认出这是同一件雕塑。

00:06:29

Its highlights may change with the room.

它的高光可以随着房间而变化。

00:06:31

The other means that selected pixel values must remain identical.

另一种是:选定像素的数值必须完全一致。

00:06:36

Those are different requirements.

这是两种不同的要求。

00:06:38

Generative conditioning can encourage recognizable identity.

生成式条件控制可以帮助保留可辨认的身份特征。

00:06:41

Copying source pixels or compositing can enforce exact preservation in a specified region.

复制原图像素或进行合成,可以强制指定区域完全不变。

00:06:47

Neither requirement automatically makes the whole image coherent: a copied region may now have the wrong lighting.

但这两种要求都不会自动保证整幅图像协调:复制过来的区域,光照可能已经不合适。

00:06:55

Before choosing a tool, settle the meaning of same.

选择工具之前,先确定“相同”到底是什么意思。

SLIDE 011 · 00:06:59 · Part 1 / Image editing

Editing as conditional generation

00:06:59

Read the expression as a distribution of possible edited results, given the source, instruction, and optional references.

把这个表达式理解为:给定原图、指令和可选参考信息后,可能产生的编辑结果的分布。

00:07:07

The symbol theta represents the model's learned parameters.

符号 theta 表示模型学习到的参数。

00:07:11

The vertical bar means 'given.' We are not asking for a unique answer in every case.

竖线表示“给定”。我们并不是在所有情况下都要求唯一答案。

00:07:17

Two different forest scenes might both satisfy the same brief.

两个不同的森林场景,都可能满足同一份创作说明。

00:07:21

However, a source image is a condition rather than a promise that every unmentioned pixel will be copied.

但原图只是条件,并不意味着每一个未提及的像素都会被复制。

00:07:27

To judge success, we need a more precise contract than 'it looks plausible.'

要判断是否成功,我们需要比“看起来合理”更精确的约定。

SLIDE 012 · 00:07:33 · Part 1 / Image editing

A condition is not a hard constraint

00:07:33

The probability expression says that the system produces a candidate given conditions.

概率表达式表示:系统在给定条件下生成一个候选结果。

00:07:38

It does not contain a certificate that the candidate passes our checks.

它并不附带证明,保证这个结果通过我们的检查。

00:07:43

The acceptance step is a separate decision that we impose on the result.

是否接受,是我们对结果做出的另一个独立判断。

00:07:47

Imagine two forest outputs.

想象两个森林版本。

00:07:50

Both look plausible, but one changes the pear's stem and one retains it.

两个都看起来合理,但一个改变了梨柄,另一个保留了它。

00:07:54

The model can assign probability to both; our contract can reject one.

模型可以给两者都分配概率,而我们的约定可以拒绝其中一个。

00:08:00

In practice, acceptance may involve human inspection, image comparisons, or explicit production constraints.

实际上,验收可能依赖人工检查、图像对比,或明确的制作约束。

SLIDE 013 · 00:08:07 · Part 1 / Image editing

A smaller workspace: image latents

00:08:07

The RGB image has 1024 by 1024 pixels and three channels: just over three million scalar values.

RGB 图像有 1024 乘 1024 个像素,每个像素有三个通道,总共略多于三百万个标量值。

00:08:15

With spatial reduction by sixteen and sixty-four latent channels, the representation becomes 64 by 64 by 64: 262,144 values.

空间尺寸缩小十六倍、潜在通道数为六十四时,表示变成 64 乘 64 乘 64,即 262,144 个数值。

00:08:29

Notice the trap.

注意这里的陷阱。

00:08:31

Sixteen times smaller along each spatial axis does not mean sixteen times fewer values overall, because the number of channels changes too.

每个空间轴缩小十六倍,并不意味着总数值量减少十六倍,因为通道数也变了。

00:08:39

Here the scalar count falls by a factor of twelve.

在这个例子中,标量数量减少了十二倍。

00:08:43

That is specific to this configuration.

这个比例只适用于这一具体配置。

00:08:45

The model works in a learned representation, and its decoder must recover visible detail.

模型在学习得到的表示中工作,解码器必须从中恢复可见细节。

00:08:52

Our next question is what that compression already loses before we request any edit.

接下来要问的是:在我们提出任何编辑要求之前,压缩已经丢失了什么?

SLIDE 014 · 00:08:57 · Part 1 / Image editing

Is the detail lost before the edit?

00:08:57

Suppose the small exhibition caption is already blurry after an edit.

假设编辑后,展览的小字号说明文字变模糊了。

00:09:02

You might spend an hour rewriting the prompt to say, preserve every letter.

你可能花一个小时重写提示词,强调要保留每一个字母。

00:09:08

There is a quicker diagnostic: encode the original image and decode it again without requesting a change.

有一个更快的诊断方法:不要求任何改变,先把原图编码,再解码回来。

00:09:15

If the letters are already damaged in that round trip, part of the problem lies in representation and reconstruction.

如果这次往返已经损坏了文字,那么问题的一部分就在表示和重建环节。

00:09:23

A more emphatic instruction cannot restore a detail the editing path never retained faithfully.

再强硬的指令,也无法恢复编辑流程从未忠实保留的细节。

00:09:29

Look at fine text, thin edges, and small textures first.

首先检查小字、细线边缘和微小纹理。

00:09:33

If reconstruction is sound but the edit damages them, investigate the editing stage.

如果重建没有问题,而编辑损坏了它们,就调查编辑阶段。

00:09:39

It also explains why exact exhibition typography deserves its own editable layer.

这也说明,要求精确的展览文字排版,应当保留在独立的可编辑图层中。

SLIDE 015 · 00:09:45 · Part 1 / Image editing

Why token count matters

00:09:45

Now consider the number of tokens a transformer processes.

现在考虑 Transformer 需要处理多少个 token。

00:09:49

In our toy example, one token per position on a 64 by 64 grid gives 4,096 tokens.

在这个简化例子里,64 乘 64 的网格上,每个位置对应一个 token,共有 4,096 个。

00:09:58

Grouping each two-by-two region gives 1,024 tokens, a reduction by four.

把每个二乘二区域组合起来,就得到 1,024 个 token,数量减少四倍。

00:10:06

Dense self-attention compares token pairs.

稠密自注意力会比较 token 对。

00:10:09

Squaring the two counts gives 16,777,216 and 1,048,576 pair scores.

将两个数量分别平方,得到 16,777,216 和 1,048,576 个配对分数。

00:10:21

That is a reduction by sixteen.

数量减少了十六倍。

00:10:24

This is why representation choices matter so much for cost.

这就是为什么表示方式会如此显著地影响成本。

00:10:28

It is not a claim that the whole system runs sixteen times faster.

但这不等于整个系统的运行速度会提高十六倍。

00:10:32

Text tokens, reference tokens, other layers, and implementation details also matter.

文本 token、参考图 token、其他网络层和实现细节,也会影响结果。

SLIDE 016 · 00:10:38 · Part 1 / Image editing

Double the resolution: count the cost

00:10:38

Before reading the second number, make a prediction.

看第二个数字之前,先预测一下。

00:10:42

We double the image width and double its height, keeping the patch scheme fixed.

我们把图像宽度和高度都加倍,同时保持分块方式不变。

00:10:47

How many image tokens do we get?

图像 token 会变成多少?

00:10:49

Four times as many.

原来的四倍。

00:10:51

In dense self-attention, each token can compare with every other token.

在稠密自注意力中,每个 token 都可以与其他所有 token 比较。

00:10:56

Four times four gives sixteen times as many pair scores.

四乘四,配对分数的数量就变成十六倍。

00:10:59

This calculation describes the image-token attention component, not a promise that the entire system becomes sixteen times slower.

这个计算只描述图像 token 的注意力部分,并不保证整个系统会慢十六倍。

00:11:08

Other operations and implementations matter.

其他操作和实现方式同样重要。

00:11:12

The useful habit is to ask what grew: pixels, tokens, comparisons, or measured runtime.

一个有用的习惯是追问:增加的究竟是像素、token、比较次数,还是实测运行时间?

00:11:19

They are related quantities, but they are not interchangeable.

这些量相互关联,但不能互换。

SLIDE 017 · 00:11:23 · Part 1 / Image editing

Inside a current image editor

00:11:23

Let us walk through this architecture from the input side.

让我们从输入端走一遍这个架构。

00:11:27

The source image contributes semantic information through a vision-language component and visual information through a VAE.

原图通过视觉语言组件提供语义信息,并通过 VAE 提供视觉信息。

00:11:35

The target stream begins with an evolving noisy representation.

目标分支从一个不断演化的带噪表示开始。

00:11:39

The transformer processes that stream together with the conditions.

Transformer 将这个分支与条件信息一起处理。

00:11:44

During training, we possess an example of the desired edited target, so we can measure what the model should learn.

训练时,我们拥有所需编辑结果的样本,因此可以衡量模型应该学习什么。

00:11:51

During inference, we have the source and request but not the desired target.

推理时,我们有原图和要求,却没有所需的目标图像。

00:11:55

The model must generate it.

模型必须把它生成出来。

00:11:57

That difference is essential.

这个区别至关重要。

00:12:00

The architecture does not quietly receive the answer when you ask it to edit your photograph.

当你让模型编辑照片时,架构并没有偷偷收到答案。

SLIDE 018 · 00:12:06 · Part 1 / Image editing

The missing target at inference

00:12:06

There is a piece of information on the left that we do not possess on the right: the desired edited image.

左边有一项信息,是右边没有的:期望得到的编辑图像。

00:12:12

During supervised training, a source, an instruction, and a target show the model what a successful change looks like.

在监督训练中,原图、指令和目标图一起,向模型展示成功的变化是什么样子。

00:12:20

During use, we provide the request because that target does not yet exist.

使用时,我们之所以提出要求,正是因为目标图像还不存在。

00:12:26

This distinction explains a surprisingly common confusion.

这个区别解释了一个相当常见的混淆。

00:12:29

Showing a reference to an editor is not automatically teaching new weights.

给编辑器看一张参考图,并不自动意味着学习新的权重。

00:12:34

It may simply be conditioning this particular generation.

它可能只是在为这一次生成提供条件。

00:12:38

Adaptation, which we will reach with LoRA, changes trainable parameters.

适配则会改变可训练参数,讲到 LoRA 时我们会展开。

00:12:43

It is available when we construct the learning signal, and absent when we ask the trained model to make our ceramic pear.

构造学习信号时,目标图像是已知的;要求训练好的模型生成陶瓷梨时,它则是未知的。

SLIDE 019 · 00:12:50 · Part 1 / Image editing

Two representations of the source

00:12:50

Think of these two representations using our pear.

用这只梨来理解这两种表示。

00:12:54

Semantic features help with questions such as: which object is the pear, what does ceramic mean, and what does 'it' refer to in the instruction?

语义特征帮助回答:哪个物体是梨,陶瓷是什么意思,以及指令中的“它”指什么?

00:13:05

Visual latents provide appearance information, including shapes and textures that help keep the particular source recognizable.

视觉潜变量提供外观信息,包括形状和纹理,帮助保留原图的独特特征。

00:13:13

This is a conceptual distinction, not a claim that the branches have perfectly isolated responsibilities.

这是一种概念上的区分,并不是说两个分支的职责完全隔离。

00:13:20

Learned representations can overlap in what they encode.

学习到的表示所编码的信息可能重叠。

00:13:24

If an edit fails, ask whether the model misunderstood the request or lost important visual information.

编辑失败时,要问:模型是误解了要求,还是丢失了重要的视觉信息?

SLIDE 020 · 00:13:31 · Part 1 / Image editing

OpenSubject: identity across new scenes

00:13:31

How do you teach a system the difference between a pear and this particular pear?

如何教会系统区分“一只梨”和“这只特定的梨”?

00:13:36

That question connects to OpenSubject, research I coauthored with Yexin Liu and our collaborators.

这个问题与 OpenSubject 有关,这是我与Yexin Liu及其他合作者共同完成的研究。

00:13:43

A video offers something valuable: repeated observations of a subject as the view changes.

视频提供了一种宝贵资源:随着视角变化,对同一主体进行反复观察。

00:13:49

The shared identity is a learning opportunity.

共同的身份特征提供了学习机会。

00:13:53

Follow the main path in the figure.

请沿着图中的主路径看。

00:13:55

We curate clips, verify subjects across frames, and select diverse pairs.

我们筛选视频片段,跨帧验证主体,并挑选具有多样性的图像对。

00:14:01

Inpainting or outpainting helps synthesize reference inputs, followed by verification.

通过图像修补或扩图来合成参考输入,再进行验证。

00:14:08

The corpus contains 2.5 million samples; the contribution is training data and a benchmark.

数据集包含二百五十万个样本;这项工作的贡献是训练数据和评测基准。

00:14:15

For our exhibition, imagine the same sculpture photographed in several rooms.

对于我们的展览,可以想象在不同房间拍摄同一件雕塑。

00:14:20

We want freedom to change the setting without losing its distinctive features.

我们希望自由改变环境,同时不丢失它的独特特征。

00:14:24

That is a classroom application of the identity problem, rather than a claim that our fictional pear was tested in the paper.

这是身份保持问题的课堂应用,并不意味着论文测试过这只虚构的梨。

SLIDE 021 · 00:14:32 · Part 1 / Image editing

Flow matching: the training target

00:14:32

The editor needs a way to learn how to move from noise toward an image.

编辑器需要学习如何从噪声走向图像。

00:14:36

In this simple flow-matching setup, we construct an intermediate latent by mixing noise, epsilon, with a known target latent, z one.

在这个简单的流匹配设定中,我们将噪声 epsilon 与已知目标潜变量 z₁ 混合,构造中间潜变量。

00:14:47

The mixing variable t tells us where we are along that training path.

混合变量 t 告诉我们处于这条训练路径的什么位置。

00:14:52

For this straight interpolation, the target velocity is the target latent minus the noise.

对于这种直线插值,目标速度等于目标潜变量减去噪声。

00:14:57

The model sees an intermediate state and learns to predict that direction, given the source and instruction.

模型看到中间状态,并在原图和指令的条件下,学习预测这个方向。

00:15:04

There is also an important clock distinction: t here is the generative process's time.

还要区分两种时钟:这里的 t 是生成过程中的时间。

00:15:09

Later, our video will have another time axis—the seconds that pass in the scene.

稍后的视频还有另一条时间轴:场景中实际流逝的秒数。

SLIDE 022 · 00:15:15 · Part 1 / Image editing

The interpolation endpoints

00:15:15

Before trusting an equation, check its easiest cases.

相信一个方程之前,先检查最简单的情况。

00:15:19

At t equals zero, the coefficient on the target becomes zero, so we recover the noise.

当 t 等于零时,目标的系数为零,所以我们得到噪声。

00:15:26

At t equals one, the coefficient on noise vanishes, so we recover the target latent.

当 t 等于一时,噪声的系数消失,所以我们得到目标潜变量。

00:15:34

Along this simple straight path, the difference between target and noise gives the direction we train the model to predict.

沿着这条简单直线路径,目标减去噪声,就是我们训练模型预测的方向。

00:15:41

You do not need to imagine a recognizable half-finished image at every intermediate point.

不必想象每个中间点都对应一张能辨认的半成品图像。

00:15:47

The calculation takes place in a learned representation, and this path is a teaching construction rather than the only possible design.

计算发生在学习得到的表示中;这条路径用于教学,并不是唯一可能的设计。

SLIDE 023 · 00:15:57 · Part 1 / Image editing

Predict the next sampling step

00:15:57

Here is the smallest possible sampling calculation.

这是一个最小的采样计算例子。

00:16:01

One coordinate currently equals 0.20.

当前某个坐标值是 0.20。

00:16:04

Its predicted velocity is 0.60, and the next step has size 0.10.

预测速度为 0.60,下一步的步长为 0.10。

00:16:12

Before looking at the result, say what operation should happen.

看结果之前,先说说应该做什么运算。

00:16:17

We add a small amount of motion to the current state: 0.20 plus 0.10 times 0.60, giving 0.26.

在当前状态上加一小步:0.20 加上 0.10 乘 0.60,得到 0.26。

00:16:27

In a real latent, many coordinates update together.

在真实的潜变量中,很多坐标会同时更新。

00:16:31

These numbers are invented to show an Euler step, and practical samplers can use more elaborate solvers.

这些数字是为了演示欧拉步而编造的,实际采样器可以采用更复杂的求解器。

00:16:37

But the sequence is now less mysterious: predict a direction, take a step, repeat, then decode.

但这个过程现在不再那么神秘:预测方向,迈出一步,重复,然后解码。

SLIDE 024 · 00:16:45 · Part 1 / Image editing

A better solver cannot clarify the brief

00:16:45

Imagine following a beautifully accurate set of directions to the wrong gallery.

想象你沿着极其精确的路线指引,却走向了错误的画廊。

00:16:50

Taking smaller steps will not repair the destination.

把步子迈小一点,并不能修正目的地。

00:16:54

The same distinction helps with generative sampling.

同样的区分也有助于理解生成采样。

00:16:57

A better numerical solver can follow a learned field more accurately.

更好的数值求解器,可以更准确地沿着学到的向量场前进。

00:17:02

That alone does not settle an ambiguous editing request.

但仅凭这一点,无法解决含糊的编辑要求。

00:17:06

If the pear's boundary is unstable across sampling settings, numerical or model behavior may matter.

如果不同采样设置下梨的边界不稳定,数值计算或模型行为可能有影响。

00:17:13

If we never specified whether the glass reflection should become ceramic, we also have a brief problem.

如果我们从未说明玻璃倒影是否应变成陶瓷倒影,那么创作说明本身也有问题。

00:17:20

Ask whether the system struggled to realize a clear instruction, or whether we have not agreed on what success means.

要问:系统难以实现一个清楚的指令,还是我们尚未就成功的含义达成一致?

00:17:27

Those situations call for different next actions.

这两种情况需要不同的下一步行动。

SLIDE 025 · 00:17:31 · Part 1 / Image editing

Where edit supervision comes from

00:17:31

Where does the ability to follow an edit instruction come from?

遵循编辑指令的能力从何而来?

00:17:35

A training example can contain a source image, an instruction, and the corresponding edited target.

一个训练样本可以包含原图、指令,以及对应的编辑目标图像。

00:17:42

InstructPix2Pix is an early example that used synthetic editing data to teach this relationship.

InstructPix2Pix 是一个较早的例子,它用合成编辑数据来学习这种关系。

00:17:48

Suppose a training pair says 'make the object ceramic,' but its target also moves the camera and replaces the background.

假设某个训练对的指令是“把物体变成陶瓷”,但目标图还移动了相机并替换了背景。

00:17:56

The learning signal no longer cleanly identifies the requested transformation.

这时,学习信号就不再能清晰地指向所要求的变换。

00:18:01

The model may learn unwanted associations.

模型可能学到不希望出现的关联。

00:18:05

This gives us an important data-quality question: did the target accomplish the requested change while preserving what should remain?

因此,一个重要的数据质量问题是:目标图是否完成了要求的变化,同时保留了应该保留的内容?

00:18:14

Attractive targets alone are not enough to teach dependable editing.

仅有好看的目标图,还不足以教会可靠的编辑。

SLIDE 026 · 00:18:18 · Part 1 / Image editing

What a mismatched training pair teaches

00:18:18

A training pair is a lesson, and lessons can accidentally teach the wrong thing.

一个训练对就是一堂课,而课程可能无意中教错东西。

00:18:23

Imagine a source glass pear and a target ceramic pear that also has a different background.

想象原图是一只玻璃梨,目标图是一只陶瓷梨,但背景也变了。

00:18:29

The written request says only to change the material.

文字要求却只说改变材质。

00:18:32

Which visible changes should the model associate with that request?

模型应当把哪些可见变化与这个要求联系起来?

00:18:36

If this mismatch recurs in the training data, the model can learn an unwanted association between a material change and a scene change.

如果训练数据中反复出现这种不匹配,模型可能把材质变化与场景变化错误地联系起来。

00:18:44

Compare the instruction with the actual difference between source and target.

请将指令与原图、目标图之间的实际差异进行比较。

00:18:49

Ask what supervision is rewarding, including changes nobody meant to label.

要问监督信号在奖励什么,包括那些没人打算标注的变化。

00:18:54

The preservation contract begins in the examples used to teach the editor.

保留约定,应当从用来教会编辑器的样本开始。

SLIDE 027 · 00:18:59 · Part 1 / Image editing

Classifier-free guidance (CFG)

00:18:59

Sometimes a conditioned prediction moves in the right direction but too weakly.

有时,条件预测的方向是对的,但力度不足。

00:19:04

Classifier-free guidance uses the difference between a conditioned prediction and a baseline prediction to steer the result.

无分类器引导利用条件预测与基线预测之间的差异,来引导结果。

00:19:12

Read the formula as baseline, plus a scaled change in direction.

可以把公式读成:基线,加上经过缩放的方向变化。

00:19:16

When s equals one, the baseline terms cancel and we recover the conditioned prediction.

当 s 等于一时,基线项抵消,我们得到条件预测。

00:19:22

Above one, we extrapolate beyond it.

大于一时,我们会外推到条件预测之外。

00:19:25

That can strengthen the requested change, but it can strengthen errors as well.

这可以强化所要求的变化,但也可能强化错误。

00:19:30

In an editing system, the baseline may still retain image information; baseline does not always mean no information at all.

在编辑系统中,基线可能仍保留图像信息;基线并不总意味着完全没有信息。

00:19:39

Think of guidance as a specific operation on predictions, rather than a universal quality dial.

把引导理解为对预测值的一种具体运算,而不是万能的质量旋钮。

SLIDE 028 · 00:19:46 · Part 1 / Image editing

Guidance goes beyond the prediction

00:19:46

The baseline predicts 0.2 and the conditioned version predicts 0.5.

基线预测是 0.2,条件预测是 0.5。

00:19:52

With guidance scale two, do we get a value somewhere between them?

引导强度设为二,结果会落在两者之间吗?

00:19:57

No.

不会。

00:19:58

We take 0.2 plus twice the difference of 0.3, which gives 0.8.

我们用 0.2 加上两倍的差值 0.3,得到 0.8。

00:20:05

We have moved beyond the conditioned prediction.

我们已经越过了条件预测。

00:20:08

If the useful direction includes a slight mistake, we can amplify both.

如果有效方向中包含一点错误,那么两者都可能被放大。

00:20:13

For the pear, the new material might become clearer while the stem or boundary becomes less faithful.

对于这只梨,新材质可能更明显,但柄或边界可能变得不那么忠实。

00:20:19

Compare the requested change and preservation separately.

请分别比较所要求的变化和保留效果。

00:20:24

One slider can move those judgments in opposite directions; a single overall impression can hide the tradeoff.

同一个滑块可能让两项评价朝相反方向变化;单一的总体印象会掩盖这种取舍。

SLIDE 029 · 00:20:31 · Part 1 / Image editing

Different controls carry different information

00:20:31

Different controls communicate different kinds of information.

不同控制手段传达不同类型的信息。

00:20:35

Text describes a requested change.

文本描述所要求的变化。

00:20:38

A mask specifies a region.

蒙版指定区域。

00:20:41

A spatial condition can describe pose, depth, or edges.

空间条件可以描述姿态、深度或边缘。

00:20:46

An image reference can supply appearance information that would be difficult to express precisely in words.

参考图可以提供难以用语言精确表达的外观信息。

00:20:52

These are not interchangeable knobs, and every product does not expose all of them.

这些控制并不能互相替代,也不是每款产品都提供全部选项。

00:20:57

ControlNet and IP-Adapter are examples of distinct architectural approaches.

ControlNet 和 IP-Adapter 就是两种不同架构思路的例子。

00:21:01

If exact untouched pixels are essential, a generation mask alone may be insufficient; explicit copying or compositing can enforce that requirement.

如果必须精确保留未编辑像素,仅有生成蒙版可能不够;明确复制像素或合成,可以强制实现这一要求。

00:21:12

Choose a control by asking what information is missing, then check whether the resulting image actually respected it.

选择控制手段时,先问缺少什么信息,再检查结果是否真正遵循了它。

00:21:18

A mask indicates a permitted region; it does not by itself solve seams or the reflection of a changed material.

蒙版指出允许编辑的区域,但它本身不能解决接缝,也不能处理材质改变后的倒影。

00:21:25

The next slide names an architecture that learns to use a spatial condition.

下一页介绍一种学习使用空间条件的架构。

SLIDE 030 · 00:21:30 · Part 1 / Image editing

ControlNet: add spatial conditioning

00:21:30

Suppose your sentence is understood, but the pear keeps changing shape.

假设模型理解了你的话,但梨的形状总在变化。

00:21:34

You can describe its outline with more adjectives, or supply spatial evidence.

你可以用更多形容词描述轮廓,也可以直接提供空间证据。

00:21:39

ControlNet gives an edge map, depth map, or pose a learned route into a compatible diffusion model.

ControlNet 为边缘图、深度图或姿态信息,提供了一条进入兼容扩散模型的可学习路径。

00:21:46

Trace the two paths in this original architecture.

请追踪原始架构中的两条路径。

00:21:49

The pretrained path stays frozen.

预训练主路径保持冻结。

00:21:51

A trainable copy processes the added condition, and zero-initialized one-by-one convolutions connect its features to the main path.

一个可训练副本处理新增条件,零初始化的一乘一卷积把它的特征连接到主路径。

00:22:00

Initially those connections contribute zero; training learns their contribution.

初始时,这些连接的贡献为零;训练会学出它们应有的贡献。

00:22:06

The copied branch itself starts from pretrained weights, not all zeros.

副本分支本身从预训练权重开始,而不是全部从零开始。

00:22:11

Now return to our request.

现在回到我们的要求。

00:22:13

Edges can help specify the outline.

边缘可以帮助指定轮廓。

00:22:16

They cannot, by themselves, tell us whether the reflected ceramic looks convincing or the blue stem retains its identity.

但仅凭边缘,无法判断陶瓷倒影是否可信,也无法保证蓝色梨柄保留身份特征。

SLIDE 031 · 00:22:24 · Part 1 / Image editing

Multiple references need distinct identities

00:22:24

With several references, the system must know which reference contributes which information.

当有多张参考图时,系统必须知道每张图分别提供什么信息。

00:22:30

Imagine asking for the sculpture from image one and the atmosphere from image two.

想象你要求使用图一的雕塑,以及图二的氛围。

00:22:37

If the references become confused, you may get the wrong object with the right lighting.

如果参考图的角色混淆,可能会得到光照正确、物体却错误的结果。

00:22:43

The illustrated research approach uses separators and image-index information to distinguish inputs.

图示研究方法使用分隔符和图像索引信息来区分输入。

00:22:48

In an art brief, name the role of each image rather than presenting a pile of vaguely related inspiration.

在艺术创作说明中,要说明每张图的角色,而不是堆放一批关系模糊的灵感图。

00:22:55

Then inspect for leakage: did a composition reference accidentally replace the subject?

然后检查信息是否串用:构图参考是否意外替换了主体?

00:23:02

This figure describes one proposed mechanism, not a universal design used by every editor.

这张图描述的是一种提出的机制,并不是所有编辑器通用的设计。

SLIDE 032 · 00:23:08 · Part 1 / Image editing

A reference-role failure is visible in the output

00:23:08

Look at these published multi-image examples with reference roles in mind.

请带着“参考图角色”的概念,观察这些已发表的多图示例。

00:23:13

Before judging the output, identify what each input was supposed to contribute.

评价输出之前,先确定每个输入原本应该贡献什么。

00:23:18

Then trace the subject and the requested transformation into the result.

然后在结果中追踪主体和所要求的变换。

00:23:23

A reference-role failure can look superficially attractive.

参考图角色混淆的结果,表面上也可能很好看。

00:23:27

The system may borrow the wrong object's appearance or import a background that was never intended.

系统可能借用了错误物体的外观,或引入了原本不需要的背景。

00:23:33

We are examining qualitative examples from a particular paper, not conducting a broad comparison between products.

我们正在查看某篇论文的定性示例,并不是对产品做全面比较。

00:23:40

The figure helps us practice an inspection method: follow the intended contribution of each input and look for unintended transfers between them.

这张图帮助我们练习一种检查方法:追踪每个输入的预期贡献,寻找意外的信息转移。

SLIDE 033 · 00:23:49 · Part 1 / Image editing

Attention as weighted information gathering

00:23:49

Attention lets a token gather information from other tokens with different weights.

注意力让一个 token 以不同权重,从其他 token 收集信息。

00:23:55

Our simplified scalar example gives weight 0.8 to a value of 0.9, and weight 0.2 to a value of 0.1.

在这个简化的标量例子中,数值 0.9 的权重为 0.8,数值 0.1 的权重为 0.2。

00:24:05

The weighted result is 0.74.

加权结果是 0.74。

00:24:08

The real mechanism operates on learned vectors, so this is arithmetic intuition rather than a literal description of artistic decision-making.

真实机制处理的是学习到的向量,因此这里只是在建立算术直觉,并非字面描述艺术决策。

00:24:17

The important structure is selective combination: every available source need not contribute equally.

关键结构是选择性组合:不是每个可用来源都必须贡献同样多的信息。

00:24:23

In a multi-reference edit, the system also needs to distinguish where information came from.

在多参考图编辑中,系统还需要区分信息来自哪里。

00:24:29

Otherwise, gathering information successfully can still produce the wrong mixture of subject identity and visual style.

否则,即使成功收集了信息,也可能错误地混合主体身份与视觉风格。

SLIDE 034 · 00:24:36 · Part 1 / Image editing

Softmax weights sum to one

00:24:36

The weights in this simplified attention example sum to one.

在这个简化注意力例子中,权重之和为一。

00:24:41

That makes the result a weighted combination of the values.

因此,结果是各个数值的加权组合。

00:24:45

A larger weight makes the associated value contribute more to this particular calculation.

权重越大,对应数值在这次计算中的贡献就越大。

00:24:51

Now be cautious about jumping from the calculation to an explanation of the final image.

但要谨慎,不能直接从这个计算跳到对最终图像的解释。

00:24:57

Real networks contain many layers, heads, and transformations.

真实网络包含许多层、注意力头和变换。

00:25:01

A high weight at one point may not tell us why a visible feature ultimately appeared.

某处权重很高,未必能解释某个可见特征最终为何出现。

00:25:06

For practical reference editing, the useful question remains whether the intended information survives in the output, regardless of how compelling an attention visualization looks.

在实际参考图编辑中,关键仍是预期信息是否出现在输出中,无论注意力可视化看起来多么有说服力。

SLIDE 035 · 00:25:18 · Part 1 / Image editing

LoRA: low-rank adaptation

00:25:18

What if a useful concept must recur across many requests?

如果一个有用的概念需要在许多请求中反复出现,怎么办?

00:25:22

A reference can condition a generation, while LoRA adapts selected learned weights using a low-rank update.

参考图可以为一次生成提供条件;LoRA 则通过低秩更新,调整选定的模型权重。

00:25:29

The update is a product of two smaller matrices.

更新量是两个较小矩阵的乘积。

00:25:33

Instead of learning every entry of a large update matrix independently, we express that update as a product of two smaller matrices.

我们不独立学习大型更新矩阵的每个元素,而是用两个较小矩阵的乘积来表示更新。

00:25:41

For a 4096 by 4096 weight matrix and rank sixteen, the full matrix has 16,777,216 entries.

对于一个 4096 乘 4096、秩设为十六的例子,完整矩阵有 16,777,216 个元素。

00:25:52

The two factors together have 131,072, a factor of 128 fewer for this update.

两个因子合计只有 131,072 个元素,这次更新的参数量减少了一百二十八倍。

00:26:01

That is not a claim that the entire model shrinks by 128.

这并不意味着整个模型缩小了一百二十八倍。

00:26:05

The base weights still exist.

基础权重仍然存在。

00:26:08

Low rank makes adaptation economical, but does not certify that the learned subject survives a new pose.

低秩让适配更经济,但不能证明学到的主体在新姿态下仍能保持一致。

00:26:14

That question requires examples the adaptation did not already see.

这个问题需要用适配时未见过的样例来检验。

SLIDE 036 · 00:26:19 · Part 1 / Image editing

Did it learn the subject or the setting?

00:26:19

An adapted model reproduces your favorite training portrait perfectly.

适配后的模型完美复现了你最喜欢的训练肖像。

00:26:24

Is that enough to use the character in a new film?

这就足以让这个角色出演新电影了吗?

00:26:27

Consider what else it may have learned: the familiar camera angle, the background, even the lighting that always accompanied the subject.

想一想它还可能学到了什么:熟悉的机位、背景,甚至总与主体一起出现的光照。

00:26:35

A held-out check deliberately changes those circumstances.

留出测试会有意改变这些条件。

00:26:39

Ask for a new view or a different setting and inspect the distinctive features.

要求一个新视角或不同环境,并检查独特特征。

00:26:45

We want a reusable concept, rather than a narrow ability to reproduce familiar combinations.

我们需要的是可复用的概念,而不只是复现熟悉组合的狭窄能力。

00:26:51

For the pear, keep the blue stem recognizable while moving the exhibition outdoors.

对于这只梨,要在把展览移到户外时,仍保留可辨认的蓝色柄。

00:26:57

If identity collapses there, more praise for the training examples will not solve the problem.

如果身份特征在那里崩溃,再称赞训练样例也解决不了问题。

SLIDE 037 · 00:27:03 · Part 1 / Image editing

What does the reward favor?

00:27:03

A reward tells a learning process what kinds of outputs to favor.

奖励告诉学习过程应当偏好哪些输出。

00:27:07

The figure shows a recent approach that separates task-specific reward training and then distills what is learned.

图中是一种近期方法:分别训练任务特定的奖励,再蒸馏学到的能力。

00:27:14

Look at the categories: editing quality involves more than a single judgment of attractiveness.

看这些类别:编辑质量不只是单一的美观判断。

00:27:21

Now imagine an artwork whose purpose is to feel awkward or disturbing.

现在想象一件艺术作品,其目的就是让人感到别扭或不安。

00:27:25

A generic preference for polished images could work against that purpose.

对精致图像的普遍偏好,可能反而违背这个目的。

00:27:29

Human preference signals are useful, but they do not define artistic merit for every project.

人类偏好信号很有用,但不能为所有项目定义艺术价值。

00:27:35

This is the bridge to the next section: technical optimization can help produce candidates, while an artist still needs to decide which properties serve the work.

这正是通向下一部分的桥梁:技术优化可以帮助生成候选作品,而艺术家仍要决定哪些属性服务于作品。

SLIDE 038 · 00:27:46 · Part 1 / Image editing

One reward cannot stand in for every intention

00:27:46

These edit examples invite a question about evaluation.

这些编辑示例引出了一个评价问题。

00:27:50

Which properties would you reward separately?

你会分别奖励哪些属性?

00:27:53

You might ask whether the instruction was carried out, whether identity survived, and whether the image is visually convincing.

可以问:指令是否完成、身份是否保留,以及图像在视觉上是否可信。

00:28:01

Those judgments can disagree.

这些判断可能彼此冲突。

00:28:04

A highly polished output might erase an awkward feature that is central to the artwork.

高度精致的输出,可能抹掉作品中至关重要的别扭特征。

00:28:09

An unusual composition might serve the brief while attracting a lower generic preference score.

不寻常的构图可能符合创作要求,却得到较低的通用偏好分数。

00:28:15

We should understand what a reward encourages before treating a high score as artistic approval.

在把高分视为艺术认可之前,我们应当理解奖励究竟鼓励什么。

00:28:20

The published examples illustrate the research setting; our classroom task is to articulate the intention against which we would judge a particular result.

已发表示例展示的是研究情境;课堂任务则是明确创作意图,并据此判断具体结果。

SLIDE 039 · 00:28:30 · Part 1 / Image editing

Attractiveness and fidelity are separate tests

00:28:30

Which failure would you notice first?

你会先注意到哪种失败?

00:28:33

On one side, the output is beautiful, but it has quietly replaced our particular sculpture with a generic decorative pear.

一边的结果很美,却悄悄把我们的特定雕塑换成了普通的装饰梨。

00:28:40

On the other, the right sculpture is present, but its reflection still behaves like the old material.

另一边保留了正确雕塑,但倒影仍像原来的材质。

00:28:46

The first may win a quick aesthetic vote.

第一种可能赢得快速的审美投票。

00:28:49

The second may pass a checklist that only asks whether the requested object changed.

第二种可能通过只检查“目标物体是否改变”的清单。

00:28:53

Neither satisfies the full brief.

两者都没有满足完整要求。

00:28:56

Has the editor lost identity, broken the scene's optical logic, or missed the work's intention?

编辑器是丢失了身份、破坏了场景光学逻辑,还是偏离了作品意图?

00:29:02

Naming the mismatch gives us a next action.

说清不匹配之处,才能决定下一步。

00:29:05

Saying only that an output looks a little wrong leaves the diagnosis unfinished.

只说输出“看起来有点不对”,诊断还没有完成。

SLIDE 040 · 00:29:11 · Part 1 / Image editing

A useful diagnosis changes the next action

00:29:11

A diagnosis becomes useful when it changes the next action.

只有能改变下一步行动的诊断,才有用。

00:29:15

If the wrong object changes, first suspect an ambiguous reference or region specification.

如果改错了物体,首先怀疑参考对象或区域描述有歧义。

00:29:22

If the correct object becomes another instance, appearance grounding may be the issue.

如果改对了物体,却变成另一个个体,问题可能在外观信息的约束。

00:29:28

If the material looks disconnected from its reflection, the dependent change may be missing.

如果材质与倒影脱节,可能缺少必要的连带变化。

00:29:34

These are candidate explanations, not automatic conclusions.

这些只是候选解释,不是自动成立的结论。

00:29:38

Choose a small revision that tests one of them.

选择一个小调整,检验其中一种解释。

00:29:40

If it does not help, reconsider the hypothesis.

如果没有帮助,就重新考虑假设。

00:29:44

This is more informative than changing the prompt, seed, reference, and model all at once, because then a successful result would leave us unsure which intervention mattered.

这比同时改提示词、随机种子、参考图和模型更有信息量;否则,即使成功,也不知道哪个改动起了作用。

SLIDE 041 · 00:29:54 · Part 1 / Image editing

Discuss: the ceramic edit

00:29:54

We now have several ways to communicate an edit.

现在,我们有了几种表达编辑要求的方法。

00:29:58

Would you preserve the reflection or let it change, and why?

你会保留倒影,还是允许它变化?为什么?

00:30:02

When would a mask help more than a longer prompt?

什么时候蒙版比更长的提示词更有帮助?

00:30:06

What would edges or depth control still leave uncertain?

边缘或深度控制仍会留下哪些不确定性?

00:30:10

Discuss these three questions using the pear already on screen.

请用屏幕上的这只梨,讨论这三个问题。

00:30:14

For each control you favor, name the failure it addresses and something it cannot settle.

对于你支持的每种控制手段,指出它解决什么失败,以及它不能决定什么。

00:30:19

Pause the video here.

请在这里暂停视频。

00:30:21

The next slide offers one possible contract; compare it with your reasoning rather than treating it as the only artistic answer.

下一页给出一种可能的约定;请与你的推理比较,不要把它当成唯一的艺术答案。

SLIDE 042 · 00:30:33 · Part 1 / Image editing

A worked ceramic-edit contract

00:30:33

Here is one defensible answer to the ceramic puzzle.

对于陶瓷难题,这是一种有理有据的答案。

00:30:37

Preserve the blue stem and recognizable silhouette.

保留蓝色柄和可辨认的轮廓。

00:30:40

Allow transmission, highlights, and the reflection to change with the new material.

允许透光、高光和倒影随着新材质改变。

00:30:47

Then inspect the boundary and water, because that is where the request reaches beyond the object.

然后检查边界和水面,因为要求的影响会在那里超出物体本身。

00:30:53

Notice the wording: allow a justified consequence, rather than permit arbitrary drift.

注意措辞:允许有理由的连带变化,而不是允许任意漂移。

00:30:58

We have not given the model permission to redesign the gallery.

我们没有授权模型重新设计画廊。

00:31:02

We have named the changes necessary to make this material transformation coherent.

我们明确指出了,让这次材质变换协调一致所必需的变化。

00:31:07

A mask could help localize work, while compositing might protect a region that truly must remain exact.

蒙版可以帮助限定工作区域;合成则可以保护真正必须完全不变的区域。

00:31:15

The contract tells us how to use those tools.

编辑约定告诉我们如何使用这些工具。

SLIDE 043 · 00:31:18 · Part 1 / Image editing

The brief, the model, and the review

00:31:18

We can now explain our opening puzzle at three levels.

现在,我们可以在三个层面解释开场难题。

00:31:22

The brief decides what should change and what should survive.

创作说明决定什么应改变、什么应保留。

00:31:26

The model uses representations, conditioning, and sampling to propose a result.

模型通过表示、条件控制和采样,提出候选结果。

00:31:31

Our review checks whether that proposal actually satisfies the brief.

我们的审查检验这个候选结果是否真正满足要求。

00:31:36

Keep those levels separate when you diagnose a failure.

诊断失败时,要区分这三个层面。

00:31:39

Stronger guidance will not choose the exhibition's purpose.

更强的引导不会替你选择展览的目的。

00:31:43

A more eloquent artistic statement will not enforce identical pixels.

更动人的艺术陈述不会强制像素完全一致。

00:31:48

A beautiful preview will not prove preservation.

精美预览不会证明保留成功。

00:31:51

The useful skill is connecting the right intervention to the observed problem.

有用的能力,是把正确的干预对应到观察到的问题。

00:31:56

If your partner can tell when to use it and what it leaves uncertain, you have understood more than its name.

如果同伴能说清何时使用它、它还留下什么不确定性,你就不只是记住了它的名字。

SLIDE 044 · 00:32:03 · Part 1 / Image editing

After Part 1: explain the control

00:32:03

We can now explain our controls rather than just name them.

现在,我们能够解释控制手段,而不只是说出名称。

00:32:06

How does ControlNet differ from a mask or a text prompt?

ControlNet 与蒙版或文本提示词有什么区别?

00:32:10

Why can stronger guidance make an edit worse?

为什么更强的引导可能让编辑变差?

00:32:14

What evidence would show that an edit preserved identity?

什么证据能够表明编辑保留了身份特征?

00:32:18

Discuss these three questions.

请讨论这三个问题。

00:32:20

Choose a concrete requirement—a silhouette, an untouched region, or a distinctive stem—and connect your explanation to it.

选一个具体要求——轮廓、未编辑区域,或独特的梨柄——并把解释与它联系起来。

00:32:27

Then consider whether the control guarantees that requirement or only helps express it.

然后考虑:这种控制是保证实现要求,还是仅仅帮助表达要求?

00:32:32

Pause the video here before the break.

休息之前,请在这里暂停视频。

SLIDE 045 · 00:32:39 · Part 1 / Image editing

Break

00:32:39

We will take ten minutes and resume when the class is ready.

我们休息十分钟,等大家准备好再继续。

00:32:43

When you come back, think about a series of images you would recognize as belonging to one artwork.

回来时,请想一组你能认出属于同一件艺术作品的图像。

00:32:48

What makes them belong together: the subject, the palette, the treatment of space, or something else?

是什么让它们属于一体:主体、色彩、空间处理,还是别的因素?

00:32:56

Leave that question open for now.

暂时先保留这个问题。

00:32:59

We will use it to move from individual editing operations to a coherent visual language.

我们将用它,从单次编辑操作转向连贯的视觉语言。

SLIDE 046 · 00:33:08 · Part 1 / Image editing

One exhibition, different visual worlds

00:33:08

Look again at the artwork we have been using as a technical test.

再看看这件一直被我们用来做技术测试的作品。

00:33:12

After the break, it has a different job.

休息之后,它有了不同的任务。

00:33:15

We are no longer asking only whether the edit obeys a request.

我们不再只问编辑是否遵循了要求。

00:33:19

We are asking what a visitor might feel, and which visual decisions create that feeling.

我们要问:观众可能感受到什么,哪些视觉决定创造了这种感受?

00:33:26

Two images can have different surfaces and still belong to one exhibition.

两张图像的表面可以不同,却仍属于同一个展览。

00:33:30

Two images can share a palette and feel unrelated.

两张图像可以使用相同色彩,却感觉毫不相关。

00:33:33

The difference is worth arguing about.

这种差别值得讨论。

00:33:36

In the next section, I will show prepared alternatives for AFTER RAIN.

下一部分,我会展示为《雨后》准备的不同方案。

00:33:41

Choose a direction in your mind, and be ready to explain a visible reason.

请在心里选择一个方向,并准备说出可见的理由。

00:33:46

Your neighbor may choose the other one.

你的邻座可能会选另一个。

SLIDE 047 · 00:33:49 · Part 2 / Art direction

Art direction

00:33:49

Now take the art director’s seat.

现在,请坐到艺术总监的位置。

00:33:51

The model can produce many plausible outputs; we decide which differences matter for AFTER RAIN.

模型能生成许多合理的结果;我们来决定哪些差异对《雨后》重要。

00:33:57

As we compare the prepared images, name the visible relationship that carries the idea.

比较这些准备好的图像时,请指出承载创意的可见关系。

00:34:03

That could be scale, light, or the treatment of space.

它可能是尺度、光线,或空间处理。

00:34:09

Your explanation will be more useful than simply calling an image impressive.

这样的解释,比简单地说图像很惊艳更有用。

SLIDE 048 · 00:34:13 · Part 2 / Art direction

Before Part 2: artistic intention

00:34:13

Before seeing the poster alternatives, consider what you want a visitor to experience.

看海报方案之前,先想想你希望观众经历什么。

00:34:18

What makes an image feel fragile rather than merely attractive?

什么让图像显得脆弱,而不仅仅是好看?

00:34:22

Can very different images belong to the same artwork, and why?

差别很大的图像,能否属于同一件作品?为什么?

00:34:27

Which artistic decisions would you keep for yourself?

哪些艺术决定,你会留给自己?

00:34:31

Discuss these three questions.

请讨论这三个问题。

00:34:33

You might disagree about the effect of empty space or the meaning of a material.

你们可能对留白的效果或材质的含义有不同看法。

00:34:37

Locate the visual evidence behind that disagreement.

请找出分歧背后的视觉证据。

00:34:41

Pause the video here.

请在这里暂停视频。

00:34:43

Keep your starting position in mind when we compare the prepared directions.

比较预先准备的方向时,请记住你的初始立场。

SLIDE 049 · 00:34:51 · Part 2 / Art direction

The exhibition brief

00:34:51

Here is our commission: AFTER RAIN, a fictional exhibition about fragile objects in changing environments.

这是我们的委托:《雨后》,一个关于变化环境中脆弱物体的虚构展览。

00:35:00

Imagine a visitor seeing its poster across a corridor before encountering a five-second moving image inside.

想象观众先在走廊另一端看到海报,再进入展厅看到一段五秒的动态图像。

00:35:07

What should that visitor expect to feel?

你希望这位观众期待怎样的感受?

00:35:10

The prepared examples keep the amber pear and blue stem recognizable.

准备好的示例都保留了可辨认的琥珀色梨和蓝色柄。

00:35:14

We will examine photographic and collage directions, then consider how a moving reveal changes the experience.

我们将研究摄影和拼贴两个方向,再考虑运动中的揭示如何改变体验。

00:35:22

Nobody needs to generate an image during this lecture.

这节课不需要任何人现场生成图像。

00:35:25

Fragile is a useful beginning, but it is not yet an art direction.

“脆弱”是一个有用的起点,但还不算艺术方向。

00:35:30

We need to translate that word into something visible enough to compare and precise enough to revise.

我们需要把这个词转化为足够可见、能够比较,又足够精确、能够修改的东西。

SLIDE 050 · 00:35:36 · Part 2 / Art direction

Making fragility visible

00:35:36

Try replacing fragile with expensive in the brief.

试着把创作说明中的“脆弱”换成“昂贵”。

00:35:40

You might still choose glass, but would you use the same composition?

你可能仍然选择玻璃,但还会用同样的构图吗?

00:35:44

A large centered object and assertive lighting might suggest a luxury product.

居中放大的物体配合强势光照,可能让人联想到奢侈品。

00:35:49

A small object surrounded by quiet space might instead seem exposed.

安静空间中一个小小的物体,则可能显得无所庇护。

00:35:54

Scale can make the pear feel vulnerable.

尺度可以让梨显得易受伤害。

00:35:57

Restrained light can make us look closely.

克制的光线可以引导我们细看。

00:35:59

A reflection can suggest a world less stable than the object itself.

倒影可以暗示一个比物体本身更不稳定的世界。

00:36:04

These are artistic hypotheses to test against an audience's reading, not a formula that makes every image fragile.

这些是需要通过观众解读来检验的艺术假设,不是让所有图像显得脆弱的公式。

00:36:12

Point to the feature doing the work.

请指出真正起作用的特征。

00:36:15

If nobody can locate it in the image, the intention may still be living only in the prompt.

如果没人能在图像中找到它,意图可能仍然只存在于提示词里。

SLIDE 051 · 00:36:21 · Part 2 / Art direction

Hokusai: scale, rhythm, and tension

00:36:21

In Hokusai's Great Wave, look first for Mount Fuji.

看葛饰北斋的《神奈川冲浪里》,先找富士山。

00:36:24

It is small and distant.

它很小,也很远。

00:36:26

Now follow the curves of the wave and the boats beneath it.

现在沿着巨浪的曲线,以及下方的小船看。

00:36:30

Scale and rhythm create a tension we can discuss without turning the artwork into a style label.

尺度和节奏创造了张力;我们可以讨论这种张力,而不是把作品变成一个风格标签。

00:36:36

For AFTER RAIN, we might borrow a relationship: a small, vulnerable form facing a much larger environment.

对《雨后》来说,我们可以借鉴一种关系:一个小而脆弱的形体,面对远大于自身的环境。

00:36:43

We do not need to reproduce the wave or ask for a generic imitation of the artist.

不需要复制巨浪,也不需要要求泛泛地模仿这位艺术家。

00:36:48

The reference becomes useful when we can say which visual decision we are studying.

当我们能够说明正在研究哪个视觉决定时,参考作品才真正有用。

00:36:53

Then we must test its translation.

然后,还必须检验这种转化。

00:36:55

Our quiet flooded gallery has a different subject and emotional register.

我们安静而积水的画廊,有着不同的主体和情感基调。

SLIDE 052 · 00:37:00 · Part 2 / Art direction

Reference as a relationship, not a label

00:37:00

Instead of using a reference as a label, describe a relationship you can observe.

不要把参考作品当作标签,而要描述你能观察到的关系。

00:37:06

You might want a small stable form beneath a large dynamic curve, or a rhythm that moves the eye through the composition.

你可能想要一个巨大动态曲线下的小型稳定形体,或引导视线穿过构图的节奏。

00:37:14

Then translate that relationship into the new subject and brief.

然后,把这种关系转化到新的主体和创作要求中。

00:37:18

The result should be judged as its own artwork, not as a contest to resemble the reference.

应把结果作为独立作品评价,而不是比赛谁更像参考图。

00:37:23

This approach makes references more useful to collaborators because they can understand what you are borrowing.

这种方法让参考资料对合作者更有用,因为他们能理解你借鉴的是什么。

00:37:29

It also makes iteration more focused: if the intended tension is missing, you can revise scale or rhythm rather than vaguely asking for more influence from the source.

它也让迭代更聚焦:如果缺少预期张力,可以调整尺度或节奏,而不是含糊地要求“更受原作影响”。

SLIDE 053 · 00:37:41 · Part 2 / Art direction

Where does your eye arrive?

00:37:41

Before admiring the surface detail, decide where your eye arrives.

欣赏表面细节之前,先判断你的视线最先落在哪里。

00:37:45

Is it the pear, the reflection, or the space waiting above them?

是梨、倒影,还是它们上方留出的空间?

00:37:50

A poster has to organize that first encounter.

海报必须组织这第一次相遇。

00:37:54

If every region is equally busy, the audience has to invent a hierarchy we have not provided.

如果每个区域都同样繁忙,观众就得自行建立我们没有提供的层次。

00:38:00

Here, the open area can make the object feel small and give the title somewhere to live.

这里的开放区域能让物体显得小,也给标题留下位置。

00:38:05

We can ruin the sense of quiet by filling every available gap with decorative detail, even if each addition is attractive on its own.

如果用装饰细节填满每一处空隙,就可能破坏安静感,即使每个新增细节单独看都很漂亮。

00:38:13

When revising this image, I would first ask whether the composition expresses the intended scale and silence.

修改这张图时,我会先问:构图是否表达了预期的尺度感与寂静?

00:38:20

More texture comes later, if the work needs it.

如果作品需要,之后再增加纹理。

SLIDE 054 · 00:38:24 · Part 2 / Art direction

Negative space has a job

00:38:24

Negative space does two jobs in this poster.

留白在这张海报中承担两项任务。

00:38:27

It changes how large or isolated the object feels, and it creates a place for typography.

它改变物体显得多大、多孤立,也为文字排版创造位置。

00:38:34

These are connected design decisions rather than separate finishing steps.

这是相互关联的设计决定,而不是彼此独立的收尾步骤。

00:38:39

If you fill the pale area with dramatic detail, the image may become more visually active but leave no calm route for reading the title.

如果在浅色区域填满戏剧性细节,画面可能更活跃,却没有平静的路径让人阅读标题。

00:38:47

If you reserve too much space, the subject might lose the presence you wanted.

如果留白过多,主体又可能失去你想要的存在感。

00:38:51

Evaluate the space in relation to the finished use.

请结合最终用途来评价空间。

00:38:55

A generative background is part of a larger composition when it must carry text, branding, or other exact elements.

当背景需要承载文字、品牌或其他精确元素时,生成背景只是更大构图的一部分。

SLIDE 055 · 00:39:03 · Part 2 / Art direction

A visual language has rules

00:39:03

The collage direction changes the rules of the world.

拼贴方向改变了这个世界的规则。

00:39:06

Torn edges replace smooth contours.

撕裂边缘取代了平滑轮廓。

00:39:09

Flat layers replace optical depth.

平面层次取代了光学纵深。

00:39:12

Amber and indigo keep a connection to our recurring subject.

琥珀色和靛蓝色,让它仍与反复出现的主体保持联系。

00:39:17

Look at how those decisions affect the pear's apparent weight and vulnerability.

看看这些决定如何影响梨看起来的重量与脆弱感。

00:39:22

If the body looks like paper but the water behaves like a photograph, do we accept that collision?

如果主体像纸,水却像照片,我们是否接受这种碰撞?

00:39:28

We might, if it is deliberate and supports the work.

如果它是有意的,并服务于作品,我们可能接受。

00:39:31

We might reject it as an unresolved mixture.

我们也可能认为它是尚未解决的混杂,因而拒绝。

00:39:35

The answer depends on the intended visual language.

答案取决于预期的视觉语言。

00:39:38

The next comparison asks you to locate precisely where that language holds together or breaks.

下一组比较,请你精确指出:这种语言在哪里成立,又在哪里断裂。

SLIDE 056 · 00:39:44 · Part 2 / Art direction

When a reflection breaks the collage

00:39:44

A photographic reflection inside a paper collage can be a mistake, or the most interesting decision in the image.

纸质拼贴中的摄影式倒影,可能是错误,也可能是整张图最有意思的决定。

00:39:51

We need to know whether the mismatch is an intentional disruption and whether it produces the intended effect.

我们需要知道:这种不匹配是否是有意的打破,以及是否产生了预期效果。

00:39:59

Suppose the exhibition explores unstable memories.

假设展览探索的是不稳定的记忆。

00:40:02

An impossibly photographic reflection could make sense.

一个不可能的摄影式倒影,就可能说得通。

00:40:06

Suppose the brief calls for a coherent world assembled from torn paper.

假设要求是构建一个由撕纸组成的统一世界。

00:40:10

The same reflection might weaken it.

同样的倒影可能削弱它。

00:40:13

This does not mean every accident deserves an explanation after the fact.

这并不意味着每个意外都值得事后找理由。

00:40:18

Make the intention specific, examine what the viewer can actually see, and compare alternatives.

请明确意图,检查观众实际能看见什么,并比较不同方案。

00:40:25

A strong critique can distinguish a productive contradiction from an excuse for an unresolved result.

有力的评论能区分富有成效的矛盾,与为未完成结果寻找的借口。

SLIDE 057 · 00:40:32 · Part 2 / Art direction

Reference roles

00:40:32

A reference becomes easier to use when we assign it a role.

为参考图分配角色后,它就更容易使用。

00:40:36

One image might define the subject, another the composition, and another the material treatment.

一张图可以定义主体,另一张定义构图,再一张定义材质处理。

00:40:43

Name the property you want from each.

请说清你想从每张图中得到什么属性。

00:40:45

The model may not isolate those roles perfectly.

模型未必能完美隔离这些角色。

00:40:49

That is why you should look for unwanted transfers, such as importing a reference's background when you only wanted its palette.

因此,要检查意外的信息转移,例如只想借用配色,却把参考图的背景也引入了。

00:40:56

Try removing one reference and observing what disappears.

试着移除一张参考图,观察什么随之消失。

00:41:00

Start with the smallest useful set.

从最小但有用的参考集合开始。

00:41:03

If five references contradict one another, adding a sixth may make the problem harder to diagnose rather than solve it.

如果五张参考图互相矛盾,再加第六张可能让问题更难诊断,而不是解决问题。

SLIDE 058 · 00:41:11 · Part 2 / Art direction

What did the reference contribute?

00:41:11

We have added a reference because we think it helps.

我们添加参考图,是因为认为它有帮助。

00:41:14

How would we discover whether it actually does?

怎样才能知道它是否真的有帮助?

00:41:16

Remove it and compare.

移除它,再比较。

00:41:19

First state its intended role: perhaps the pear's silhouette, perhaps the composition's sense of scale.

先说明它的预期作用:可能是梨的轮廓,也可能是构图的尺度感。

00:41:26

Then look for that quality in outputs with and without the reference.

然后,在有无这张参考图的输出中寻找这种品质。

00:41:30

Also inspect what came along uninvited.

也要检查哪些信息不请自来。

00:41:33

A useful subject reference can carry a background or lighting scheme we did not want.

有用的主体参考,也可能携带我们不想要的背景或光照。

00:41:38

This is an ablation: change one component to understand its contribution.

这就是消融实验:改变一个组件,以理解它的贡献。

00:41:44

Random generation makes one lucky comparison weak evidence, so several samples can help.

随机生成会让一次幸运的对比缺乏说服力,因此多个样本会有帮助。

00:41:51

If we cannot explain what a reference contributes, we may be making the request more complicated without making the direction clearer.

如果说不清一张参考图的贡献,我们可能只是把要求变复杂,却没有让方向更清楚。

SLIDE 059 · 00:41:59 · Part 2 / Art direction

An art-direction prompt

00:41:59

Listen to how this prompt distributes responsibility.

听听这个提示词如何分配职责。

00:42:03

The pear reference supplies silhouette and the blue stem.

梨的参考图提供轮廓和蓝色柄。

00:42:06

Placement low on the right organizes the composition.

将主体放在右下方,用来组织构图。

00:42:10

A pale upper-left area reserves space for the title.

左上方的浅色区域,为标题留出空间。

00:42:14

Soft daylight and restrained reflections support the mood.

柔和日光与克制的倒影,支撑整体情绪。

00:42:17

Every phrase points toward something we could inspect in an output.

每个短语,都指向输出中可以检查的东西。

00:42:22

Compare that with asking for a stunning masterpiece with beautiful lighting.

与“用漂亮光照创作一幅惊艳杰作”这样的要求比较一下。

00:42:26

This prompt is still a proposal, not a binding contract enforced by the model.

这个提示词仍然只是提议,不是模型强制执行的约束协议。

00:42:31

If an output fails, we can now name the failed relationship: the title area is crowded, the light is too assertive, or the reference's background has leaked into the scene.

输出失败时,我们现在可以指出哪种关系失败了:标题区太拥挤、光线太强势,或参考图背景渗入了场景。

SLIDE 060 · 00:42:43 · Part 2 / Art direction

Group Editing: one decision, many views

00:42:43

A single poster can look coherent by itself.

单张海报本身可能看起来很统一。

00:42:47

A series exposes a harder problem: does the same material decision survive across viewpoints?

系列作品暴露出更难的问题:同样的材质决定,能否在不同视角下保持?

00:42:53

Our Group Editing collaboration studies related images that should be edited consistently.

我们合作开展的 Group Editing 研究,关注应当保持一致编辑的关联图像。

00:42:58

Treating each image independently can produce slightly different costumes or materials.

独立处理每张图,可能生成略有不同的服装或材质。

00:43:03

The method arranges related images as pseudo-video frames to use a video model's consistency prior.

该方法把相关图像排列成伪视频帧,利用视频模型的一致性先验。

00:43:10

VGGT provides geometric correspondences.

VGGT 提供几何对应关系。

00:43:14

Geometry-enhanced rotary positional embeddings connect geometry features with image latents, while Identity-RoPE supports identity preservation.

几何增强的旋转位置编码,将几何特征与图像潜变量联系起来;Identity-RoPE 则支持身份保持。

00:43:24

Follow the penguin examples through the figure before reading every label.

阅读每个标签之前,先沿着图中的企鹅示例看。

00:43:28

Reliable correspondence is part of the technical problem: when views cannot be matched well, repeating the same instruction alone does not establish a coherent series.

可靠对应关系是技术问题的一部分:视角无法准确匹配时,仅重复相同指令,并不能保证系列统一。

SLIDE 061 · 00:43:40 · Part 2 / Art direction

Iteration as a controlled comparison

00:43:40

Iteration becomes informative when we know what changed between attempts.

当我们知道两次尝试之间改了什么,迭代才有信息量。

00:43:44

Suppose we compare two lighting treatments.

假设我们比较两种灯光处理。

00:43:47

Keep the subject and composition instructions stable, and record the references and model version.

保持主体和构图指令不变,并记录参考图和模型版本。

00:43:54

If a seed control exists, holding it fixed can help the first comparison.

如果可以控制随机种子,固定种子有助于初步比较。

00:43:59

A shared seed still does not guarantee identical composition after a prompt change.

但即使种子相同,改变提示词后,也不能保证构图完全一致。

00:44:04

If results vary widely, examine several outputs for each condition.

如果结果差异很大,就检查每种条件下的多个输出。

00:44:09

Otherwise, one lucky sample may decide the whole direction.

否则,一个幸运样本就可能决定整个方向。

00:44:13

The table is a proposed experiment, not measured results.

这张表是拟议实验,不是实测结果。

00:44:17

Its purpose is to make each generation answer a question about the artwork.

它的目的,是让每次生成都回答一个关于作品的问题。

SLIDE 062 · 00:44:21 · Part 2 / Art direction

A lucky image or a reliable direction?

00:44:21

Imagine you compare two prompts once, and the second gives a wonderful poster.

想象你只比较了两条提示词各一次,第二条生成了精彩海报。

00:44:28

Did its wording cause the improvement, or did it receive a favorable random sample?

改进是措辞造成的,还是它碰巧抽到了有利的随机样本?

00:44:33

From one output, it can be difficult to tell.

只看一个输出,往往很难判断。

00:44:36

Several outputs per condition reveal whether the direction is stable or merely fortunate.

每种条件生成多个结果,才能看出这个方向是稳定有效,还是仅仅幸运。

00:44:42

For art direction, we can ask two different questions.

在艺术指导中,我们可以问两个不同的问题。

00:44:45

Would I exhibit this particular image?

我愿意展出这张具体图像吗?

00:44:48

Could I reliably develop a series using this direction?

我能用这个方向,可靠地发展出一个系列吗?

00:44:52

One exceptional output may answer the first while leaving the second unresolved.

一张出色的输出可能回答了第一个问题,却没有解决第二个。

00:44:57

Repeating generations is helpful when it reveals a pattern relevant to the brief; it becomes a distraction when we keep browsing alternatives to avoid deciding what we value.

重复生成若能揭示与创作要求相关的规律,就有帮助;若只是为了逃避价值判断而不断浏览方案,就成了干扰。

SLIDE 063 · 00:45:08 · Part 2 / Art direction

Typography is part of the image

00:45:08

Now the image becomes a poster.

现在,图像变成了海报。

00:45:11

The title is an editable typographic layer, which lets us choose its wording, line breaks, and placement exactly.

标题是可编辑的文字图层,让我们精确选择措辞、换行和位置。

00:45:20

That matters when a work has to carry an exhibition name rather than merely resemble a poster in a generated preview.

作品需要准确呈现展览名称时,这很重要,而不是只在生成预览中“看起来像海报”。

00:45:28

Look at how the title occupies the area we deliberately left open.

看看标题如何占据我们有意留出的区域。

00:45:32

The image and type were designed to cooperate.

图像与文字是协同设计的。

00:45:35

We can adjust hierarchy without asking the model to regenerate the sculpture and risk changing it.

我们可以调整层次,而不必让模型重新生成雕塑、冒着改变它的风险。

00:45:41

This is a useful division of labor: generation develops the visual material, and direct layout controls exact communication.

这是一种有用的分工:生成负责发展视觉素材,直接排版负责精确传达信息。

00:45:49

The audience sees one composition.

观众看到的是一个整体构图。

00:45:52

It does not need to know which parts came from which tool.

他们不需要知道哪个部分来自哪种工具。

SLIDE 064 · 00:45:56 · Part 2 / Art direction

Typography needs a reading order

00:45:56

Typography establishes a sequence of attention.

文字排版建立了注意力的顺序。

00:45:59

Ask what the viewer should read first, what comes next, and where their eye returns to the artwork.

要问:观众应先读什么、接着读什么,以及视线在哪里回到作品?

00:46:05

In our poster, reserved space allows the title to be clear without covering the sculpture.

在我们的海报中,预留空间让标题清楚可读,又不会遮住雕塑。

00:46:10

This is also a production decision.

这也是一个制作决定。

00:46:13

Exact text can remain editable, while the image carries the material and atmosphere.

精确文字可以保持可编辑,而图像承载材质和氛围。

00:46:18

If the title feels too dominant, change size, placement, or contrast and inspect the whole composition again.

如果标题太抢眼,就调整大小、位置或对比度,再检查整个构图。

00:46:25

Do not judge the text in isolation from the image.

不要脱离图像,单独评价文字。

00:46:29

The poster is the relationship between them, including the empty space that lets each element do its job.

海报是两者之间的关系,也包括让各元素发挥作用的空白。

SLIDE 065 · 00:46:35 · Part 2 / Art direction

Two directions for the same brief

00:46:35

Take a few seconds with both directions before I describe them.

在我描述之前,先花几秒看看两个方向。

00:46:40

Which would you put outside AFTER RAIN?

你会把哪一个放在《雨后》展厅外?

00:46:42

Choose privately first, so your answer is not just a response to mine.

先自己选择,避免答案只是对我的回应。

00:46:47

Then identify the visible feature that made you choose.

然后指出促使你选择的可见特征。

00:46:51

The photographic direction can invite attention to material and atmosphere.

摄影方向可以引导观众关注材质与氛围。

00:46:56

The collage direction can make fragility feel constructed through paper edges and layers.

拼贴方向可以通过纸边和层次,让脆弱感显得被建构出来。

00:47:01

Both can serve the brief, but they make different promises to a visitor.

两者都可能符合要求,但向观众许下的承诺不同。

00:47:06

A useful defense goes beyond realism versus abstraction.

有用的辩护,不应止于写实与抽象的区别。

00:47:10

Tell us what the audience is likely to notice or feel, and how the composition produces that reading.

请告诉我们,观众可能注意或感受到什么,以及构图如何产生这种解读。

00:47:17

We will use the disagreement to decide what to revise, rather than vote for a universally better image.

我们将用分歧来决定修改方向,而不是投票选出普遍更好的图像。

SLIDE 066 · 00:47:24 · Part 2 / Art direction

Why would you exhibit this one?

00:47:24

Selection is an artistic decision.

选择本身就是艺术决定。

00:47:27

Once generation gives us many plausible alternatives, choosing one determines what the work becomes.

当生成提供许多合理方案时,选中哪一个,就决定作品最终成为什么。

00:47:34

The reason should connect to the brief: perhaps the small object feels more exposed, or the torn edge makes fragility physically legible.

理由应与创作要求相连:可能是小物体更显得无所庇护,或撕裂边缘让脆弱感变得可触可见。

00:47:43

The rejected direction is useful evidence too.

被放弃的方向也是有用证据。

00:47:47

Explain what it does well and why you are choosing something else for this exhibition.

请解释它哪里做得好,以及为什么为这个展览选择了别的方案。

00:47:52

An unexpected model output can also change the direction, if you choose to develop it deliberately.

如果你决定有意识地发展一个意外输出,它也可以改变创作方向。

00:47:58

The key question is whether you can now articulate the intention and carry it through subsequent decisions.

关键是:你现在能否说清意图,并在后续决定中贯彻它?

00:48:05

Surprise can begin a work; it does not finish the artist's judgment.

惊喜可以开启作品,却不能替艺术家完成判断。

SLIDE 067 · 00:48:11 · Part 2 / Art direction

A critique with observable evidence

00:48:11

A useful critique names a visible feature, explains its effect, and proposes a revision.

有用的评论会指出可见特征、解释效果,并提出修改建议。

00:48:16

For example: the crisp reflection makes the collage feel photographic, so I would simplify its shape to match the flatter layers.

例如:清晰倒影让拼贴显得像摄影,所以我会简化它的形状,以配合更平面的层次。

00:48:24

Compare that with 'I do not like the reflection.' The first comment gives the artist a relationship to inspect and a possible next move.

与“我不喜欢倒影”相比,前一种评论给艺术家提供了可检查的关系,以及可能的下一步。

00:48:33

You can disagree about the intended effect, but make the disagreement specific.

你们可以对预期效果有不同看法,但要把分歧说具体。

00:48:39

This is a classroom critique framework rather than an official grading rubric.

这是课堂评论框架,不是正式评分标准。

00:48:44

Use it to connect evidence in the image to the idea the work is trying to communicate.

用它把图像中的证据,与作品试图传达的观念联系起来。

SLIDE 068 · 00:48:49 · Part 2 / Art direction

LightCtrl: lighting as an artistic choice

00:48:49

A curator asks, can the same sculpture feel more vulnerable without changing its shape?

策展人问:能否不改变形状,却让同一件雕塑显得更脆弱?

00:48:55

Lighting is one way to answer.

灯光是一种回答方式。

00:48:58

Our LightCtrl collaboration studies controllable relighting from a single image, including light direction, intensity, and color temperature.

我们合作开展的 LightCtrl 研究,探索从单张图像进行可控重光照,包括光线方向、强度和色温。

00:49:07

Follow the chair through the figure.

请沿着图中的椅子示例看。

00:49:09

The latent proxy encoder extracts compact physical cues.

潜在代理编码器提取紧凑的物理线索。

00:49:13

A lighting-aware mask guides the denoiser toward regions affected by the change, and preference optimization in the proxy branch supports physical consistency.

光照感知蒙版引导去噪器关注受变化影响的区域,代理分支中的偏好优化则支持物理一致性。

00:49:24

The method connects an interpretable lighting request with image generation.

这种方法把可解释的光照要求与图像生成连接起来。

00:49:29

For our pear, look for consequences rather than the word dramatic: where do highlights move, what happens to shadow, and how does the material read?

对于梨,不要只看“戏剧性”这个词,而要看后果:高光移向哪里、阴影如何变化、材质呈现如何?

00:49:40

Inferring geometry and material from one photograph is ambiguous, so the result still needs inspection.

从单张照片推断几何和材质存在歧义,因此仍需检查结果。

00:49:47

The exhibition example is our application of the idea, not an additional result reported by the paper.

展览示例是我们对这一思路的应用,不是论文另外报告的实验结果。

SLIDE 069 · 00:49:54 · Part 2 / Art direction

Discuss: two artistic directions

00:49:54

Choose between the two prepared directions for AFTER RAIN.

请在为《雨后》准备的两个方向中做出选择。

00:49:58

Which poster communicates fragility more clearly, and why?

哪张海报更清楚地传达脆弱感?为什么?

00:50:02

What does the rejected direction reveal about your choice?

被放弃的方向,揭示了你选择中的什么?

00:50:06

Which single revision would most change the audience’s reading?

哪一个单独的修改,会最大程度改变观众的解读?

00:50:10

Discuss these three questions with a specific visual feature in view.

请看着一个具体视觉特征,讨论这三个问题。

00:50:15

You can prefer quiet space or the instability of torn paper, but explain how that choice serves the exhibition.

你可以偏爱安静空间,也可以偏爱撕纸的不稳定感,但要解释这种选择如何服务于展览。

00:50:22

Pause the video here.

请在这里暂停视频。

00:50:23

Listen for a persuasive reason to choose the direction you initially rejected.

试着听取一个有说服力的理由,支持你最初拒绝的方向。

SLIDE 070 · 00:50:32 · Part 2 / Art direction

A process record makes critique more precise

00:50:32

Consider what a process record lets us discuss.

想一想,过程记录能让我们讨论什么。

00:50:35

An artist can record the intention and invariants, the prompt and reference roles, and one observation about a rejected result.

艺术家可以记录意图与不变量、提示词与参考图角色,以及对一个被拒绝结果的观察。

00:50:44

That record explains what an attempt was testing.

这些记录说明了每次尝试在检验什么。

00:50:48

Without that context, a folder of attractive images can be difficult to interpret.

缺少背景信息,一文件夹好看的图像也可能很难解读。

00:50:53

We may not know which input changed or why a version was rejected.

我们可能不知道改了哪个输入,或为什么拒绝某个版本。

00:50:57

A few precise notes make a comparison more informative.

几条精确笔记,就能让比较更有信息量。

00:51:01

In a critique, this lets us ask about the artist’s decisions and the evidence behind them rather than guessing only from the final image.

在评论中,这让我们能够追问艺术家的决定及其证据,而不只是从最终图像猜测。

00:51:09

It also helps distinguish an intentional departure from an accidental change.

它也有助于区分有意偏离与意外变化。

SLIDE 071 · 00:51:14 · Part 2 / Art direction

A gallery conversation

00:51:14

Imagine we are standing between the two finished posters in a gallery review.

想象我们站在画廊评审现场,面前是两张完成的海报。

00:51:19

I would begin with a visible decision that carries the idea, then locate an unintended change, then propose one revision worth discussing.

我会先指出一个承载观念的可见决定,再找出一个意外变化,最后提出一个值得讨论的修改。

00:51:29

The order matters: the critique begins by understanding the work before prescribing a repair.

顺序很重要:评论应先理解作品,再给出修正方案。

00:51:35

We can hear two contrasting readings of the same poster.

对同一张海报,我们可能听到两种相反的解读。

00:51:38

One person may find the empty space quiet; another may find it emotionally distant.

一个人觉得留白安静,另一个人却觉得它情感疏离。

00:51:44

Ask which details support each reading.

请问,哪些细节支持各自的解读?

00:51:46

There is no image-making task here.

这里没有图像制作任务。

00:51:49

We are practicing how to make feedback useful to the next decision.

我们是在练习如何让反馈有助于下一次决定。

00:51:54

A good proposal names what would change and what effect we expect, so that a later version could confirm or challenge the reasoning.

好建议会说明改什么、预期产生什么效果,让后续版本能够证实或挑战这套推理。

00:52:02

Discuss these three questions, then pause the video for the gallery conversation.

请讨论这三个问题,然后暂停视频,进行画廊对话。

SLIDE 072 · 00:52:11 · Part 2 / Art direction

After Part 2: defend a decision

00:52:11

Before the artwork begins to move, defend a decision about the images.

在作品开始运动之前,请为一个图像决定提供理由。

00:52:17

How can a technically correct image fail as an artwork?

技术上正确的图像,为什么可能在艺术上失败?

00:52:20

Which visual rule should remain across a series?

一个系列中,哪条视觉规则应该保持?

00:52:24

What evidence makes a critique useful for the next revision?

什么证据能让评论对下一轮修改有用?

00:52:28

Discuss these three questions through one of the prepared posters.

请结合其中一张准备好的海报,讨论这三个问题。

00:52:31

Make words such as coherent or expressive concrete: locate the feature, describe its effect, and explain what changing it would do.

把“统一”“有表现力”等词说具体:指出特征、描述效果,并解释改变它会怎样。

00:52:41

Pause the video here.

请在这里暂停视频。

00:52:43

We will carry those artistic rules into the video section.

我们会把这些艺术规则带入视频部分。

SLIDE 073 · 00:52:50 · Topic 3 / Video editing

Generative video editing

00:52:50

Our image now has to survive time.

现在,图像必须经受时间的考验。

00:52:53

The subject can move, disappear, and return.

主体会运动、消失,然后返回。

00:52:57

We will examine selected frames from published research, follow the technical representations, and use a prepared AFTER RAIN scenario to decide what a convincing edit requires.

我们将查看已发表研究的精选帧、理解技术表示,并用准备好的《雨后》情境判断可信编辑需要什么。

00:53:09

Keep the preservation contract, but add an event that could break it.

保留编辑约定,再加入一个可能让它失效的事件。

SLIDE 074 · 00:53:14 · Topic 3 / Video editing

Before Topic 3: images in motion

00:53:14

A moving image creates new ways to break our contract.

动态图像会以新的方式破坏约定。

00:53:17

What new failures become possible when an image moves?

图像运动起来后,会出现哪些新的失败?

00:53:21

What should happen when the pear disappears behind a column?

梨消失在柱子后面时,应该发生什么?

00:53:25

Can four convincing frames prove that a video works?

四张可信的帧,能证明整段视频成立吗?

00:53:28

Discuss these three questions.

请讨论这三个问题。

00:53:31

Separate a correct change in visibility from an unwanted change in identity.

区分正确的可见性变化,与不希望发生的身份变化。

00:53:36

Name a moment you would need to see between the selected frames before accepting the shot.

接受这个镜头之前,请指出你需要在这些精选帧之间看到的某个时刻。

00:53:41

Pause the video here.

请在这里暂停视频。

00:53:43

Your prediction will give us a concrete test for the methods that follow.

你的预测会为后续方法提供一个具体测试。

SLIDE 075 · 00:53:51 · Topic 3 / Video editing

A video edit must survive the next frame

00:53:51

A still image lets us choose a flattering instant.

静态图像让我们挑选一个最好看的瞬间。

00:53:54

A video makes the object keep its promises in the next frame.

视频则要求物体在下一帧继续信守承诺。

00:53:58

Look at the source sequence.

看原始序列。

00:54:00

As viewpoint and visibility change, we continue to recognize the subject.

随着视角和可见性变化,我们仍能认出主体。

00:54:06

An edit has to preserve that relationship while introducing its requested transformation.

编辑必须在引入所要求变换的同时,保留这种关系。

00:54:12

Remember our two clocks.

记住两种时钟。

00:54:14

The sampling steps describe how a model generates a result.

采样步描述模型如何生成结果。

00:54:17

The frames here describe time passing in the depicted scene.

这里的帧描述被描绘场景中时间的流逝。

00:54:22

A system may use many sampling steps to produce a short clip, but those steps are not extra seconds of action.

系统可能用很多采样步生成短片,但这些步数不是额外的剧情秒数。

00:54:29

Now the preservation contract must hold across an event, not just inside a frame.

现在,保留约定必须跨越一个事件成立,而不只是局限于单帧内部。

SLIDE 076 · 00:54:35 · Topic 3 / Video editing

Consistency is not stillness

00:54:35

Would a video be perfectly consistent if every frame were identical?

如果每帧完全相同,视频就算完美一致吗?

00:54:39

Only if the intended scene were perfectly still.

只有目标场景本来就完全静止时才是。

00:54:43

In our moving shot, pose, viewpoint, illumination, and visibility should change.

在运动镜头中,姿态、视角、光照和可见性都应变化。

00:54:52

Consistency means that those changes remain coherent with the scene.

一致性意味着这些变化与场景保持协调。

00:54:56

The blue stem may become hidden during a turn.

蓝色柄可能在转动时被遮住。

00:55:00

That is different from its color changing without a lighting explanation.

这与没有光照原因却突然变色,是不同的。

00:55:04

A texture should move with its surface, rather than crawl independently across it.

纹理应该随表面一起移动,而不是独自在表面爬动。

00:55:09

If there is no convincing explanation, you may have found instability.

如果找不到可信解释,你可能发现了不稳定性。

00:55:14

This is why the goal cannot simply be to minimize change from one frame to the next.

因此,目标不能只是尽量减少相邻帧之间的变化。

SLIDE 077 · 00:55:20 · Topic 3 / Video editing

A moving image can reveal an idea

00:55:20

Here is a prepared storyboard, not a generated video result.

这是准备好的分镜,不是生成视频的结果。

00:55:24

In the first view, the reflection suggests an object we have not fully seen.

第一个视角中,倒影暗示了一个我们尚未看全的物体。

00:55:29

A partial view then gives us enough evidence to make a guess.

接着,局部视角提供足够线索,让我们猜测。

00:55:33

The wide view finally changes our understanding of the setting.

最后的全景改变了我们对环境的理解。

00:55:36

The audience is doing something during those five seconds: forming an expectation, testing it, and revising it.

在这五秒中,观众一直在行动:建立预期、检验预期,然后修正它。

00:55:44

Compare this sequence with revealing everything in the first frame.

把这个序列与第一帧就揭示一切的版本比较一下。

00:55:48

The same sculpture could be present, but the experience would differ.

同一件雕塑可以都在场,但体验会不同。

00:55:52

For AFTER RAIN, the timing of information can carry fragility as strongly as material or lighting.

对于《雨后》,信息揭示的时机,可以像材质或灯光一样有力地承载脆弱感。

00:55:58

We should decide that structure before asking a tool to fill in motion.

要求工具补充运动之前,我们应该先决定这种结构。

SLIDE 078 · 00:56:03 · Topic 3 / Video editing

Five seconds can have a structure

00:56:03

Five seconds is short, but it can still have a structure.

五秒虽然短,却仍然可以有结构。

00:56:06

Our proposed first interval shows a reflection, giving the viewer clues.

我们设想的第一段展示倒影,给观众线索。

00:56:11

The second offers a partial view.

第二段提供局部视角。

00:56:13

The final interval reveals the wider setting and changes how the object is understood.

最后一段揭示更广阔的环境,改变人们对物体的理解。

00:56:18

Try another allocation and predict the effect.

试着重新分配时间,并预测效果。

00:56:21

A longer reflection might create uncertainty; a quick reveal might make the piece feel like a product shot.

更长的倒影段落可能制造不确定感;快速揭示则可能让它像产品广告镜头。

00:56:28

These timings are artistic proposals, not measured outputs from a video model.

这些时间安排是艺术提案,不是视频模型的实测输出。

00:56:34

By deciding the experience first, we can evaluate whether the generated motion and cuts support it.

先决定体验,才能评价生成的运动和剪辑是否支持它。

00:56:40

A technically smooth sequence can still have the wrong rhythm.

技术上流畅的序列,节奏仍可能不对。

SLIDE 079 · 00:56:44 · Topic 3 / Video editing

An image editor can read a contact sheet

00:56:44

The research begins with a surprisingly simple experiment: arrange video frames into a contact sheet and give that image to an image editor.

这项研究始于一个出人意料的简单实验:把视频帧排成接触印样,再交给图像编辑器。

00:56:53

This asks whether an existing image-editing capability can transfer across multiple views presented together.

它想检验:现有图像编辑能力,能否迁移到一起呈现的多个视角上?

00:57:01

Treat the result as evidence that motivates a research direction.

请把结果当作启发研究方向的证据。

00:57:05

A grid can make frames available in a shared context, but a successful-looking selection does not prove reliable video editing.

网格让多帧共享上下文,但一组看似成功的精选帧,并不能证明视频编辑可靠。

00:57:12

There may be flicker between the sampled frames.

采样帧之间可能仍有闪烁。

00:57:16

The next steps investigate how to represent video more effectively and adapt the model, rather than assuming a contact sheet alone solves time.

后续研究探索更有效的视频表示和模型适配,而不是假设接触印样本身就解决了时间问题。

SLIDE 080 · 00:57:24 · Topic 3 / Video editing

Four good frames can hide a bad video

00:57:24

But the missing intervals are where some of the most revealing failures live.

恰恰是在缺失的间隔中,可能藏着最有揭示性的失败。

00:57:29

A feature can jump, disappear, and return between the frames on this page.

某个特征可能在这一页的两帧之间跳动、消失,再返回。

00:57:33

Use the contact sheet for what it does well: compare appearance across selected moments and locate regions to inspect.

发挥接触印样的长处:比较选定时刻的外观,并定位需要检查的区域。

00:57:41

Then review full playback, with closer attention around turns and occlusion.

然后完整播放,特别关注转向和遮挡附近。

00:57:45

Think of a film review based only on publicity stills.

想象只凭宣传剧照来评论电影。

00:57:49

You might judge the costume and lighting, but you have not yet seen the performance.

你或许能评价服装和灯光,却还没有看过表演。

00:57:55

In generative video, that missing performance includes whether the same object continues to exist convincingly from moment to moment.

对生成视频而言,缺失的“表演”包括:同一个物体是否在每个时刻都可信地持续存在。

SLIDE 081 · 00:58:04 · Topic 3 / Video editing

The grid can live in latent space

00:58:04

The grid can also be constructed in latent space.

网格也可以在潜在空间中构造。

00:58:07

Instead of only combining visible pixel images, the method encodes frames and arranges their tokens into a virtual image grid.

这种方法不只是组合可见像素图像,而是先编码各帧,再把 token 排列成虚拟图像网格。

00:58:16

The published example changes a bus into a graphics card.

论文中的示例把一辆公交车变成显卡。

00:58:20

The key idea is shared treatment of positions within that constructed representation.

核心思路是,在这种构造的表示中统一处理各位置。

00:58:25

Do not confuse this step with using a video VAE; they are different design decisions.

不要把这一步与使用视频 VAE 混为一谈;它们是不同的设计决定。

00:58:31

Putting frames near one another in a representation can help the model exchange information, but it does not impose a guarantee of physical continuity.

在表示中把帧放在一起,有助于模型交换信息,但不会强制保证物理连续性。

00:58:40

We still need to inspect what survives across changing views.

我们仍需检查视角变化时,哪些信息得以保留。

SLIDE 082 · 00:58:45 · Topic 3 / Video editing

A virtual grid is a representation choice

00:58:45

A virtual grid provides a way to arrange information for a model trained around image-like structure.

虚拟网格为围绕图像结构训练的模型,提供了一种信息排列方式。

00:58:51

It creates shared context and positional relationships among frame representations.

它在各帧表示之间建立共享上下文和位置关系。

00:58:56

That can be a useful bridge when reusing an image editor.

复用图像编辑器时,这可以成为有用的桥梁。

00:59:00

But arrangement alone does not enforce the laws of motion or the persistence of a hidden object.

但排列方式本身,不能强制执行运动规律或被遮挡物体的持续存在。

00:59:05

Those abilities depend on the model, adaptation, data, and other parts of the workflow.

这些能力取决于模型、适配、数据,以及工作流程中的其他部分。

00:59:12

Separate the representation choice from the capability claim.

请区分表示选择与能力主张。

00:59:16

A diagram can explain how frames enter a system without proving that the system handles every difficult event.

图表可以解释帧如何进入系统,却不能证明系统处理得了所有困难事件。

00:59:23

That proof would require appropriate outputs and evaluation.

证明这一点,需要适当的输出和评价。

SLIDE 083 · 00:59:28 · Topic 3 / Video editing

Repurposing an image model for video

00:59:28

This architecture asks whether an image editor's abilities can be reused for video.

这个架构检验图像编辑器的能力能否复用于视频。

00:59:33

The video VAE encodes a clip into a compressed latent representation.

视频 VAE 将片段编码成压缩的潜在表示。

00:59:39

Learned projections connect those video latents with the image-editing transformer, so the model can operate on a compatible arrangement of information.

学习得到的投影把视频潜变量与图像编辑 Transformer 连接起来,让模型处理兼容的信息排列。

00:59:48

Trace that main route before reading the smaller branches.

阅读较小分支之前,先沿着这条主路径看。

00:59:51

Adaptation connects them; the paper explores LoRA and full training choices, with an optional enhancement stage.

适配把它们连接起来;论文探索了 LoRA 和全量训练,并包含可选的增强阶段。

00:59:59

The attractive idea is reuse of editing knowledge.

吸引人的想法是复用编辑知识。

01:00:03

The test is whether that reuse preserves the temporal relationships our shot needs.

真正的检验是:这种复用是否保留了镜头需要的时间关系?

01:00:08

Architecture explains where information flows; the edited clip shows whether the intended event survives.

架构解释信息流向哪里;编辑后的片段展示预期事件是否得以保留。

SLIDE 084 · 01:00:15 · Topic 3 / Video editing

Adaptation connects incompatible representations

01:00:15

The image model and video VAE do not necessarily speak the same representational language.

图像模型与视频 VAE 未必使用相同的表示语言。

01:00:22

Learned projections help connect them.

学习得到的投影帮助连接二者。

01:00:25

Adaptation then allows the editing transformer to operate usefully with the video representation.

之后,适配让编辑 Transformer 能有效使用视频表示。

01:00:31

This is a common engineering pattern: reuse a capable component while learning the interface and behavior needed for another task.

这是一种常见工程模式:复用有能力的组件,同时学习另一项任务所需的接口和行为。

01:00:39

The benefit is not automatic.

收益并不会自动出现。

01:00:41

A projection must preserve useful information, and the adapted model must learn how the new structure relates to edits.

投影必须保留有用信息,适配后的模型也必须学会新结构如何与编辑相关联。

01:00:49

In the published pipeline, decoding returns the edited representation to video.

在已发表的流程中,解码将编辑后的表示还原为视频。

01:00:56

Follow the information through every stage rather than treating the name of the reused model as an explanation by itself.

请追踪信息经过每个阶段,而不要把被复用模型的名字本身当作解释。

SLIDE 085 · 01:01:04 · Topic 3 / Video editing

45 frames become 12: why?

01:01:04

Here is a counting puzzle with a small trap.

这里有一个带小陷阱的计数题。

01:01:06

In this example, forty-five pixel frames are compressed with a special first frame and a factor of four for the remaining temporal groups.

在这个例子中,四十五个像素帧采用特殊首帧处理,其余时间组按四倍压缩。

01:01:15

Forty-five is four times eleven, plus one.

四十五等于四乘十一,再加一。

01:01:19

The latent sequence therefore has eleven plus one, or twelve frames.

因此,潜在序列有十一加一,即十二帧。

01:01:24

Those twelve latent frames can be arranged in the illustrated three-by-four virtual grid.

这十二个潜在帧可以排成图中的三乘四虚拟网格。

01:01:29

Simply dividing forty-five by four would miss the first-frame convention.

简单地用四十五除以四,会忽略首帧约定。

01:01:34

This arithmetic belongs to the representation used in this example; it is not a rule for every video model.

这个算法属于本例采用的表示,并不是所有视频模型的通用规则。

01:01:41

A small detail in temporal compression changes what the editing model actually receives.

时间压缩中的一个小细节,会改变编辑模型实际收到什么。

SLIDE 086 · 01:01:47 · Topic 3 / Video editing

Temporal compression preserves a special first frame

01:01:47

Solve the frame-count relationship step by step.

一步步解出帧数关系。

01:01:50

Start with four k plus one equals forty-five.

从 4k 加一等于四十五开始。

01:01:54

Subtract one to obtain forty-four, then divide by four to get k equals eleven.

减一得到四十四,再除以四,得到 k 等于十一。

01:02:00

The latent sequence contains k plus one frames, so its length is twelve.

潜在序列包含 k 加一帧,因此长度为十二。

01:02:06

This arithmetic encodes the special handling of the first frame in the stated convention.

这个计算反映了该约定对首帧的特殊处理。

01:02:11

It is different from simply dividing the total number of pixel frames by four.

它不同于直接把像素帧总数除以四。

01:02:17

Understanding them can prevent mistakes when arranging latent frames into a virtual grid or comparing the representation with the original clip.

理解这些关系,可以避免在排列潜在帧网格、或与原始片段比较时出错。

SLIDE 087 · 01:02:26 · Topic 3 / Video editing

Guidance in a published video ablation

01:02:26

The requested edit turns the scene into a cyberpunk workshop with holographic documents.

编辑要求是把场景变成赛博朋克工作室,并加入全息文档。

01:02:32

Compare the results before focusing on the setting labels.

先比较结果,再关注设置标签。

01:02:36

Which version carries out the requested transformation more completely, and what visible evidence supports your answer?

哪个版本更完整地实现了要求的变换?有什么可见证据?

01:02:43

The paper presents this case as an example where classifier-free guidance improves edit completeness.

论文用这个案例说明,无分类器引导可以改善编辑完整性。

01:02:50

Its baseline retains source latents, and its implementation includes rescaling.

其基线保留了原始潜变量,实现中还包含重缩放。

01:02:55

This does not establish one universally best guidance value.

这并不意味着存在一个普遍最优的引导值。

01:02:59

Guidance also has a computation cost when it requires another model pass.

如果引导需要额外一次模型前向计算,也会带来计算成本。

01:03:05

Relate the example back to our equation: changing a prediction combination affects how strongly the result follows the condition.

联系之前的方程:改变预测组合,会影响结果遵循条件的力度。

SLIDE 088 · 01:03:13 · Topic 3 / Video editing

What this comparison actually shows

01:03:13

An ablation asks what changes when one component is removed.

消融实验追问:移除某个组件后,会发生什么变化?

01:03:17

In the displayed example, guidance makes more of the requested transformation visible.

在展示的例子中,引导让更多要求的变换可见。

01:03:22

That is useful evidence about this comparison.

这是关于这次比较的有用证据。

01:03:26

It is not yet a measurement of reliability across the kinds of shots we might produce for an exhibition.

但它还不是对我们可能用于展览的各类镜头的可靠性测量。

01:03:32

As a reader, separate the observation from the next question.

作为读者,请区分观察结果与下一步问题。

01:03:35

We can observe a more complete edit here.

我们可以观察到,这里的编辑更加完整。

01:03:38

We would still want to know what happens across other clips, how often preservation suffers, and what the computation costs.

但仍想知道:其他片段会怎样、保留效果多常受损,以及计算成本是多少。

01:03:46

You can learn a mechanism from a selected example while designing a stronger test for the decision you actually need to make.

你可以从精选示例中学习机制,同时为真正需要做出的决定设计更强的测试。

SLIDE 089 · 01:03:55 · Topic 3 / Video editing

Local editing across frames

01:03:55

This is a local edit: the sheep's face changes and receives a white star-shaped patch.

这是一次局部编辑:绵羊的脸发生变化,增加了一块白色星形斑纹。

01:04:00

Look at the patch across poses.

观察不同姿态下的这块斑纹。

01:04:03

Does it remain attached to the same facial region, with a plausible change in apparent shape as the head turns?

它是否始终附着在同一面部区域,并在头部转动时合理改变表观形状?

01:04:10

The most attractive frame is not enough.

仅看最漂亮的一帧是不够的。

01:04:12

An identity marker can slide, disappear, or change shape at another moment.

身份标记可能在另一个时刻滑动、消失或变形。

01:04:17

Selected frames let us ask the right questions, but the full clip is needed to check continuity between them.

精选帧让我们提出正确问题,但要检查帧间连续性,仍需完整视频。

01:04:25

For our discussion, identify a small recognizable feature that will make identity drift easier to notice.

请为讨论找一个可辨认的小特征,让身份漂移更容易被发现。

SLIDE 090 · 01:04:33 · Topic 3 / Video editing

Follow the identity marker

01:04:33

Choose a distinctive feature and follow it through the shot: its attachment to the subject, its shape, and its reappearance after partial visibility.

选一个独特特征,并在镜头中追踪它:与主体的附着关系、形状,以及部分遮挡后的再次出现。

01:04:43

For our pear, the blue stem is a useful witness.

对这只梨而言,蓝色柄是有用的见证者。

01:04:47

It may become hidden as the camera moves; we should allow that.

随着相机移动,它可能被遮住;这是应当允许的。

01:04:51

When it returns, it should still belong to the same object.

当它返回时,仍应属于同一个物体。

01:04:54

A marker that slides across the surface or reappears in a new shape tells us something a general impression of smooth motion might miss.

如果标记在表面滑动,或以新形状返回,就揭示了整体流畅感可能掩盖的问题。

01:05:02

We are using the marker as a diagnostic aid, while still judging the whole object's identity and the shot's intended motion.

我们用标记辅助诊断,同时仍评价整个物体的身份和镜头预期运动。

SLIDE 091 · 01:05:11 · Topic 3 / Video editing

Global stylization across frames

01:05:11

Here the instruction changes the whole sequence into a minimal monochrome sketch.

这里的指令把整个序列变成极简单色素描。

01:05:16

Local object replacement and global stylization allow different degrees of visual freedom.

局部物体替换与整体风格化,允许的视觉自由程度不同。

01:05:22

Even with a large style change, motion and scene structure should remain readable.

即使风格变化很大,运动和场景结构也应保持可读。

01:05:28

Look for rules across frames: line density, silhouette treatment, and the handling of depth.

寻找跨帧规则:线条密度、轮廓处理,以及深度的表现方式。

01:05:35

Do they feel like one visual language?

它们是否像同一种视觉语言?

01:05:38

This connects directly to our collage discussion.

这与我们的拼贴讨论直接相连。

01:05:42

A style is more useful to an art director when described through operations that can persist through time.

当风格被描述为能够持续存在于时间中的操作时,它对艺术总监更有用。

01:05:48

If every frame reinvents those operations, the result may feel unstable even when each still is appealing.

如果每一帧都重新发明这些操作,即使每张静帧好看,整体也可能不稳定。

SLIDE 092 · 01:05:56 · Topic 3 / Video editing

A style can flicker while motion stays correct

01:05:56

A video can preserve object motion while its style flickers.

视频可以保留物体运动,风格却仍然闪烁。

01:06:00

Line density may jump, a paper texture may crawl, or shading may switch between flat and volumetric treatments without a scene explanation.

线条密度可能跳变,纸张纹理可能爬动,明暗处理也可能无故在平面与立体之间切换。

01:06:09

Review style as a set of temporal rules, just as we reviewed it across surfaces in the collage.

像审查拼贴不同表面的风格那样,把视频风格作为一组时间规则来审查。

01:06:15

Some variation is appropriate when the camera or light changes.

相机或光线变化时,某些变化是合理的。

01:06:19

The question is whether the material language remains coherent.

问题是:材质语言是否仍然连贯?

01:06:23

If your artwork depends on a particular drawing or collage treatment, this check is as important as whether the subject stays in the same place.

如果作品依赖特定绘画或拼贴处理,这项检查与主体位置是否稳定同样重要。

01:06:32

Motion correctness alone does not establish visual consistency.

运动正确,本身并不能证明视觉一致。

SLIDE 093 · 01:06:37 · Topic 3 / Video editing

Why raw frame difference is misleading

01:06:37

If we subtract one frame from the next, a perfectly correct camera movement can create a large difference.

如果将相邻帧相减,完全正确的相机运动也会产生很大差异。

01:06:44

The same surface has simply moved to another pixel location.

同一个表面,只是移动到了另一个像素位置。

01:06:49

A more meaningful comparison first aligns corresponding visible content, then measures the remaining appearance difference.

更有意义的比较,是先对齐对应的可见内容,再测量剩余外观差异。

01:06:57

The teaching equation uses a motion warp for that alignment and a visibility mask to exclude occluded regions.

这个教学公式使用运动变换来对齐,并用可见性蒙版排除被遮挡区域。

01:07:04

We should not demand agreement for content that is hidden or newly revealed.

对于隐藏或刚被揭示的内容,不应要求直接一致。

01:07:09

This is an illustrative metric, not a claim about the exact evaluation used by the paper.

这是示意性指标,不是在声称论文使用了这一精确评测。

01:07:15

A low difference can also reward a video that barely moves.

低差异也可能奖励几乎不动的视频。

01:07:19

A metric must be checked against the task it is supposed to represent.

必须对照指标所要代表的任务,检查指标本身。

SLIDE 094 · 01:07:24 · Topic 3 / Video editing

The frozen video wins the wrong test

01:07:24

Imagine two candidates for our five-second shot.

想象五秒镜头的两个候选版本。

01:07:27

One shows a coherent camera move around the pear with modest frame differences.

一个展示相机连贯地绕梨移动,帧间差异适中。

01:07:32

The other repeats a single beautiful frame.

另一个反复显示同一张漂亮图像。

01:07:35

A naive difference score might prefer the frozen version.

简单的差异分数,可能偏爱冻结版本。

01:07:39

Our audience would immediately notice that the intended reveal never happened.

观众却会立刻发现,预期的揭示根本没有发生。

01:07:43

That is a useful example of optimizing the measurement while missing the purpose.

这是优化测量指标却错失目的的一个例子。

01:07:48

We need evidence for both coherent appearance and requested motion.

我们需要外观连贯和所要求运动两方面的证据。

01:07:52

Neither can substitute for the other.

两者不能相互替代。

01:07:55

Each time, one easy-to-observe quality tempted us to stand in for the whole brief.

每一次,我们都容易让某个易于观察的品质,代表整份创作要求。

01:08:00

Good evaluation keeps the intended result in view, especially when a convenient score seems reassuring.

良好的评测始终关注预期结果,尤其当方便的分数令人安心时。

SLIDE 095 · 01:08:08 · Topic 3 / Video editing

A keyframe can anchor the edit

01:08:08

An edited keyframe can make an art direction easier to approve.

编辑后的关键帧,可以让艺术方向更容易获得认可。

01:08:12

Instead of describing every material and lighting choice in words, we can point to an image and say, this is the appearance we want.

不必用文字描述每个材质和灯光决定,我们可以指着一张图说:这就是想要的外观。

01:08:21

The workflow then carries that approved look through the source clip.

工作流程再把获批的外观延续到原始片段中。

01:08:24

Follow the four stages: select a representative frame, edit and inspect its appearance, propagate the look, and review the resulting motion.

请跟随四个阶段:选代表帧,编辑并检查外观,传播外观,再审查运动结果。

01:08:36

Will the material persist during a turn?

转向时,材质能保持吗?

01:08:39

Will the identity survive a column passing in front of it?

柱子从前方经过时,身份能保留吗?

01:08:44

This distinction is central to the Runway reading later: an appealing interaction pattern still needs a shot review suited to the artwork.

这个区别是后面 Runway 阅读的核心:吸引人的交互方式,仍需要适合这件作品的镜头审查。

SLIDE 096 · 01:08:52 · Topic 3 / Video editing

AlignVid: when the image overrules the text

01:08:52

What if the approved image is so influential that the requested event never happens?

如果获批图像影响力太强,以至于要求的事件根本没发生,怎么办?

01:08:57

Our AlignVid collaboration studies that tension in image-to-video generation.

我们合作开展的 AlignVid 研究,探索图生视频中的这种张力。

01:09:02

In the published examples, a baseline omits a sunflower or leaves a person standing; the corresponding AlignVid results implement more of the requested event.

论文示例中,基线漏掉了向日葵,或让人物一直站着;对应的 AlignVid 结果实现了更多要求的事件。

01:09:13

The intervention scales queries or keys in selected attention blocks and denoising steps without retraining.

这种干预无需重新训练,只在选定注意力模块和去噪步骤中,缩放查询或键。

01:09:19

In a scalar form, scaling Q by gamma changes the weights to softmax of gamma times Q K transpose over square root d.

在标量缩放形式下,用 gamma 缩放 Q 后,权重变成 softmax(gamma × QKᵀ / √d)。

01:09:28

It changes attention concentration.

它改变了注意力的集中程度。

01:09:31

Classifier-free guidance instead combines predictions with different conditioning; these are distinct operations.

无分类器引导则组合不同条件下的预测;两者是不同操作。

01:09:38

For an artist, the interesting failure is a faithful-looking image that refuses to do what the scene requires.

对艺术家来说,有意思的失败是:图像看起来很忠实,却拒绝做场景要求的事。

01:09:44

The published frames illustrate that tension; a finished shot still needs temporal review.

已发表的帧展示了这种张力;完成的镜头仍需进行时间连续性审查。

SLIDE 097 · 01:09:50 · Topic 3 / Video editing

A five-second video edit brief

01:09:50

Our five-second brief changes the pear's body to ivory ceramic while retaining the blue stem, camera movement, and position.

我们的五秒要求,是把梨的主体变为象牙白陶瓷,同时保留蓝色柄、相机运动和位置。

01:09:58

Reflections may adapt to the new material.

倒影可以适应新材质。

01:10:00

The pear must remain the same object after passing behind a column.

梨经过柱子后面再出现时,必须仍是同一个物体。

01:10:05

Which clause is hardest to verify?

哪一条最难验证?

01:10:08

The reappearance is a strong candidate, because the system must maintain identity through a period of invisibility.

再次出现很可能最难,因为系统必须在一段不可见期间保持身份。

01:10:14

This is more demanding than transferring a visible color from frame to frame.

这比逐帧传递可见颜色更具挑战。

01:10:19

Begin with one clear transformation in a short clip.

先在短片中做一个明确的变换。

01:10:23

Then make the critical visibility event part of the review, rather than discovering it only after choosing a favorite result.

然后把关键可见性事件纳入审查,而不是选好最喜欢的结果后才发现它。

SLIDE 098 · 01:10:31 · Topic 3 / Video editing

The moment the pear returns

01:10:31

The column is the moment of truth.

柱子就是关键考验。

01:10:34

Before the pear disappears, we can inspect its silhouette, material, and stem.

梨消失前,我们能检查它的轮廓、材质和柄。

01:10:39

During occlusion, there may be nothing visible to compare.

遮挡期间,可能没有可见内容可供比较。

01:10:43

When it returns, the model has to make the same object convincing again.

返回时,模型必须再次让同一个物体显得可信。

01:10:48

Do not reject correct invisibility as a failure.

不要把正确的不可见状态当作失败。

01:10:51

Instead, compare the identity before and after the event, including how the object emerges at the boundary.

而应比较事件前后的身份,包括物体如何从边界处显露。

01:10:58

A smooth-looking clip can still reveal a subtly redesigned pear on the other side.

看起来流畅的视频,也可能在另一侧露出一只被悄悄重新设计的梨。

01:11:03

We are not merely watching for anything strange; we are testing the preservation requirement at the point where the shot makes it hardest to satisfy.

我们不只是寻找怪异之处,而是在镜头最难满足要求的地方,检验保留约束。

SLIDE 099 · 01:11:12 · Topic 3 / Video editing

A small addition creates new relationships

01:11:12

Adding a small object creates several new relationships.

添加一个小物体,会创造多种新关系。

01:11:16

In this published example, the edit adds a drone.

在这个已发表示例中,编辑添加了一架无人机。

01:11:20

Even if the drone looks convincing by itself, its scale, placement, and movement must fit the scene.

即使无人机本身很可信,它的尺度、位置和运动也必须适合场景。

01:11:27

Imagine a camera moving forward while the added object changes size in the wrong direction.

想象相机向前移动,新增物体的大小却朝错误方向变化。

01:11:32

The object might be beautifully rendered, but its relationship with the camera would expose the edit.

物体可能渲染得很漂亮,但它与相机的关系会暴露编辑。

01:11:38

Or it might hover at an unintended location relative to other objects.

或者,它可能相对其他物体悬停在错误位置。

01:11:42

When you review an addition, evaluate those relationships across time.

审查新增物体时,要跨时间评价这些关系。

01:11:47

A close-up of the inserted object cannot answer every question about whether it belongs.

仅看新增物体的特写,无法回答它是否融入场景的所有问题。

SLIDE 100 · 01:11:53 · Topic 3 / Video editing

An addition must belong to camera motion

01:11:53

An added object must fit the camera as well as the scene.

新增物体既要符合场景,也要符合相机。

01:11:57

A convincing texture and shape in one still do not establish a coherent trajectory.

单张静帧中可信的纹理和形状,不能证明运动轨迹连贯。

01:12:02

If the camera approaches, the apparent size and position of the addition should evolve in a compatible way.

如果相机靠近,新增物体的表观大小和位置应相应变化。

01:12:09

The exact expectation depends on whether the object is stationary or moving independently.

具体预期取决于物体是静止,还是独立运动。

01:12:15

That is why the brief should specify the intended relationship.

因此,创作说明应该明确预期关系。

01:12:19

Inspect the addition relative to nearby objects and the background, not only in a crop.

要相对附近物体和背景检查新增内容,而不只看裁剪图。

01:12:26

A production-quality edit is a collection of relationships that survive over time, rather than a new object that looks impressive in isolation.

制作级编辑是一组经得起时间考验的关系,而不是一个单独看很惊艳的新物体。

SLIDE 101 · 01:12:35 · Topic 3 / Video editing

The shot review

01:12:35

Review the whole shot at normal speed, then inspect moments where failure is most likely.

先以正常速度看完整个镜头,再检查最可能失败的时刻。

01:12:41

Pay attention to identity, intended motion, occlusion, boundaries, and the ending.

关注身份、预期运动、遮挡、边界和结尾。

01:12:46

High-motion content is a limitation the research authors specifically identify.

高运动量内容,是研究作者明确指出的局限。

01:12:51

If a clip fails, choose a revision based on the failure.

片段失败时,要根据具体失败选择修改。

01:12:54

You might reduce the transformation, shorten the segment, add a reference frame where supported, or repair part of the result through conventional compositing.

可以减小变换幅度、缩短片段、在支持时添加参考帧,或用传统合成修复部分结果。

01:13:05

Another generation is useful when it tests a reasoned hypothesis.

当新一轮生成用于检验有理由的假设时,它才有用。

01:13:09

Do not let repeated attempts distract you from whether the piece still communicates the artistic intention you began with.

不要让反复尝试分散注意力,忘记作品是否仍在传达最初的艺术意图。

SLIDE 102 · 01:13:16 · Topic 3 / Video editing

A failed shot suggests a different workflow

01:13:16

A failed shot is information about the workflow.

失败镜头提供了关于工作流程的信息。

01:13:20

If the appearance is wrong from the beginning, revising the keyframe or the material direction may help.

如果一开始外观就不对,修改关键帧或材质方向可能有帮助。

01:13:26

If the first frame is convincing but identity breaks after occlusion, the problem calls for temporal review and a more suitable propagation or editing approach.

如果首帧可信,但遮挡后身份破坏,就需要时间连续性审查,以及更合适的传播或编辑方法。

01:13:36

If a region must remain exact, direct compositing may be part of the solution.

如果某个区域必须完全不变,直接合成可以是解决方案的一部分。

01:13:41

We can also split a complicated transformation into stages, provided the transitions remain coherent.

也可以把复杂变换拆成多个阶段,但必须保证过渡连贯。

01:13:48

The important move is to connect the observed failure to the next intervention.

关键是把观察到的失败,与下一步干预联系起来。

01:13:53

Repeating a vague request with more enthusiasm does not use what the failed shot has taught us.

更热情地重复模糊要求,并没有利用失败镜头教给我们的东西。

SLIDE 103 · 01:13:59 · Topic 3 / Video editing

Discuss: the difficult moment

01:13:59

Return to the proposed five-second ceramic-pear shot.

回到拟议的五秒陶瓷梨镜头。

01:14:03

Which moment would you inspect most closely?

你会最仔细检查哪个时刻?

01:14:05

What should remain recognizable after the column?

经过柱子后,什么应该仍然可辨认?

01:14:08

What failure would make you reject an attractive clip?

什么失败会让你拒绝一段好看的视频?

01:14:12

Discuss these three questions by describing an event and the evidence you would seek.

请通过描述事件和所需证据,讨论这三个问题。

01:14:17

If your answer is temporal consistency, make it visible: what changes, when, and why would that violate the intended shot?

如果你的答案是“时间一致性”,请让它可见:什么在何时改变,为什么违反了镜头意图?

01:14:27

Pause the video here.

请在这里暂停视频。

01:14:29

Compare your acceptance criteria with another person’s before we examine the compact review plan.

在查看简要审查方案之前,先与另一位同学比较验收标准。

SLIDE 104 · 01:14:38 · Topic 3 / Video editing

A five-second test plan

01:14:38

Here is a compact test plan.

这是一个简要测试方案。

01:14:41

The request is an ivory ceramic body with the blue stem retained.

要求是象牙白陶瓷主体,并保留蓝色柄。

01:14:45

Camera motion and object motion should follow the source.

相机运动和物体运动应遵循原片。

01:14:49

The clip includes a column so that we can inspect a difficult reappearance.

片段中包含柱子,让我们检查困难的再次出现过程。

01:14:53

The evidence includes normal-speed playback and a closer look before and after that event.

证据包括正常速度播放,以及事件前后的仔细观察。

01:14:59

We also inspect reflections because a material change has optical consequences.

还要检查倒影,因为材质变化会产生光学后果。

01:15:04

This plan does not require a long benchmark suite.

这个方案不需要庞大的基准测试集。

01:15:08

It simply makes the request, the likely failure, and the acceptance evidence explicit.

它只是明确要求、可能的失败,以及验收证据。

01:15:14

You can use the same structure to test a different subject or visual transformation.

你可以用同样的结构测试其他主体或视觉变换。

SLIDE 105 · 01:15:19 · Topic 3 / Video editing

Three decisions without a model

01:15:19

Let us check whether the mechanisms are now useful without a model in front of us.

现在没有模型在面前,让我们检验这些机制是否真的有用了。

01:15:24

Consider the three questions on screen.

请考虑屏幕上的三个问题。

01:15:27

A material change reaches the reflection: explain why.

材质变化会影响倒影:请解释原因。

01:15:32

A subject reference and LoRA both help represent a concept: explain the difference.

主体参考图和 LoRA 都有助于表达概念:请解释两者区别。

01:15:39

Four frames look excellent: explain why the clip can still fail.

四帧看起来很出色:请解释整段视频为什么仍可能失败。

01:15:44

Take a short silent beat before answering.

回答前,先安静思考片刻。

01:15:47

The next slide names the decisions hiding inside the questions.

下一页会指出这些问题中隐藏的决定。

01:15:51

If you get stuck, locate the type of problem first: the physical scene, the model's information, or evidence across time.

如果卡住了,先定位问题类型:物理场景、模型获得的信息,还是跨时间的证据。

01:16:00

That is often enough to begin a clear answer.

通常,做到这一步就足以开始给出清晰答案。

SLIDE 106 · 01:16:03 · Topic 3 / Video editing

Match the problem to the intervention

01:16:03

Here are the decisions those questions conceal.

这些问题隐藏着以下决定。

01:16:06

If the reflection must change with the new material, we need to define the permitted region and consequences of the edit.

如果倒影必须随新材质变化,就需要定义编辑允许的区域和连带后果。

01:16:14

If we want a concept for this request, a reference can condition generation; if we want learned adaptation, LoRA changes selected weights.

如果只为这次要求表达概念,参考图可提供生成条件;如果需要学习式适配,LoRA 会改变选定权重。

01:16:25

If continuity matters, the evidence must include the intervals between our chosen stills.

如果重视连续性,证据就必须包含精选静帧之间的间隔。

01:16:30

Each decision rules out a tempting shortcut.

每个决定,都排除了一条诱人的捷径。

01:16:34

Freezing every surrounding pixel may freeze the wrong reflection.

冻结周围每一个像素,可能冻结错误的倒影。

01:16:38

Calling every supplied image training confuses conditioning with adaptation.

把所有输入图像都称为训练,会混淆条件控制与适配。

01:16:43

Calling four stills a successful video omits time.

把四张静帧称为成功视频,就遗漏了时间。

01:16:47

Use this slide to repair your explanation, rather than memorize a preferred sentence.

请用这一页修正你的解释,而不是背诵某句标准答案。

SLIDE 107 · 01:16:53 · Topic 3 / Video editing

The mechanisms behind your decisions

01:16:53

The material changes how light interacts with the pear, so its visible consequences can extend into the reflection.

材质改变光与梨的相互作用,因此可见后果会延伸到倒影。

01:17:01

A reference conditions an output; LoRA learns a low-rank update to selected model weights.

参考图为输出提供条件;LoRA 为选定模型权重学习低秩更新。

01:17:07

Four good frames leave unobserved transitions, including events such as occlusion where identity can fail.

四张好帧仍留下未观察的过渡,包括可能导致身份失败的遮挡事件。

01:17:14

Those are compact answers, but each has an application.

这些答案很简短,但每个都有实际用途。

01:17:18

They help us write a better preservation contract, choose between two kinds of intervention, and design a more revealing review.

它们帮助我们写出更好的保留约定、在两类干预中选择,并设计更有揭示性的审查。

01:17:26

If your answer used different words and preserved those distinctions, it works.

如果你用了不同措辞,但保留了这些区别,答案就成立。

01:17:31

Tomorrow's interface may rename its controls, but the difference between specifying a request, adapting a model, and checking its result will still matter.

明天的界面可能给控件改名,但提出要求、适配模型与检查结果之间的区别,仍然重要。

SLIDE 108 · 01:17:41 · Topic 3 / Video editing

Three levels of checking an edit

01:17:41

We can check an edit at three levels.

我们可以在三个层面检查编辑。

01:17:43

First, did the requested transformation occur?

第一,所要求的变换发生了吗?

01:17:47

Second, did the necessary invariants survive?

第二,必要的不变量保留了吗?

01:17:50

Third, does the result serve the artwork's intention?

第三,结果是否服务于作品意图?

01:17:54

A result can pass the first level and fail the second, as when the correct material appears on a different object.

结果可能通过第一层却不通过第二层,例如正确材质出现在了另一个物体上。

01:18:01

It can pass both technical levels and still fail the artistic one, as when the chosen effect undermines the intended mood.

也可能通过两个技术层面,却在艺术层面失败,例如所选效果削弱了预期情绪。

01:18:10

Keeping the levels separate makes critique more precise.

区分这些层面,会让评论更精确。

01:18:14

It also explains why neither a generic aesthetic score nor a strict pixel comparison can answer every question we care about.

这也解释了为什么通用审美分数和严格像素比较,都无法回答我们关心的所有问题。

SLIDE 109 · 01:18:22 · Topic 3 / Video editing

Who decides?

01:18:22

We have spent the lecture making the editing request more precise.

整堂课,我们都在让编辑要求更精确。

01:18:27

Now consider who gets to choose the request in the first place.

现在想想:一开始由谁来选择提出什么要求?

01:18:31

A system can offer convincing alternatives, but somebody still decides which intention matters and which evidence counts as success.

系统可以提供可信方案,但仍需要有人决定什么意图重要,以及什么证据才算成功。

01:18:41

The three readings let us examine that responsibility from different positions.

三篇阅读让我们从不同位置审视这种责任。

01:18:46

Pachocki raises questions about capable AI systems and human values.

Pachocki 提出了能力强大的 AI 系统与人类价值观之间的问题。

01:18:51

Koe asks readers to reconsider their own goals and habits.

Koe 请读者重新审视自己的目标与习惯。

01:18:55

Runway presents a workflow for turning an approved image into a video edit.

Runway 展示把获批图像转化为视频编辑的工作流程。

01:18:59

Keep our pear in mind as we read them: we can delegate parts of its production while still arguing about what the artwork should become.

阅读时请记住这只梨:我们可以委托部分制作,同时仍讨论作品应该成为什么。

SLIDE 110 · 01:19:08 · Topic 3 / Video editing

After Topic 3: judge the whole shot

01:19:08

We have an approved appearance and a moving shot to judge.

现在,我们有了获批外观,以及需要评价的运动镜头。

01:19:11

What can an approved keyframe establish, and what can it not?

获批关键帧能确定什么,又不能确定什么?

01:19:16

Why might a low frame-difference score reward a bad video?

为什么低帧差分数可能奖励一段坏视频?

01:19:20

Which evidence would convince you that identity survived motion?

什么证据能说服你,身份经受住了运动?

01:19:24

Discuss these three questions.

请讨论这三个问题。

01:19:26

Your answer should account for the requested event as well as the subject’s appearance.

答案既要考虑主体外观,也要考虑所要求的事件。

01:19:30

Use the column, a turn, or the sketch treatment as a concrete example.

可以用柱子、转向或素描处理作为具体例子。

01:19:35

Pause the video here.

请在这里暂停视频。

01:19:36

We will next ask how the readings change our view of responsibility for those decisions.

接下来,我们将追问阅读如何改变我们对这些决定所涉及责任的看法。

SLIDE 111 · 01:19:45 · Readings and technical extensions

AI goals and human direction

01:19:45

These first two readings ask about direction from different sides.

前两篇阅读从不同方面追问方向。

01:19:49

An Alien Mind considers the relationship between an AI system achieving a goal and respecting human values.

《An Alien Mind》思考 AI 系统实现目标与尊重人类价值观之间的关系。

01:19:57

Dan Koe's essay invites people to examine the goals and habits directing their own lives.

Dan Koe 的文章邀请人们审视引导自己生活的目标和习惯。

01:20:02

We will use its proposal as material for critical discussion of creative practice.

我们把他的主张作为批判性讨论创作实践的材料。

01:20:07

Find a specific claim, explain your interpretation, and connect it to a decision from the lecture.

找出一个具体主张,解释你的理解,并联系课堂中的一个决定。

01:20:14

A fictional artist is fine for discussing personal direction.

讨论个人方向时,可以使用虚构艺术家的例子。

01:20:18

The fifth quiz question asks you to reflect on one reading of your choice.

测验第五题要求你从自选的一篇阅读出发进行反思。

01:20:23

Our class discussion can compare all three perspectives without requiring everyone to disclose personal experiences.

课堂可以比较三种视角,而不要求每个人披露个人经历。

SLIDE 112 · 01:20:30 · Readings and technical extensions

Reading 3: image-guided video editing

01:20:30

The third reading moves from questions about direction to a concrete editing workflow.

第三篇阅读从方向问题,转向具体编辑流程。

01:20:36

Runway's announcement of Aleph 2.0 and Edit Studio describes establishing an appearance in an edited image and applying that change through video.

Runway 关于 Aleph 2.0 和 Edit Studio 的公告,描述了先在编辑图像中确定外观,再把变化应用到视频的方法。

01:20:45

It connects directly to our keyframe discussion.

它与关键帧讨论直接相连。

01:20:48

Read it with the ceramic pear and column in mind.

阅读时,请想着陶瓷梨和柱子。

01:20:51

Which decision does the interface make easier to express?

这个界面让哪项决定更容易表达?

01:20:55

Which result would you still need to watch before approval?

批准之前,哪些结果仍需要亲眼播放检查?

01:20:59

The page is a vendor's account, so its demonstrations illustrate claimed capabilities rather than supply an independent comparative evaluation.

这是一篇厂商介绍,因此演示说明的是其声称的能力,而不是独立的比较评测。

01:21:08

We can learn from the interaction design and still propose an occlusion test that asks whether the workflow meets this exhibition's particular needs.

我们可以借鉴交互设计,同时提出遮挡测试,检验流程是否满足这个展览的具体需要。

SLIDE 113 · 01:21:17 · Readings and technical extensions

Optional technical reading

01:21:17

These technical papers are optional companions to the discussion readings.

这些技术论文,是讨论阅读之外的可选配套材料。

01:21:21

Choose according to the question you want to investigate.

请根据你想研究的问题选择。

01:21:24

InstructPix2Pix helps explain edit supervision; Flow Matching develops the generative training framework.

InstructPix2Pix 帮助解释编辑监督;Flow Matching 发展了生成训练框架。

01:21:32

Qwen-Image-2.0 and Qwen-Video-Edit provide recent architecture examples.

Qwen-Image-2.0 和 Qwen-Video-Edit 提供了近期架构示例。

01:21:38

You do not need to read all of them before trying the class discussion.

参加课堂讨论之前,不必读完全部论文。

01:21:42

Start with a question, locate the part of the paper that addresses it, and distinguish the proposed method from the authors' evidence.

从一个问题开始,找到论文中回应它的部分,并区分提出的方法与作者的证据。

01:21:50

The slide notes retain the other references on multi-image inputs, rewards, composition, and video editing.

幻灯片备注保留了多图输入、奖励、构图和视频编辑方面的其他参考资料。

SLIDE 114 · 01:21:58 · Readings and technical extensions

Three source types, three kinds of evidence

01:21:58

The three readings provide different kinds of material.

三篇阅读提供不同类型的材料。

01:22:02

A research leader's perspective develops an argument about alignment and future development.

研究领导者的观点文章,提出关于对齐和未来发展的论证。

01:22:07

A reflective essay offers a way to examine personal direction.

反思性文章提供一种审视个人方向的方法。

01:22:12

A vendor announcement presents a workflow and promotes its capabilities.

厂商公告展示工作流程,并推广其能力。

01:22:16

Read each according to what it can support.

阅读每种材料时,要考虑它能够支持什么。

01:22:19

Identify forecasts, personal claims, and demonstrations rather than treating every sentence as the same kind of evidence.

识别预测、个人主张和演示,不要把每句话都当作同类证据。

01:22:27

You may disagree with an author and still find a useful question.

你可以不同意作者,却仍从中找到有用的问题。

01:22:31

Our synthesis is practical: what do you want to make, what will you delegate, and what evidence will you use to decide whether the collaboration served that intention?

我们的综合问题很实际:你想创作什么、委托什么,以及用什么证据判断合作是否服务于这一意图?

SLIDE 115 · 01:22:42 · Readings and technical extensions

Separate image and text guidance

01:22:42

This technical extension separates image guidance and text guidance in InstructPix2Pix.

这个技术延伸,区分 InstructPix2Pix 中的图像引导和文本引导。

01:22:49

Begin with the prediction using neither condition.

从不使用任何条件的预测开始。

01:22:52

Add a scaled difference for including the image, then another scaled difference for adding the text alongside that image.

加上引入图像后的缩放差值,再加上在图像之外引入文本后的另一个缩放差值。

01:22:59

Set both scales to one and follow the cancellation: the intermediate terms disappear, leaving the fully conditioned noise prediction.

把两个强度都设为一,观察抵消:中间项消失,只留下完全条件化的噪声预测。

01:23:07

This is a useful check that you understand the expression.

这是检查你是否理解表达式的有用方法。

01:23:11

Epsilon here denotes a noise prediction, unlike the velocity in our flow-matching slides.

这里的 epsilon 表示噪声预测,与流匹配页面中的速度不同。

01:23:18

These controls belong to this formulation and should not be assumed to map directly onto every current editor's interface.

这些控制属于这一公式,不应假设它们直接对应所有当前编辑器的界面。

SLIDE 116 · 01:23:26 · Readings and technical extensions

What remains when the terms cancel?

01:23:26

Now set both guidance scales to one and follow the cancellation.

现在把两个引导强度都设为一,观察抵消。

01:23:30

The initial no-condition prediction cancels its negative copy in the first difference.

初始无条件预测,与第一个差值中的负项抵消。

01:23:36

The image-only prediction then cancels its negative copy in the second difference.

随后,仅图像条件的预测,与第二个差值中的负项抵消。

01:23:41

The result is the prediction with both image and text conditions.

最终得到同时使用图像和文本条件的预测。

01:23:45

That gives us a reference point for understanding the two controls.

这为理解两个控制提供了参考点。

01:23:49

If the text scale were zero while the image scale stayed one, we would instead recover the image-only prediction.

如果文本强度为零、图像强度仍为一,就会得到仅使用图像条件的预测。

01:23:57

It shows what information each difference adds.

这说明了每个差值添加什么信息。

01:24:00

Keep epsilon's meaning explicit here: this InstructPix2Pix expression predicts noise, whereas the earlier flow-matching example predicted velocity.

请明确 epsilon 的含义:这里的 InstructPix2Pix 公式预测噪声,而前面的流匹配例子预测速度。

SLIDE 117 · 01:24:11 · Readings and technical extensions

Training: the target supplies the lesson

01:24:11

Read the pseudocode in two groups.

把伪代码分成两组来读。

01:24:14

First we prepare an example: source, instruction, target, conditions, and target latent.

首先准备样本:原图、指令、目标、条件,以及目标潜变量。

01:24:23

Then we sample noise and time, construct the intermediate latent, predict velocity, and compare that prediction with the known training target.

然后采样噪声和时间,构造中间潜变量,预测速度,并将预测与已知训练目标比较。

01:24:32

The final update changes trainable weights based on the loss.

最后根据损失更新可训练权重。

01:24:37

At inference, there is no known edited target to supply in this way.

推理时,并没有已知编辑目标可以这样提供。

01:24:41

We instead follow the learned generative predictions from a starting state.

我们从初始状态出发,遵循学习得到的生成预测。

01:24:46

This pseudocode explains the conceptual loop; it omits practical engineering such as batches, precision, and scheduling.

这段伪代码解释概念循环,省略了批处理、精度和调度等实际工程细节。

01:24:55

Its most important distinction is what information is available during learning versus use.

最重要的区别,是学习与使用时分别有哪些信息可用。

SLIDE 118 · 01:25:01 · Readings and technical extensions

Inference: where did the target go?

01:25:01

Look for the line that disappeared.

请找出消失的那一行。

01:25:03

The inference loop has no known edited target and no loss that updates weights.

推理循环没有已知编辑目标,也没有用于更新权重的损失。

01:25:08

Instead, we encode the source and instruction, begin with noise, repeatedly predict a velocity and move the latent, then decode the result.

我们编码原图和指令,从噪声开始,反复预测速度并移动潜变量,最后解码结果。

01:25:18

Training uses examples to adjust the learned model.

训练用样本调整学习到的模型。

01:25:21

In this simplified inference loop, we use that model to construct an output we do not yet possess.

在这个简化推理循环中,我们用该模型构造尚未拥有的输出。

01:25:28

The code is conceptual; practical systems use their own representations and sampling schedules.

代码是概念性的;实际系统使用各自的表示和采样调度。

01:25:35

But you can now explain why providing another reference changes the information available for a request without automatically becoming a model-training operation.

但你现在可以解释:增加参考图会改变请求可用的信息,却不会自动变成模型训练。

01:25:44

It is the same distinction we used when choosing between references and LoRA.

这与选择参考图还是 LoRA 时使用的区别相同。

SLIDE 119 · 01:25:50 · Readings and technical extensions

Choose the approach by the requirement

01:25:50

Before we turn to the readings' discussion questions, use this table as a compact decision aid.

转向阅读讨论问题之前,请把这张表当作简明决策辅助。

01:25:56

Exact untouched pixels suggest a role for masks and compositing.

如果未编辑像素必须精确不变,可以考虑蒙版与合成。

01:26:00

A new visual language may benefit from reference conditioning.

新的视觉语言,可能受益于参考图条件控制。

01:26:04

A reusable learned concept may motivate adaptation.

可复用的学习概念,可能需要模型适配。

01:26:08

A transformation that must survive motion needs a video workflow and temporal evidence.

必须经受运动的变换,需要视频工作流程和时间证据。

01:26:14

These approaches can cooperate in the same artwork.

这些方法可以在同一件作品中协作。

01:26:17

Choose according to the contract, then inspect the failure most likely to undermine it.

根据约定选择,再检查最可能破坏约定的失败。

01:26:23

That leaves us with a larger question for the readings: once production becomes easier, how do we choose a worthwhile direction and retain meaningful judgment over the result?

这给阅读留下更大的问题:制作更容易之后,如何选择值得追求的方向,并对结果保留有意义的判断?

SLIDE 120 · 01:26:34 · Closing reading discussions

An Alien Mind

01:26:34

An Alien Mind distinguishes achieving an assigned goal from generalizing human values in unfamiliar circumstances.

《An Alien Mind》区分了完成指定目标,与在陌生情境中泛化人类价值观。

01:26:41

Pachocki also raises questions about monitoring increasingly capable systems.

Pachocki 也提出了如何监测日益强大系统的问题。

01:26:47

Treat the essay as an argument containing claims and forecasts that we can examine.

请把文章视为包含主张与预测的论证,我们可以对其检视。

01:26:52

For our discussion, imagine an exhibition team delegating production while retaining responsibility for what the work communicates.

讨论时,想象一个展览团队委托制作,但仍对作品传达的内容负责。

01:27:00

Discuss the three questions on screen: explain the goal-and-values distinction, identify a forecast and the evidence it would need, and defend a boundary for human control.

讨论屏幕上的三个问题:解释目标与价值观的区别,找出一个预测及所需证据,并为人类控制的边界辩护。

01:27:12

Pause the video here.

请在这里暂停视频。

01:27:14

Our art-direction example is an analogy; it does not give an image editor the same agency or risk profile as an autonomous research system.

艺术指导只是类比;它并不赋予图像编辑器与自主研究系统相同的能动性或风险特征。

SLIDE 121 · 01:27:27 · Closing reading discussions

How to fix your entire life in 1 day

01:27:27

Koe's title makes a dramatic promise: How to fix your entire life in one day.

Koe 的标题作出了强烈承诺:《如何在一天内修复你的整个生活》。

01:27:32

The essay proposes examining identity and goals, interrupting habitual behavior, and turning reflection into action.

文章建议审视身份与目标、打断习惯行为,并把反思转化为行动。

01:27:41

We can discuss the usefulness of that proposal without accepting the title as a guarantee.

我们可以讨论这些建议是否有用,而不必把标题当作保证。

01:27:46

Imagine an artist who can generate a hundred attractive images but cannot choose what to make.

想象一个艺术家,能生成一百张好看的图像,却无法决定要创作什么。

01:27:52

Does easy production help clarify a direction, or make avoiding the decision easier?

轻松制作会帮助明确方向,还是让逃避决定更容易?

01:27:58

Discuss the three questions on screen.

请讨论屏幕上的三个问题。

01:28:01

Choose an idea worth using or challenging, explain the reason, and connect it to creative intention.

选一个值得采用或质疑的观点,解释理由,并把它与创作意图联系起来。

01:28:08

Pause the video here.

请在这里暂停视频。

01:28:10

You can use a fictional artist or public example; no personal disclosure is needed.

可以使用虚构艺术家或公开案例,不需要披露个人经历。

SLIDE 122 · 01:28:20 · Closing reading discussions

Introducing Aleph 2.0 and Edit Studio

01:28:20

Runway's reading proposes approving an edited image before applying its appearance through a video.

Runway 的阅读提出:先批准编辑图像,再把它的外观应用到整段视频。

01:28:26

It offers a concrete answer to a communication problem: an art director can point to the desired look, rather than describe every feature in words.

它具体回答了一个沟通问题:艺术总监可以指出目标外观,而不必用文字描述每个特征。

01:28:35

Now bring back the column.

现在,让柱子重新出现。

01:28:37

A convincing keyframe does not tell us what the pear will look like after it reappears.

可信关键帧并不能告诉我们,梨再次显露后会是什么样子。

01:28:42

Discuss the three questions on screen: what the frame establishes, which claim deserves a harder test, and how the workflow serves AFTER RAIN's intention.

讨论屏幕上的三个问题:这一帧确定了什么,哪个主张需要更严格测试,以及流程如何服务于《雨后》的意图。

01:28:52

Pause the video.

请暂停视频。

01:28:54

The source is a vendor announcement; our job is to distinguish a useful demonstrated workflow from a reliability claim still needing evaluation.

来源是厂商公告;我们的任务是区分有用的已演示流程,与仍需评价的可靠性主张。

SLIDE 123 · 01:29:07 · Closing reading discussions

Your final judgment

01:29:07

Return to your first judgment of the glass pear.

回到你最初对玻璃梨的判断。

01:29:11

We began with a small request and discovered that it touched the scene's physics, the artwork's intention, and the audience's experience over time.

我们从一个小要求出发,发现它触及场景物理、作品意图,以及观众随时间展开的体验。

01:29:20

You now have more precise ways to say what should change, what should survive, and how to judge the result.

现在,你有更精确的方法说明什么应改变、什么应保留,以及如何判断结果。

01:29:27

Finish with the three questions on screen.

最后,请讨论屏幕上的三个问题。

01:29:30

Where would you place the boundary between control and surprise?

你会把控制与惊喜之间的边界放在哪里?

01:29:33

How would you balance technical success and artistic purpose?

你会如何平衡技术成功与艺术目的?

01:29:37

What would you delegate, and what evidence would you require?

你会委托什么,又要求什么证据?

01:29:41

Pause the video for the final discussion.

请暂停视频,进行最终讨论。

01:29:44

Connect one mechanism or reading to a concrete artistic decision.

把一种机制或一篇阅读,与具体艺术决定联系起来。

01:29:49

Listen for an answer that makes you revise your own.

留意一个能让你修正自己看法的回答。

01:29:52

That revision is a fitting last act for a class about editing.

对一堂关于编辑的课来说,这样的修正,是恰当的最后一笔。