MiniMaxH3全参考模式改写输出格式指南

2026-08-19 MiniMaxH3

全参考模式改写输出格式指南

英语原文:https://github.com/MiniMax-AI/MiniMax-H3/blob/main/skills/h3-prompt-writing/references/ref-en.txt

本指南说明全参考模式下,改写输出内容的组织方式与撰写规范。

全部六个改写板块均使用英文撰写。仅保留 <d> 标签内的对话、歌词,以及场景中实际可见文字的原始语言。

描述细节要求detailed_description(详细描述)需要尽可能详尽、表述明确。对每一个镜头,清晰交代当前构图、主体外貌与位置、环境与光照、动作与状态变化、镜头运动、当下音效,以及参考内容实际出现、生效的位置。禁止把描述简化成剧情梗概,或是单纯罗列参考对应关系

镜头、镜头运动、说话人、对白、普通音效的基础格式,与《视频提示词撰写指南(T2VA / I2VA / FL2VA / L2VA)》保持一致。本指南重点讲解全参考模式特有的参考标签、分析板块以及格式差异。

1. 整体结构

一份完整改写输出按以下顺序包含六个板块:

板块 用途
subject_definitions 定义参考内容以及对应的参考标签
summary 概括任务类型、目标视频、主要参考对应关系
retention_analysis 说明参考内容如何被保留、迁移、复用
detailed_description 按照播放顺序描述画面、动作、镜头、音效、对白
overall_soundscape 概括环境音与物体动作音效
non_diegetic_music 描述仅观众能够听见的背景音乐

2. 参考标签与定义(subject_definitions

全参考改写使用四类标签,标记参考素材的来源与作用:

标签 含义
<Subject N> 从参考素材抽象出来的可视内容,可在目标视频中复用或修改
<Picture N> 参考图片,用作具体目标帧或是镜头规划锚点
<Video N> 参考视频,提供剪辑源、画面接续起点,或是整段视频的时序结构
<Audio N> 被复制或作为参考依据的音频信号

一份内容一旦分配参考标签,该标签含义在 subject_definitionssummaryretention_analysisdetailed_description 以及音频相关板块中全程保持不变。

subject_definitions 对每一项后续需要单独追踪的参考内容做定义,可以是人物、环境、源视频结构、音轨。每项单独占一行,说明标签指代对象、参考作用、需要遵循的核心特征;如果需要明确来源素材名称,则写明对应源文件。
如果 <Picture N> / <Video N> 仅用于定义其他参考对象,后续不会单独分析、调用,则在对应对象的定义里引用该标签,不再单独新开一行。
retention_analysis 记录每一项参考内容出现的位置,以及它是完整保留、部分保留、迁移使用还是复用。

2.1 <Subject N>

<Subject N> 用于可复用的可视内容,包括:

  • 人物、动物、物体
  • 场景、背景、环境
  • 服饰、道具、界面、视觉特效
  • 画风、动作、表情、姿态

它代表目标视频实际会使用的内容单元,不等同于源文件本身。一个主体可以由多份参考素材共同定义;一份参考素材也可以提供多个主体。

<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.

<主体1>为<图片1>中的年轻女性,黑色长发,蓝色开衫,佩戴细银项链。

同一主体来自多份素材时,合并来源,并写明每份素材提供什么信息:

<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.

<主体1>:人物外貌取自<图片1>,行走动作取自<视频1>的女性。

2.2 <Picture N>

当参考图片本身作为某镜头的首帧、关键帧、末帧、剪辑关键帧,或是构图锚点时,使用独立的 <Picture N>

<Picture 2> is the first frame of [Shot 1], showing a woman seated beside a café window.

<图片2>是[镜头1]的首帧,展示一名女子坐在咖啡馆窗边。

如果图片只用来定义人物、场景、服装、画风,不要新建独立图片条目,直接在对应 <Subject N> 的定义中引用图片来源。

图片用作分镜稿、镜头规划参考时,写明它对应哪些镜头,提供哪些规划信息:

<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.

<图片3>是[镜头1]、[镜头2]的分镜参考,规定镜头视角、主体摆放位置与镜头顺序。

2.3 <Video N>

<Video N> 专用于视频级整体关系,例如:

  • 对原始视频做剪辑
  • 从原始视频末尾接续生成新内容
  • 参考原视频的镜头运动、剪辑点、节奏、时序结构
<Video 1> is the source video for the target video edit.

<视频1>是目标视频剪辑所使用的源视频。

如果参考视频里的人物、物体、场景、动作、特效作为可视内容被复用,仍然归属于 <Subject N><Video N> 标记素材或结构来源,不能替代主体标签。

2.4 <Audio N>

<Audio N> 代表独立音频素材,或是参考视频中启用的同步音轨。常见用途:

  • 复制全部或部分音频信号
  • 参考背景音乐风格
  • 参考说话人的音色、演绎方式
  • 使用原音频中的对白、歌词、音效
  • 参考节拍、节奏、音频连贯性

<Audio N> 明确对应目标视频里的某个发声人时,定义中复用该发声人的全局ID:如果发声人已经定义为主体,则写 <Subject N> (Sx);否则使用一段稳定的声音描述,后跟 (Sx)。ID 遵循目标视频全局发声人编号顺序,不要在音频定义内独立分配、重新编号。发声人编号规则详见第5.4节。

<Audio 1> is the voice‑timbre reference for <Subject 1> (S1).

<音频1>是<主体1>(S1)的音色参考。

一份音频素材承担多重作用时,用一句通顺英文描述全部作用,不要额外拆分小节。

2.5 同一份参考视频的画面轨与音频轨

<Video N><Audio N> 的编号相互独立。编号只代表自身类别内的序号,不代表两类标签之间存在绑定配对。
因此同一份参考视频可以同时对应 <Video 1><Audio 2>;编号不同,不代表二者不能来自同一个源文件。

普通参考视频,不会仅仅因为文件自带声音就自动生成 <Audio N>
<Audio N> 的定义重点说明音频起到的作用,不必强制写明它来自哪一个 <Video N>。只有来源存在歧义时,才补充共同源素材信息,示例:

<Video 1> is the source video for the target video edit.
<Audio 2> is the synchronized audio track of <Video 1> and is reused in the target video.

<视频1>是目标视频剪辑的源视频。
<音频2>是<视频1>的同步音轨,在目标视频中复用。

3. summary(摘要)

本部分用简短英文段落概括目标视频与参考对应关系。开头使用方括号包裹的任务类型前缀。

[reference generation] …
[video editing + reference generation + audio reuse] …

根据参考素材在目标视频里的实际作用选择任务类型:

任务类型 使用场景
keyframe completion 图片作为目标视频首帧、关键帧、末帧、剪辑关键帧或其他具体帧锚点
reference generation 图片、视频、音频为人物、场景、画风、动作、镜头运动、分镜等提供生成引导,但不作为实际画面帧,也不是被剪辑/接续的源视频
video editing 直接修改一份已存在的源视频;编辑图片、静态关键帧之间生成内容不属于该类型
video continuation 在已有源视频的基础上接续、延伸、续写、转场生成新内容
audio reuse 完整或部分直接复用原有音频信号
audio reference 不直接复制音频信号,仅参考音乐风格、音色、对白歌词内容、音效质感、节拍、音频连贯性

当任务同时满足多种关系,使用 +拼接任务类型,不要重复。
例:从源视频接续,同时用一张图片作为末帧,写作 [video continuation + keyframe completion];剪辑源视频同时保留原音频,写作 [video editing + audio reuse]

仅仅存在视频、音频素材,不会自动生成对应任务类型。
如果参考视频只提供镜头运动、剪辑、节奏,一般归为 reference generation;只有对视频直接剪辑或接续续写时,才选用 video editing / video continuation
剪辑源视频,如果原音频保留播放,同时加上 audio reuse;接续源视频但不直接复制音频,新音频只沿袭原有声音特征,则使用 audio reference

摘要部分使用前面已经定义好的 <Subject N><Picture N><Video N><Audio N>,描述主要主体、镜头流程、参考素材作用。本部分不要引入新的参考标签

视频剪辑任务,任务类型前缀之后按如下格式开头:

The target video is an edited version of <Video 1>.

目标视频是<视频1>的剪辑版本。

4. retention_analysis(保留度分析)

该部分描述每一份参考内容在目标视频中如何被保留、迁移、复制、参考。每个参考标签占一行,遵循 subject_definitions 确立的含义。

4.1 可视内容

<Subject N><Picture N><Video N> 使用下面固定标记,输出必须为英文原关键字:

关系标记 含义
fully_preserved 参考内容的定义特征完整保留
partially_preserved 仍在使用参考内容,但部分定义特征被改动、仅部分留存
attribute_transfer 将参考特征迁移到另一个可识别的目标主体上
weak_reference 仅保留画风、类别、构图、氛围层面大致相似

主体条目示例:

<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - …

<主体1>(出现于[镜头1]、[镜头3]):完整保留 — …

图片条目示例:

<Picture 2> ([Shot 1] first frame): fully_preserved - …

<图片2>([镜头1]首帧):完整保留 — …

视频结构条目示例:

<Video 1> (cut and pacing structure): weak_reference - …

<视频1>(剪辑点与节奏结构):弱参考 — …

4.2 音频

<Audio N> 使用如下关系标记:

关系标记 含义
fully_copy 源音频完整作为目标视频最终音轨
partially_copy 仅复制部分时间片段或部分音频层;复制后新增、删减、替换部分声音
reference 不直接复制信号,仅参考音色、节奏、曲风、对白内容、音效质感
weak_reference 仅保留类别、氛围层面大致相似
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.

<音频1>:完整复制 — <音频1>原样复用,作为目标视频完整最终音轨。

<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.

<音频2>:参考 — 目标发声人沿用<音频2>的音色与沉稳演绎,不复制原始音频信号。

关系标记只能在该标签于subject_definitions中已经定义的参考作用范围内选用。目标视频新增的动作、背景、剧情事件,不视作参考保真度损失。

5. detailed_description(详细描述)

这是全参考改写的主体。按照目标视频播放顺序逐镜头描述画面、动作、音效、对白,在对应位置插入参考标签。

5.1 基础格式

基础格式遵循《视频提示词撰写指南(T2VA / I2VA / FL2VA / L2VA)》:

  • 主体内容使用英文;对白、歌词、画面可见文字保留原始语言。
  • [Shot 1]标记开篇镜头,不带时间戳;后续镜头格式:[Shot N] At MM:SS.mmm, …写明剪辑时刻。
  • 镜头运动自然写在镜头段落内部,需要说明时写明运动类型、幅度、速度。
  • 发声源分配稳定ID:(S1)(S2)以此类推;对白歌词格式 <d>[语言] …</d>
  • 对白跨剪辑、语音被视频末尾截断、音频跨镜头延续,使用 <scenetrans><cutoff>以及配套连续性描述。

镜头运动词汇、多人同时说话、画外音、跨剪辑对白、画面文字的完整规则与示例,请查阅《视频提示词撰写指南(T2VA / I2VA / FL2VA / L2VA)》。

5.2 全参考模式格式差异

维度 T2VA模式 全参考模式
主字段 integrated_multimodal_description detailed_description
画风说明位置 写在[Shot 1]之后 [Shot 1]之前,用一到两句英文说明整体画风
参考信息 不使用全参考标签 在首次出现以及生效位置插入 <Subject N><Picture N><Video N><Audio N>
音频相关描述 只描述目标视频自身声音 在对应镜头/音频段落引用<Audio N>,说明是复制还是参考

开篇示例:

The target video is in a cinematic, literary music‑video style with soft lighting and a slightly desaturated color palette.
[Shot 1] The scene opens in a crowded urban street…
[Shot 2] At 00:09.000, the shot cuts to an extreme close‑up…

目标视频为电影感文艺MV风格,柔和光照,色彩轻微低饱和。
[镜头1]场景开篇于一条喧闹城市街道……
[镜头2]在00:09.000时刻,镜头切至大特写……

生成类任务,detailed_description英文篇幅一般350‑500词。对白密集内容优先保证完整时序,不必机械凑字数。视频剪辑任务描述篇幅随源视频复杂度浮动,不受生成任务字数约束。不能因为只有单个镜头就简化描述;依据镜头信息体量分配细节。

5.3 在镜头描述内使用参考标签

重要 <Subject N> 第一次清晰出现时,结合镜头实际可视画面,描述它的参考特征、画面内位置、当前动作。后续镜头继续复用同一标签,不要再重复定义标签含义。

画面帧锚点使用自然句式:

the shot begins from <Picture 1>
the shot's keyframe corresponds to <Picture 2>
the shot ends on <Picture 3>

镜头起始画面来自<图片1>
镜头关键帧对应<图片2>
镜头结束定格于<图片3>

剪辑或接续源视频时,在涉及源状态、结构、接续逻辑的位置自然引用 <Video N>。音频生效的镜头或语义段落,引用 <Audio N>

5.4 发声人、音源与对白

发声人ID与 <d>基础格式和T2VA保持一致。如果被参考的主体实际开口说话,同时写上可视参考标签与发声人ID:

<Subject 2> (S1) turns toward the woman and says, <d>[English] Last summer, I went to my grandfather's house. He talked about you.</d>

<主体2>(S1)转向女子说道:<d>[英语] Last summer, I went to my grandfather's house. He talked about you.</d>

<Subject N>标记被参考的对象;(Sx)标记实际发声源。人物发声写 <Subject N> (Sx);同一主体画外发声,格式不变,标注 off‑screen。发声对象不属于已定义主体时,用一段稳定声音描述后跟 (Sx)

如果语音只是直接复用的BGM/完整音轨中的片段,不存在真人、角色、旁白这类独立发声源,则音源写 <Audio N>,不要凭空新增 (Sx)
如果声音来自具体人物、角色、旁白等独立发声主体,则分配并复用对应的 (Sx)

When <Audio 1> reaches the phrase <d>[English] I'm lonely lonely lonely lonely lonely I'm lonely</d>, <Subject 1> performs the corresponding hand gesture without becoming a separate speaker source.

当<音频1>播放到这句 <d>[英语] I'm lonely lonely lonely lonely lonely I'm lonely</d>,<主体1>做出配套手势,本身不作为独立发声源。

直接复用参考音频的对白、旁白、歌词,或是输入提示明确要求复刻演绎时,<d>内部完整保留原文文字与原始语言。听不清的片段写 [unclear],不要自行脑补改写。标点统一为基础书面标点 , . ? !;删除波浪线、emoji、项目符号、多余装饰符号。完整陈述句、问句、感叹句,在 </d>之前分别以 .?!结尾。

仅参考音色、节奏、情绪、演绎方式时,不要把参考音频中原对白带入目标视频。

(Sx)按照目标视频实际发声事件顺序一次性分配;在detailed_description每一处发声事件复用该ID。在subject_definitions绑定发声人的<Audio N>也复用同一个(Sx)不独立新建编号retention_analysis中不要写(Sx)
BGM、完整音轨自带的人声片段使用 <Audio N>;由真实人物、角色、旁白实体发出的声音使用 (Sx)

6. overall_soundscapenon_diegetic_music

两类音频字段定义与《视频提示词撰写指南(T2VA / I2VA / FL2VA / L2VA)》保持一致。

overall_soundscape概括全片环境音与物体音效。对白、演唱、绑定特定镜头的声音事件保留在detailed_description中,不在此处重复。

overall_soundscape: Quiet indoor room tone and a low ventilation hum continue throughout the video.

整体环境音:全程持续安静室内底噪与微弱通风嗡鸣。

non_diegetic_music描述角色听不到、仅观众可听见的背景音乐。存在配乐时写明乐器、速度、动态变化:

non_diegetic_music: A restrained solo‑piano score at a slow tempo, with sustained low cello underneath and no swell.

非画内背景音乐:克制的慢节奏钢琴独奏,底层铺绵长低沉大提琴,无音量爆发。

使用参考音频时,复制/参考关系写在对应音频层所属板块:环境音、音效写进overall_soundscape;仅观众收听的配乐写进non_diegetic_music。同一份音频同时包含两类内容,则分别在两个板块描述对应关系。

overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.
non_diegetic_music: <Audio 2> is directly reused as the complete audience‑only score.

整体环境音:取自<音频1>的环境音层复制后贯穿目标视频全程。
非画内背景音乐:<音频2>直接复用,作为完整观众向配乐。

完整对白、歌词只能放在detailed_description<d>标签内,不要在这两个音频板块重复。

7. 完整示例

subject_definitions:
<Subject 1> is the coffee‑shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.
<Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.
<Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light‑pink button‑down shirt with rolled‑up sleeves.
<Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark‑grey hoodie with drawstrings.
<Audio 1> is the voice‑timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.

summary:
[reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three‑shot exchange uses <Audio 1> as the voice‑timbre reference for <Subject 3> and ends with a canned audience laugh.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.
<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light‑pink shirt are retained.
<Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark‑grey hoodie are retained.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.

detailed_description:
The target video uses a realistic multi‑camera sitcom style with warm indoor lighting.
[Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light‑pink button‑down shirt with rolled‑up sleeves, sits on the sofa holding a chocolate‑chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark‑grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick‑furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
[Shot 2] At 00:03.000, the shot cuts to a close‑up of <Subject 4> (S2), the young man in the dark‑grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.
[Shot 3] At 00:05.000, the shot cuts to a close‑up of <Subject 3> (S1), the blonde woman in the light‑pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast‑like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.

overall_soundscape:
Soft indoor coffee‑shop room tone continues throughout the scene.

non_diegetic_music:
N/A