MiniMaxH3全参考模式改写输出格式指南
全参考模式改写输出格式指南
英语原文:https://github.com/MiniMax-AI/MiniMax-H3/blob/main/skills/h3-prompt-writing/references/ref-en.txt
本指南说明全参考模式下,改写输出内容的组织方式与撰写规范。
全部六个改写板块均使用英文撰写。仅保留 <d> 标签内的对话、歌词,以及场景中实际可见文字的原始语言。
描述细节要求:detailed_description(详细描述)需要尽可能详尽、表述明确。对每一个镜头,清晰交代当前构图、主体外貌与位置、环境与光照、动作与状态变化、镜头运动、当下音效,以及参考内容实际出现、生效的位置。禁止把描述简化成剧情梗概,或是单纯罗列参考对应关系。
镜头、镜头运动、说话人、对白、普通音效的基础格式,与《视频提示词撰写指南(T2VA / I2VA / FL2VA / L2VA)》保持一致。本指南重点讲解全参考模式特有的参考标签、分析板块以及格式差异。
1. 整体结构
一份完整改写输出按以下顺序包含六个板块:
| 板块 | 用途 |
|---|---|
subject_definitions |
定义参考内容以及对应的参考标签 |
summary |
概括任务类型、目标视频、主要参考对应关系 |
retention_analysis |
说明参考内容如何被保留、迁移、复用 |
detailed_description |
按照播放顺序描述画面、动作、镜头、音效、对白 |
overall_soundscape |
概括环境音与物体动作音效 |
non_diegetic_music |
描述仅观众能够听见的背景音乐 |
2. 参考标签与定义(subject_definitions)
全参考改写使用四类标签,标记参考素材的来源与作用:
| 标签 | 含义 |
|---|---|
<Subject N> |
从参考素材抽象出来的可视内容,可在目标视频中复用或修改 |
<Picture N> |
参考图片,用作具体目标帧或是镜头规划锚点 |
<Video N> |
参考视频,提供剪辑源、画面接续起点,或是整段视频的时序结构 |
<Audio N> |
被复制或作为参考依据的音频信号 |
一份内容一旦分配参考标签,该标签含义在
subject_definitions、summary、retention_analysis、detailed_description以及音频相关板块中全程保持不变。
subject_definitions 对每一项后续需要单独追踪的参考内容做定义,可以是人物、环境、源视频结构、音轨。每项单独占一行,说明标签指代对象、参考作用、需要遵循的核心特征;如果需要明确来源素材名称,则写明对应源文件。
如果 <Picture N> / <Video N> 仅用于定义其他参考对象,后续不会单独分析、调用,则在对应对象的定义里引用该标签,不再单独新开一行。
retention_analysis 记录每一项参考内容出现的位置,以及它是完整保留、部分保留、迁移使用还是复用。
2.1 <Subject N>
<Subject N> 用于可复用的可视内容,包括:
- 人物、动物、物体
- 场景、背景、环境
- 服饰、道具、界面、视觉特效
- 画风、动作、表情、姿态
它代表目标视频实际会使用的内容单元,不等同于源文件本身。一个主体可以由多份参考素材共同定义;一份参考素材也可以提供多个主体。
<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.
<主体1>为<图片1>中的年轻女性,黑色长发,蓝色开衫,佩戴细银项链。
同一主体来自多份素材时,合并来源,并写明每份素材提供什么信息:
<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.
<主体1>:人物外貌取自<图片1>,行走动作取自<视频1>的女性。
2.2 <Picture N>
当参考图片本身作为某镜头的首帧、关键帧、末帧、剪辑关键帧,或是构图锚点时,使用独立的 <Picture N>:
<Picture 2> is the first frame of [Shot 1], showing a woman seated beside a café window.
<图片2>是[镜头1]的首帧,展示一名女子坐在咖啡馆窗边。
如果图片只用来定义人物、场景、服装、画风,不要新建独立图片条目,直接在对应 <Subject N> 的定义中引用图片来源。
图片用作分镜稿、镜头规划参考时,写明它对应哪些镜头,提供哪些规划信息:
<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.
<图片3>是[镜头1]、[镜头2]的分镜参考,规定镜头视角、主体摆放位置与镜头顺序。
2.3 <Video N>
<Video N> 专用于视频级整体关系,例如:
- 对原始视频做剪辑
- 从原始视频末尾接续生成新内容
- 参考原视频的镜头运动、剪辑点、节奏、时序结构
<Video 1> is the source video for the target video edit.
<视频1>是目标视频剪辑所使用的源视频。
如果参考视频里的人物、物体、场景、动作、特效作为可视内容被复用,仍然归属于 <Subject N>。<Video N> 标记素材或结构来源,不能替代主体标签。
2.4 <Audio N>
<Audio N> 代表独立音频素材,或是参考视频中启用的同步音轨。常见用途:
- 复制全部或部分音频信号
- 参考背景音乐风格
- 参考说话人的音色、演绎方式
- 使用原音频中的对白、歌词、音效
- 参考节拍、节奏、音频连贯性
当 <Audio N> 明确对应目标视频里的某个发声人时,定义中复用该发声人的全局ID:如果发声人已经定义为主体,则写 <Subject N> (Sx);否则使用一段稳定的声音描述,后跟 (Sx)。ID 遵循目标视频全局发声人编号顺序,不要在音频定义内独立分配、重新编号。发声人编号规则详见第5.4节。
<Audio 1> is the voice‑timbre reference for <Subject 1> (S1).
<音频1>是<主体1>(S1)的音色参考。
一份音频素材承担多重作用时,用一句通顺英文描述全部作用,不要额外拆分小节。
2.5 同一份参考视频的画面轨与音频轨
<Video N> 和 <Audio N> 的编号相互独立。编号只代表自身类别内的序号,不代表两类标签之间存在绑定配对。
因此同一份参考视频可以同时对应 <Video 1> 和 <Audio 2>;编号不同,不代表二者不能来自同一个源文件。
普通参考视频,不会仅仅因为文件自带声音就自动生成 <Audio N>。
<Audio N> 的定义重点说明音频起到的作用,不必强制写明它来自哪一个 <Video N>。只有来源存在歧义时,才补充共同源素材信息,示例:
<Video 1> is the source video for the target video edit.
<Audio 2> is the synchronized audio track of <Video 1> and is reused in the target video.
<视频1>是目标视频剪辑的源视频。
<音频2>是<视频1>的同步音轨,在目标视频中复用。
3. summary(摘要)
本部分用简短英文段落概括目标视频与参考对应关系。开头使用方括号包裹的任务类型前缀。
[reference generation] …
[video editing + reference generation + audio reuse] …
根据参考素材在目标视频里的实际作用选择任务类型:
| 任务类型 | 使用场景 |
|---|---|
keyframe completion |
图片作为目标视频首帧、关键帧、末帧、剪辑关键帧或其他具体帧锚点 |
reference generation |
图片、视频、音频为人物、场景、画风、动作、镜头运动、分镜等提供生成引导,但不作为实际画面帧,也不是被剪辑/接续的源视频 |
video editing |
直接修改一份已存在的源视频;编辑图片、静态关键帧之间生成内容不属于该类型 |
video continuation |
在已有源视频的基础上接续、延伸、续写、转场生成新内容 |
audio reuse |
完整或部分直接复用原有音频信号 |
audio reference |
不直接复制音频信号,仅参考音乐风格、音色、对白歌词内容、音效质感、节拍、音频连贯性 |
当任务同时满足多种关系,使用 +拼接任务类型,不要重复。
例:从源视频接续,同时用一张图片作为末帧,写作 [video continuation + keyframe completion];剪辑源视频同时保留原音频,写作 [video editing + audio reuse]。
仅仅存在视频、音频素材,不会自动生成对应任务类型。
如果参考视频只提供镜头运动、剪辑、节奏,一般归为 reference generation;只有对视频直接剪辑或接续续写时,才选用 video editing / video continuation。
剪辑源视频,如果原音频保留播放,同时加上 audio reuse;接续源视频但不直接复制音频,新音频只沿袭原有声音特征,则使用 audio reference。
摘要部分使用前面已经定义好的 <Subject N>、<Picture N>、<Video N>、<Audio N>,描述主要主体、镜头流程、参考素材作用。本部分不要引入新的参考标签。
视频剪辑任务,任务类型前缀之后按如下格式开头:
The target video is an edited version of <Video 1>.
目标视频是<视频1>的剪辑版本。
4. retention_analysis(保留度分析)
该部分描述每一份参考内容在目标视频中如何被保留、迁移、复制、参考。每个参考标签占一行,遵循 subject_definitions 确立的含义。
4.1 可视内容
<Subject N>、<Picture N>、<Video N> 使用下面固定标记,输出必须为英文原关键字:
| 关系标记 | 含义 |
|---|---|
fully_preserved |
参考内容的定义特征完整保留 |
partially_preserved |
仍在使用参考内容,但部分定义特征被改动、仅部分留存 |
attribute_transfer |
将参考特征迁移到另一个可识别的目标主体上 |
weak_reference |
仅保留画风、类别、构图、氛围层面大致相似 |
主体条目示例:
<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - …
<主体1>(出现于[镜头1]、[镜头3]):完整保留 — …
图片条目示例:
<Picture 2> ([Shot 1] first frame): fully_preserved - …
<图片2>([镜头1]首帧):完整保留 — …
视频结构条目示例:
<Video 1> (cut and pacing structure): weak_reference - …
<视频1>(剪辑点与节奏结构):弱参考 — …
4.2 音频
<Audio N> 使用如下关系标记:
| 关系标记 | 含义 |
|---|---|
fully_copy |
源音频完整作为目标视频最终音轨 |
partially_copy |
仅复制部分时间片段或部分音频层;复制后新增、删减、替换部分声音 |
reference |
不直接复制信号,仅参考音色、节奏、曲风、对白内容、音效质感 |
weak_reference |
仅保留类别、氛围层面大致相似 |
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
<音频1>:完整复制 — <音频1>原样复用,作为目标视频完整最终音轨。
<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.
<音频2>:参考 — 目标发声人沿用<音频2>的音色与沉稳演绎,不复制原始音频信号。
关系标记只能在该标签于subject_definitions中已经定义的参考作用范围内选用。目标视频新增的动作、背景、剧情事件,不视作参考保真度损失。
5. detailed_description(详细描述)
这是全参考改写的主体。按照目标视频播放顺序逐镜头描述画面、动作、音效、对白,在对应位置插入参考标签。
5.1 基础格式
基础格式遵循《视频提示词撰写指南(T2VA / I2VA / FL2VA / L2VA)》:
- 主体内容使用英文;对白、歌词、画面可见文字保留原始语言。
[Shot 1]标记开篇镜头,不带时间戳;后续镜头格式:[Shot N] At MM:SS.mmm, …写明剪辑时刻。- 镜头运动自然写在镜头段落内部,需要说明时写明运动类型、幅度、速度。
- 发声源分配稳定ID:
(S1)、(S2)以此类推;对白歌词格式<d>[语言] …</d>。 - 对白跨剪辑、语音被视频末尾截断、音频跨镜头延续,使用
<scenetrans>、<cutoff>以及配套连续性描述。
镜头运动词汇、多人同时说话、画外音、跨剪辑对白、画面文字的完整规则与示例,请查阅《视频提示词撰写指南(T2VA / I2VA / FL2VA / L2VA)》。
5.2 全参考模式格式差异
| 维度 | T2VA模式 | 全参考模式 |
|---|---|---|
| 主字段 | integrated_multimodal_description |
detailed_description |
| 画风说明位置 | 写在[Shot 1]之后 |
在[Shot 1]之前,用一到两句英文说明整体画风 |
| 参考信息 | 不使用全参考标签 | 在首次出现以及生效位置插入 <Subject N>、<Picture N>、<Video N>、<Audio N> |
| 音频相关描述 | 只描述目标视频自身声音 | 在对应镜头/音频段落引用<Audio N>,说明是复制还是参考 |
开篇示例:
The target video is in a cinematic, literary music‑video style with soft lighting and a slightly desaturated color palette.
[Shot 1] The scene opens in a crowded urban street…
[Shot 2] At 00:09.000, the shot cuts to an extreme close‑up…
目标视频为电影感文艺MV风格,柔和光照,色彩轻微低饱和。
[镜头1]场景开篇于一条喧闹城市街道……
[镜头2]在00:09.000时刻,镜头切至大特写……
生成类任务,detailed_description英文篇幅一般350‑500词。对白密集内容优先保证完整时序,不必机械凑字数。视频剪辑任务描述篇幅随源视频复杂度浮动,不受生成任务字数约束。不能因为只有单个镜头就简化描述;依据镜头信息体量分配细节。
5.3 在镜头描述内使用参考标签
重要 <Subject N> 第一次清晰出现时,结合镜头实际可视画面,描述它的参考特征、画面内位置、当前动作。后续镜头继续复用同一标签,不要再重复定义标签含义。
画面帧锚点使用自然句式:
the shot begins from <Picture 1>
the shot's keyframe corresponds to <Picture 2>
the shot ends on <Picture 3>
镜头起始画面来自<图片1>
镜头关键帧对应<图片2>
镜头结束定格于<图片3>
剪辑或接续源视频时,在涉及源状态、结构、接续逻辑的位置自然引用 <Video N>。音频生效的镜头或语义段落,引用 <Audio N>。
5.4 发声人、音源与对白
发声人ID与 <d>基础格式和T2VA保持一致。如果被参考的主体实际开口说话,同时写上可视参考标签与发声人ID:
<Subject 2> (S1) turns toward the woman and says, <d>[English] Last summer, I went to my grandfather's house. He talked about you.</d>
<主体2>(S1)转向女子说道:<d>[英语] Last summer, I went to my grandfather's house. He talked about you.</d>
<Subject N>标记被参考的对象;(Sx)标记实际发声源。人物发声写 <Subject N> (Sx);同一主体画外发声,格式不变,标注 off‑screen。发声对象不属于已定义主体时,用一段稳定声音描述后跟 (Sx)。
如果语音只是直接复用的BGM/完整音轨中的片段,不存在真人、角色、旁白这类独立发声源,则音源写 <Audio N>,不要凭空新增 (Sx)。
如果声音来自具体人物、角色、旁白等独立发声主体,则分配并复用对应的 (Sx)。
When <Audio 1> reaches the phrase <d>[English] I'm lonely lonely lonely lonely lonely I'm lonely</d>, <Subject 1> performs the corresponding hand gesture without becoming a separate speaker source.
当<音频1>播放到这句 <d>[英语] I'm lonely lonely lonely lonely lonely I'm lonely</d>,<主体1>做出配套手势,本身不作为独立发声源。
直接复用参考音频的对白、旁白、歌词,或是输入提示明确要求复刻演绎时,<d>内部完整保留原文文字与原始语言。听不清的片段写 [unclear],不要自行脑补改写。标点统一为基础书面标点 , . ? !;删除波浪线、emoji、项目符号、多余装饰符号。完整陈述句、问句、感叹句,在 </d>之前分别以 .、?、!结尾。
仅参考音色、节奏、情绪、演绎方式时,不要把参考音频中原对白带入目标视频。
(Sx)按照目标视频实际发声事件顺序一次性分配;在detailed_description每一处发声事件复用该ID。在subject_definitions绑定发声人的<Audio N>也复用同一个(Sx),不独立新建编号。retention_analysis中不要写(Sx)。
BGM、完整音轨自带的人声片段使用 <Audio N>;由真实人物、角色、旁白实体发出的声音使用 (Sx)。
6. overall_soundscape与non_diegetic_music
两类音频字段定义与《视频提示词撰写指南(T2VA / I2VA / FL2VA / L2VA)》保持一致。
overall_soundscape概括全片环境音与物体音效。对白、演唱、绑定特定镜头的声音事件保留在detailed_description中,不在此处重复。
overall_soundscape: Quiet indoor room tone and a low ventilation hum continue throughout the video.
整体环境音:全程持续安静室内底噪与微弱通风嗡鸣。
non_diegetic_music描述角色听不到、仅观众可听见的背景音乐。存在配乐时写明乐器、速度、动态变化:
non_diegetic_music: A restrained solo‑piano score at a slow tempo, with sustained low cello underneath and no swell.
非画内背景音乐:克制的慢节奏钢琴独奏,底层铺绵长低沉大提琴,无音量爆发。
使用参考音频时,复制/参考关系写在对应音频层所属板块:环境音、音效写进overall_soundscape;仅观众收听的配乐写进non_diegetic_music。同一份音频同时包含两类内容,则分别在两个板块描述对应关系。
overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.
non_diegetic_music: <Audio 2> is directly reused as the complete audience‑only score.
整体环境音:取自<音频1>的环境音层复制后贯穿目标视频全程。
非画内背景音乐:<音频2>直接复用,作为完整观众向配乐。
完整对白、歌词只能放在detailed_description的<d>标签内,不要在这两个音频板块重复。
7. 完整示例
subject_definitions:
<Subject 1> is the coffee‑shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.
<Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.
<Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light‑pink button‑down shirt with rolled‑up sleeves.
<Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark‑grey hoodie with drawstrings.
<Audio 1> is the voice‑timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.
summary:
[reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three‑shot exchange uses <Audio 1> as the voice‑timbre reference for <Subject 3> and ends with a canned audience laugh.
retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.
<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light‑pink shirt are retained.
<Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark‑grey hoodie are retained.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
The target video uses a realistic multi‑camera sitcom style with warm indoor lighting.
[Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light‑pink button‑down shirt with rolled‑up sleeves, sits on the sofa holding a chocolate‑chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark‑grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick‑furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
[Shot 2] At 00:03.000, the shot cuts to a close‑up of <Subject 4> (S2), the young man in the dark‑grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.
[Shot 3] At 00:05.000, the shot cuts to a close‑up of <Subject 3> (S1), the blonde woman in the light‑pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast‑like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.
overall_soundscape:
Soft indoor coffee‑shop room tone continues throughout the scene.
non_diegetic_music:
N/A