MiniMaxH3视频提示词撰写指南(T2VA、I2VA、FL2VA、L2VA)

2026-08-18 MiniMaxH3,提示词

英语原文:https://github.com/MiniMax-AI/MiniMax-H3/blob/main/skills/h3-prompt-writing/references/base-en.txt

1. 任务概述

  • T2VA(文本生成音视频):基于文本构建完整的音视频时间线。
  • I2VA(图起始生成音视频):T2VA主体内容 + 首帧指令 + 从首帧向前推演发展的视觉路径。
  • FL2VA(首尾帧生成音视频):T2VA主体内容 + 首末帧指令 + 从首帧延续至末帧的连贯画面路径。
  • L2VA(末帧生成音视频):T2VA主体内容 + 末帧指令 + 从合理前置状态逐步收敛到末帧的画面路径。

2. 最终提示词结构

2.1 第一部分:指令部分

T2VA 无图像对齐指令,直接以三大核心字段开头。

I2VA 固定使用如下语句:

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

译文:对于目标视频,在目标视频0.00秒时刻,完整参考<图片1>(来自[镜头1])。

FL2VA 固定使用如下语句:

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00‑second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS‑second mark of the target video.

译文:参考图片与目标视频的对齐规则——图片1(来自镜头1)对齐目标视频0.00秒时间点;图片2(来自镜头N)对齐目标视频S.SS秒时间点。

L2VA 固定使用如下语句:

How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS‑second mark of the target video.

译文:参考图片与目标视频的对齐规则——<图片1>(来自镜头N)对齐目标视频S.SS秒时间点。

其中,N代表实际最后一个镜头的编号;S.SS为视频有效时长,格式严格保留两位小数。指令必须放在最终提示词的第一行,指令结束后空一行,再写核心字段。

2.2 第二部分:三大核心字段

integrated_multimodal_description: [Shot 1] …

overall_soundscape: …

non_diegetic_music: …
  • integrated_multimodal_description(综合多模态描述):按时间线描述画面、动作、镜头、说话人、对话、演唱以及画内音。
  • overall_soundscape(整体环境音):概括全视频的环境音、物体动作音效、人类非语言发声。
  • non_diegetic_music(非画内背景音乐):描述角色听不到、仅观众能够听见的背景音乐。

3. 如何将关键帧融入多模态描述

3.1 I2VA:从参考图片起始,向前推演画面

<Picture 1>对应视频0.00秒的真实首帧,归属[Shot 1]。描述需要先确立参考图中的画风、主体、构图、场景基准,再续写后续动作。人物身份、服饰、色彩、关键物体、空间位置关系全程保持一致。

推荐行文结构:首帧基准锚定 → 动作启动 → 画面连续发展 → 结果或人物反应

3.2 FL2VA:描述首帧到末帧之间的演变路径

图片1是视频开篇,图片2是视频结尾。重点描述主体如何运动、姿态如何变化、物体如何被操作、构图如何演变、场景与光影如何过渡。

FL2VA一般优先使用单个镜头,便于模型在首帧与末帧之间做连续插值;只有任务明确要求时,才使用多镜头。视频末尾的最后一个镜头[Shot N]必须抵达末帧画面。

推荐行文结构:首帧状态 → 可观测的中间变化 → 差异逐步缩小 → 末帧最终状态

3.3 L2VA:推演开篇画面,最终落定到参考末帧

<Picture 1>是视频的最后一帧,归属最后一个镜头[Shot N],它不属于镜头1。结合用户意图和末帧,推演一套逻辑合理的前期画面状态,再描述人物、物体、镜头、场景如何逐步过渡逼近参考图片。

推荐行文结构:合理的前置画面状态 → 明确的动作与转场路径 → 末段镜头逐步收敛 → 最终定格到末帧画面

4. 三大通用核心板块撰写方法

4.1 沿着时间线撰写多模态描述

integrated_multimodal_description是改写后提示词的主体。所有描述内容都必须对应可视画面或可听见的声音:画面风格、初始构图、主体外观与位置、场景与关键道具、动作与反应、镜头切换、人物对白、同步画内音效。

[Shot 1]开头写明整体画风与初始构图。常见画风:Cinematic(电影质感)live‑action(实拍真人)2D‑animated(二维动画)3D CG(三维计算机渲染)claymation(黏土动画)watercolor(水彩)vintage film(复古胶片)。关键帧任务从参考图片提取画风;T2VA任务从用户输入文本选取画风。

示例:

[Shot 1] Live‑action, cinematic, a medium‑wide shot frames…

[镜头1] 实拍真人,电影质感,中宽景框住……

4.2 镜头与剪辑

第一个镜头不要写时间戳。后续镜头按顺序编号,每个镜头开头写明严格递增的剪辑时间,时间需落在视频总时长范围内。

[Shot 2] At 00:03.500, the camera cuts to…

[镜头2] 在00:03.500时刻,镜头切至……

普通硬切可使用:the camera cuts tothe shot cuts tothe shot transitions tothe shot changes tothe shot switches to。用户明确要求时,可以使用叠化、淡入淡出、划像等转场。每一次剪辑需要带来主体、空间、状态、视角、时间上的新信息。如果只需要改变镜头距离或微小角度,优先使用镜头运动,不要做剪辑。

4.3 镜头运动:运动类型 + 幅度 + 速度

完整镜头运动描述包含三要素:运动类型定义镜头如何移动;幅度定义构图变化范围;速度定义画面变化节奏。只有存在明显区分度时才补充幅度与速度;中等幅度、正常速度通常省略不写。

要素 可用表达式 说明
运动类型 Zoom In / Zoom Out 机位不动,焦距放大/缩小
运动类型 Push In / Pull Out 摄像机向前推进 / 向后拉远
运动类型 Pan Left / Pan Right 机位固定,镜头水平左右摇
运动类型 Truck Left / Truck Right 摄像机整体水平平移
运动类型 Tilt Up / Tilt Down 机位固定,镜头垂直上下俯仰
运动类型 Pedestal Up / Pedestal Down 摄像机整体向上 / 向下升降
运动类型 Arc Shot 摄像机环绕主体做弧形运动
运动类型 Tracking Shot 镜头跟随运动主体拍摄
运动类型 Static Shot 机位与镜头保持静止不动
运动类型 Shake Slightly / Shake Strongly 轻微抖动 / 强烈抖动
运动类型 POV 第一人称主观视角
运动类型 Roll Clockwise / Roll Counterclockwise 镜头绕光轴顺时针 / 逆时针旋转
幅度 with small amplitude 小幅变化
幅度 with large amplitude 大幅变化
速度 at slow speed 慢速运动
速度 at fast speed 快速运动

镜头运动要自然嵌入镜头描述的语句之中,不要在句尾堆砌标签。

The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.

摄像机小幅慢速向她手中折好的信件推近。

The camera pans right with large amplitude at fast speed, revealing the open doorway.

摄像机大幅快速向右摇镜,露出敞开的门口。

The camera holds a static shot as the runner exits the frame.

跑步者跑出画面,镜头保持静止。

4.4 说话人、对话与演唱

说话、歌唱、画外发声的人物使用固定编号标识,例如(S1)(S2)。多名已编号人物同时发声,使用复合编号,如(S1,S2)。同一个人物在不同镜头下保持同一个编号;全程不发声的角色不分配说话人编号。

人物首次出现时,结合画面、声音信息交代人物特征:人物类型、年龄、性别、是否出镜、音调、音色、语速、口音。人物特征描述、编号、动作、说话状态写在<d>标签外部;<d>内部只保留语言标记和原始台词。原文字词、标点原样保留,不要翻译、改写。

The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>

嗓音轻柔气声的年轻女子(S1)说道:<d>[英语] I get off at the next station.</d>

The two children (S1,S2) shout together, <d>[English] Wait for us!</d>

两个孩子(S1,S2)一同大喊:<d>[英语] Wait for us!</d>

画外音固定使用短语 says in an off‑screen voiceover。每一段画外音<d>之后,必须补充说明画面对应角色嘴唇没有开合。

The man (S1) says in an off‑screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.

男子(S1)以画外音叙述:<d>[英语] I still remember that road.</d>,同时他的嘴唇完全闭合。

如果同一段对白、歌词跨越剪辑点,两段描述的衔接位置都加上<scenetrans>,明确说明音频跨剪辑延续。如果语音在视频结束处被截断,使用<cutoff>。描述音频延续可用短语:continues seamlessly across the cut(剪辑处无缝延续)continues uninterrupted into the next shot(无中断延续到下一个镜头)carries over from the previous shot(承接上一镜头)remains audible across the transition(转场过程持续可闻)

4.5 屏幕内文字

画面中实际可见横幅、标牌、标签、字幕、霓虹文字,全部放在英文双引号内。原文文字、标点原样保留,不做翻译。

A red neon sign reading "营业中" glows above the doorway.

门口上方,一块写着“营业中”的红色霓虹灯牌散发光芒。

4.6 overall_soundscape(整体环境音)

使用1‑4句英文,写成一个完整段落,概括全片环境音、物体动作音效、人类非语言声音,例如风声、雨声、车流、脚步声、衣物摩擦、撞击、呼吸、笑声、喘息。对白、演唱、画内音乐写在多模态描述中,此处不要重复。只有用户明确要求全程完全静音时,才填写N/A

示例:

overall_soundscape: Steady rain taps against the café windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair.

整体环境音:持续的雨点敲打咖啡馆玻璃窗,背景伴随微弱室内环境底噪。门口铃铛响一声,随后传来湿漉漉的脚步声与椅子轻微摩擦声。

4.7 non_diegetic_music(非画内背景音乐)

使用1‑3句英文描述角色听不见、仅观众听到的背景音乐。重点写乐器、速度、节奏、音量动态变化,不要使用抽象情绪形容词,不要解释配乐的情感作用。角色能够听见的歌声、乐器、收音机、电视、手机音乐都属于画内音,放在多模态描述部分。无背景音乐填写N/A

示例:

non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.

非画内背景音乐:缓慢节奏下稀疏的钢琴音符,加入绵长低沉弦乐,音量逐步抬升后缓缓淡出。

5. 示例样例

样例1:T2VA

无参考图,直接基于文本搭建完整时间线。可以补充符合用户意图的场景、人物、动作、音效细节。

integrated_multimodal_description: [Shot 1] Live‑action, cinematic, a medium‑wide shot frames a baker opening the shutters of a small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as the middle‑aged baker with a calm, slightly raspy voice (S1) places a fresh loaf on the wooden counter and says: <d>[English] First batch of the morning.</d> [Shot 2] At 00:05.000, the camera cuts to a close‑up of steam rising from the sliced bread while the baker's final words carry over from the previous shot.

overall_soundscape: Wooden shutters scrape open over a quiet street as trays clink softly inside the bakery. The doorbell rings once, followed by light footsteps and the crisp sound of bread being sliced.

non_diegetic_music: A soft acoustic‑guitar pattern at a moderate tempo, joined by sparse upright‑bass notes and a gentle fade at the end.

综合多模态描述:[镜头1]实拍真人,电影质感,中宽镜头拍摄日出之前面包店店主打开街边小面包店的百叶窗。镜头小幅慢速向前推近,嗓音沉稳略带沙哑的中年面包师(S1)将一条新鲜面包摆上木质柜台,说道:<d>[英语] First batch of the morning.</d>。[镜头2]在00:05.000时刻,镜头切到切片面包升腾热气的特写,面包师最后的台词承接上一镜头延续播放。

整体环境音:寂静街道上木质百叶窗摩擦开启,店内托盘轻轻碰撞。门铃响一次,接着传来轻快脚步声以及面包切片的清脆声响。

非画内背景音乐:节奏舒缓柔和的原声吉他旋律,搭配稀疏低音提琴音符,结尾缓缓淡出。

样例2:I2VA

先写首帧对齐指令,把图片1的主体、构图、场景作为镜头1的起点,之后续写画面演变。

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live‑action, cinematic, the young woman shown in <Picture 1> remains beside the rain‑covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: <d>[English] I get off at the next station.</d> She folds the letter along its existing crease.

overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands.

non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.

对于目标视频,在目标视频0.00秒时刻,完整参考<图片1>(来自[镜头1])。

综合多模态描述:[镜头1]实拍真人,电影质感,<图片1>中的年轻女子待在布满雨珠的列车窗边,人物样貌、衣着、座位位置、车厢环境全部保留。镜头小幅慢速水平右移,她目光从折好的信件抬起,望向窗外掠过的城市灯火。玻璃上浮动着她的倒影,嗓音轻柔带气声的年轻女子(S1)说道:<d>[英语] I get off at the next station.</d>,顺着原有折痕把信件再次对折。

整体环境音:列车车轮持续发出金属行进节奏,叠加微弱通风系统嗡鸣。雨点敲打车窗,手中纸张发出轻微沙沙声。

非画内背景音乐:绵长大提琴长音搭配间隔疏朗的钢琴音符,节奏缓慢,音量逐步降低。

样例3:FL2VA

两张图片分别锁定开篇与结尾。正文不要重复描述两张静态图片,重点补充连接二者的运动路径。本例为8秒单镜头。

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00‑second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00‑second mark of the target video.

integrated_multimodal_description: [Shot 1] Live‑action, cinematic, a rain‑soaked cyclist begins in the position and framing established by Picture 1, holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the bicycle handle, raises the umbrella above her shoulder, and presses the runner upward until the canopy opens. Water rolls from the expanding fabric while she steps beneath it, rotates the handle into the final angle, and settles into the pose, spacing, and composition established by Picture 2 at the end of the shot.

overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes.

non_diegetic_music: N/A

参考图片与目标视频的对齐规则——图片1(来自镜头1)对齐目标视频0.00秒时间点;图片2(来自镜头1)对齐目标视频8.00秒时间点。

综合多模态描述:[镜头1]实拍真人,电影质感,浑身淋雨的骑行者起始姿态与取景完全和图片1一致,手持收拢的黑色雨伞,站在银色自行车旁。镜头小幅慢速向后拉远;她松开自行车车把,将雨伞举到肩头上方,向上推动伞滑套,直至伞面撑开。雨水顺着张开的伞布滑落,她移步站到伞下,转动伞柄调整到最终角度,镜头末尾人物姿态、空间距离、画面构图定格为图片2。

整体环境音:雨水持续打在路面,接着传来伞滑套金属咔哒声与伞面弹开的轻响。水珠从车架滴落,远处有车流驶过。

非画内背景音乐:无。

样例4:L2VA

仅用一张图片锁定视频结尾瞬间。先构建逻辑通顺的前期画面,再通过动作、物体状态、构图逐步收敛,在末尾镜头定格至图片1。本例为6秒单镜头。

How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 6.00‑second mark of the target video.

integrated_multimodal_description: [Shot 1] Live‑action, cinematic, a close shot begins with an intact drinking glass near the edge of a dark wooden table, while the same hand and sleeve visible in <Picture 1> approach from the right. The camera pushes in with small amplitude at slow speed as the fingertips strike the rim. The glass tips, falls, and hits the floor with a sharp impact; cracks spread through it as fragments slide outward. Toward the end, the moving pieces lose momentum and settle into the exact broken arrangement, hand position, camera angle, lighting, and final composition established by <Picture 1>.

overall_soundscape: Fingertips tap the glass before it scrapes across the tabletop, falls, and breaks with a sharp crash. Small fragments scatter and gradually stop sliding across the floor.

non_diegetic_music: A low electronic pulse at a slow tempo, ending immediately after the glass breaks.

参考图片与目标视频的对齐规则——<图片1>(来自镜头1)对齐目标视频6.00秒时间点。

综合多模态描述:[镜头1]实拍真人,电影质感,特写镜头起始画面:深色木桌桌边放着完好玻璃杯,<图片1>中出现的那只手连同衣袖从画面右侧伸入。镜头小幅慢速向前推近,指尖磕碰到杯沿;玻璃杯倾斜、坠落,重重砸向地面,裂纹蔓延开,碎片向外四散滑动。临近结尾,碎片运动停止,碎裂状态、手部位置、拍摄角度、光影、整体构图完全定格为<图片1>。

整体环境音:指尖轻叩玻璃杯,杯子在桌面摩擦、坠落,随后传来清脆破碎巨响。细小碎片飞溅,在地面滑动后慢慢停下。

非画内背景音乐:节奏缓慢的低频电子脉冲音效,玻璃杯破碎瞬间立刻终止。

注:代码块内部提示词原始英文为模型输入文本,保留原文;代码块外为中文翻译说明。