AI 视频资产怎么做,画面才稳定?
从人物身份、妆造、表情、动作、场景、道具和色板入手,建立可复用的 AI 视频制作资产系统。
角色换一个角度就变脸,衣服在运动中改版,道具忽大忽小,反打镜头像换了一个房间——这些问题通常不是画质词不够,而是缺少能约束模型的视觉证据。
提示词负责下达任务,资产负责提供事实与边界。先把证据做完整,再安排镜头。
| 资产 | 必须固定什么 | 最小交付内容 | 不要继承什么 |
|---|---|---|---|
| 人物身份 | 五官、发型、肤色、体型 | 正脸、侧脸、服装正面、完整背面 | 动作、场景、镜头 |
| 妆造 | 版型、材质、配色、鞋子 | 每套造型一张独立全身资产 | 表情和剧情 |
| 表情 | 情绪状态与强度 | 同机位、同光线的表情矩阵 | 身体动作 |
| 动作 | 路径、重心、接触和顺序 | 连续关键姿态 | 身份和场景风格 |
| 场景 | 空间、光线、材质和家具 | 多角度空镜加平面图 | 无关人物 |
| 道具 | 形状、比例、结构和材质 | 正面、侧面、背面和细节 | 手势和剧情 |
| 色板 | 颜色角色与使用比例 | 角色、HEX 色值和用途 | 物体形状 |
1. 为什么最终提示词之前要先做资产
一张图只能证明一个角度、一次光线和一个瞬间。完整资产包要告诉模型:镜头、服装、情绪或动作改变以后,哪些东西仍然不能变。
把每份参考素材当作制作证据,并且只给它一项明确职责。




2. 人物资产:让身份可以被检查
先做一张干净的基础人物图,再扩展成设定板。检查侧脸鼻梁、下颌、耳饰、后脑发型、衣服缝线、鞋子和身体比例,不要只看脸好不好看。


模板 A:一张头像加三张对齐的全身视图
Create a photorealistic four-view character reference sheet from Image 1. Use a clean white studio background. On the left, show one highly detailed, perfectly front-facing live-action head-and-shoulders portrait. Preserve the exact facial structure, expression, hairstyle, hair accessories, pores, fine lines, eyelashes, flyaway hairs, and natural skin texture. No beauty retouching and no over-sharpening. On the right, show the exact same character in a standard A-pose as three aligned full-body views: front, side, and back. Preserve the same body proportions, complete wardrobe design, fabric layers, shoes, accessories, and realistic material behavior in every view. No scene, no action, no extra people, no logo, no watermark.模板 B:双头像加放大的服装视图
Create a professional photorealistic character reference board from Image 1 on a neutral grey studio background. The layout must have two sections. The upper third contains two large face close-ups: a perfectly front-facing portrait on the left and a clean side-profile portrait on the right. Preserve the exact facial identity, bone structure, hair, makeup, age, skin texture, and accessories. The lower two-thirds contains two wardrobe views. On the left, show a magnified front-facing A-pose wardrobe view cropped from just below the neck to the shoes so the garment construction fills the frame. On the right, show a complete rear full-body view. Keep lighting, lens, body proportions, clothing, and materials identical across all panels. This is a production reference sheet, not a poster or fashion editorial. No action, no complex background, no logo, no watermark.最重视身份和全身一致性时用宽松四视图;最重视脸部细节和服装结构时用高信息密度模板。
3. 妆造资产:建立同一个角色的造型库
不要让一张人物图同时负责身份和所有未来服装。先做中性全身底图,再一次只替换一套完整造型,并保持人物、姿势、镜头和光线不变。


Create a full-body studio image of the person in Image 1. Preserve the exact face, hairstyle, glasses, age, skin tone, body proportions, and wardrobe. Use a neutral grey background, standard relaxed A-pose, even studio lighting, and realistic anatomy. Keep the full body and shoes visible. No logo, watermark, or extra text.Use Image 1 for the character's identity and body proportions. Use Image 2 only for the complete wardrobe, shoes, and accessories. Replace the outfit in Image 1 with the full outfit from Image 2 while preserving the same face, hairstyle, glasses, pose, body shape, lighting, camera, and neutral studio background. Do not change the person. No logo, watermark, or extra text.每套确认后的服装都要生成清晰的正面、侧面和背面资产,并命名为 LOOK_DAILY、LOOK_SPORT、LOOK_FORMAL,后续直接调用。

4. 表情和动作资产
表情资产让情绪变化时不换脸;动作资产让模型不会跳过起点与终点之间的过程。
表情提示词

Using the exact same person from Image 1, create a photorealistic facial-expression asset sheet on a neutral grey studio background. Use a clean 4×4 grid with sixteen consistent head-and-shoulder portraits. Preserve the same facial structure, hairstyle, glasses, makeup, age, skin texture, lens, angle, and lighting in every cell. Show restrained, realistic variations: neutral, focused, soft smile, curious, surprised, tense, tired, calm, frown, angry, relieved, sad, thinking, happy, shy, and unbothered. Avoid cartoon acting, face drift, beauty retouching, and plastic skin. English labels only. No logo or watermark.动作图鉴提示词

An unarmed female agent is ambushed by a larger attacker in a rain-soaked underground parking garage. She survives by evading, breaking a grip, using a restrained knee strike, sweeping, rolling, using a parked car as cover, disarming the attacker, and exiting. Convert this into a coherent 16-panel cinematic action-beat library. Each panel must show one readable action instant with a clear number, short English beat name, body direction, balance, contact point, and screen direction. Keep both character identities, wardrobe, garage geography, wet lighting, and camera language consistent. Modern grounded action-film tone, dark, cool, tense, realistic, non-graphic, no injury detail, no logo, no watermark.表情要服务剧情,克制的变化比大量夸张表情更实用。动作图的每一格只表达一个可读的瞬间、接触点、重心状态和运动方向。
5. 场景资产:建立场景圣经
一张好看的房间图无法证明反打镜头里应该出现什么。场景圣经要记录入口、反打、工作区、材质细节、常驻道具、光线方向和各区域的平面关系。
先生成空镜。人物和剧情动作属于最终镜头,不属于场景证据。

Create a complete production scene bible for an empty modern community table-tennis hall used in a gentle slice-of-life comedy. The location must feel ordinary, authentic, slightly awkward, and lived-in rather than professional or spectacular. Use a clean 3×3 grid. Show: entrance wide shot; reverse wide shot; left-side view; right-side view; player eye-level view; low view along the net; bench and water station; close-up of table, paddle, balls and worn floor markings; and a readable top-down floor plan. Preserve the exact room geometry, table positions, windows, doors, lockers, benches, water dispenser, dark-green lower walls, warm-white upper walls, fluorescent lighting and scattered balls across every panel. No people, logos, subtitles, decorative poster design, or invented rooms. Use small English view labels only.6. 道具资产:固定几何,而不只是类别
“黑色运动包”或“红色路锥”只是类别。道具资产要证明长宽比、接缝、拉链、五金、材质和功能细节,让它转动或被遮挡后仍然保持同一结构。

Create a physical prop turnaround sheet on a white or light-grey background. Show the selected prop in front, side, and rear views, plus one close-up of its most important functional detail. Preserve the exact shape, color, material, scale, seams, hardware, and structural features across every view. Use clean English labels only. No hands, people, scene lighting, brand logo, or watermark.7. 色板资产:保存颜色职责,而不只是色块
每个颜色都有职责与允许的使用比例,色板才有价值。要写清哪些颜色可以主导环境,哪些支持服装和材质,哪些只能作为小面积点缀。


You are a film colorist, visual art director, and AI-video palette designer. Analyze all uploaded frame references as one visual system rather than as separate images. Extract exactly seven shared production colors: BASE, SUPPORT, SHADOW, HIGHLIGHT, SKIN, REFLECTION, and ACCENT. For each color, provide one #RRGGBB HEX value and one concise use case. Then create one clean professional palette asset image containing only the seven swatches, English role label, HEX value, and use case. Explain which colors may cover large areas, which must remain small accents, and which hues should be avoided. No people, scene illustration, long theory, logo, watermark, subtitles, neon saturation, or random colors.上传多份素材时,先标注每个输入的职责,再写时间轴。不要让模型猜某张图负责身份、动作、空间、光线还是颜色。
8. 把资产包交给视频模型
先写素材分工,再写剧情。每一行都要回答:素材属于谁、具体参考什么、哪些内容不能继承。
| 输入 | 对应主体 | 唯一职责 | 明确排除 |
|---|---|---|---|
图片 1 | CHARACTER_A | 人脸、发型、肤色、体型 | 不参考摄影棚背景 |
图片 2 | LOOK_SPORT | 服装、鞋子和配饰 | 不重新定义人脸 |
图片 3 | EXPRESSION_A | 情绪与强度 | 不改变身体动作 |
图片 4 | ACTION_A | 动作顺序与接触关系 | 不参考身份和环境 |
图片 5 | LOCATION_A | 空间、材质、家具与光线 | 忽略偶然物品 |
图片 6 | PROP_A | 形状、比例与结构 | 忽略参考背景 |
图片 7 | PALETTE_A | 颜色职责与比例 | 不要给肤色染色 |
REFERENCE ASSIGNMENTS
Image 1 defines CHARACTER_A's identity only: face, hairstyle, glasses, skin tone, and body proportions.
Image 2 defines LOOK_SPORT only: sports top, skirt, shoes, and accessories. Do not redefine the face.
Image 3 defines EXPRESSION_A only: focused but slightly embarrassed. Keep the performance restrained.
Image 4 defines ACTION_A only: the sequence, balance, contact points, and screen direction.
Image 5 defines LOCATION_A only: room geometry, tables, windows, doors, benches, and light direction.
Image 6 defines PROP_A only: shape, scale, material, seams, and hardware.
Image 7 defines PALETTE_A only: color roles and usage ratios.
Do not inherit unrelated backgrounds, people, text, logos, poses, or composition from any reference.完成绑定后,再写镜头时间轴、摄影机、声音、连续性和排除项。同一素材不要在提示词中途改名。
提交前检查
- 每个常驻角色只有一份身份母版
- 换服装不会改变人脸
- 表情强度有清晰视觉边界
- 动作包含路径、重心、接触和结束状态
- 场景包含反打角度与平面关系
- 重要道具有稳定几何
- 每份资产只有一个编号、一个主体和一项职责
- 最终提示词包含时间轴、连续性、声音与排除项
提示词告诉模型要做什么,资产系统告诉它这个项目究竟是什么。事实稳定以后,复杂镜头才更容易导演,也更容易修改。