So wirken KI-Personen in Videos realistisch
Ein praktischer Workflow für konsistente Figuren, glaubwürdige Schauplätze, Requisiten, Handlungsketten und zeitbasierte Video-Prompts.
Scharfe Gesichter und detaillierte Haut machen ein KI-Video nicht automatisch glaubwürdig. Entscheidend ist die Beziehung zwischen Identität, Umgebung, Motivation und Zeit.
1. Realismus ist eine Beziehung, keine Auflösung
Begriffe wie 8K, ultradetailliert und filmische Hauttextur verbessern die Oberfläche, führen aber keine Darstellung. Zuschauer achten außerdem darauf, ob die Figur dieselbe bleibt, das Licht zum Raum passt, Handlungen einen Grund haben und Requisiten ihren Zustand korrekt verändern.
2. Zuerst die Identität sichern
Bereite für wiederkehrende Figuren ein kompaktes Referenzblatt vor: frontale und seitliche Gesichtsaufnahme, klare Vorderansicht der Kleidung und vollständige Rückansicht. Frisur, Accessoires, Outfit und Körperproportionen müssen in allen Ansichten gleich bleiben.

Nummeriere anschließend jedes Material und gib ihm genau eine Aufgabe:
Image 1 = Character A identity and wardrobe only.
Image 2 = Lakeside café location and window-light direction only.
Image 3 = Coffee cup prop only.
Image 4 = Phone prop only.
Image 5 = Eye-level composition only.
Preserve Character A's face, hairstyle, outfit and body proportions.
Do not inherit unrelated people, objects or text from any reference.
Mit diesem Prompt wird das Referenzboard zum Produktionsmaterial statt zum Poster:
Create a photorealistic four-view character reference board on a clean gray studio background.
Use a precise, orderly layout with no story action, decorative typography, or complex environment.
Top section: two detailed face close-ups, front view on the left and profile view on the right.
Show the same facial structure, skin texture, hairstyle, makeup, earrings, and expression in both views.
Keep pores, fine lines, eyelashes, flyaway hair, and natural light response. No beauty retouching or plastic skin.
Bottom section: two wardrobe references.
Left: a close front A-pose wardrobe view from below the neck to the shoes. Keep both arms, hands, torso,
legs, and shoes visible; crop the head completely so the clothing occupies more of the frame.
Right: a normal full rear A-pose from head to shoes, showing the back of the hairstyle and garment construction.
The four panels must show the same fictional adult, the same outfit, accessories, body proportions,
materials, and styling. Use clean three-point studio lighting and a seamless gray background.
No poster layout, no text, no mismatched clothing, no warped proportions, no glossy CG skin.3. Schauplätze mit glaubwürdiger Aufnahmelogik wählen
Eine perfekte Architekturfotografie ist nicht immer die beste Vorlage für ein lockeres Lifestyle-Video. Leichte Überbelichtung am Fenster, Falten, Kabel, Bücher oder ein halbvolles Glas liefern oft genau die Alltagsspuren, die einen Raum glaubwürdig machen.


Prüfe Lichtquelle, Kamerahöhe und Perspektive. Fensterlicht und Motivlicht müssen aus derselben Richtung kommen; für Alltagsszenen funktionieren Augen- oder Brusthöhe meist besser als extreme Winkel.


Sichtbare Nutzungsspuren
Ein glaubwürdiger Raum zeigt, dass er benutzt wird: eine halbvolle Tasse, ein Ladekabel, ein offenes Buch oder leicht zerknitterter Stoff.

4. Requisiten und kausale Handlungsketten verwenden
Gewöhnliche, leicht bedienbare Requisiten geben Händen und Blick ein klares Ziel.

Weak: The woman sits naturally at the table.
Stronger: She looks down at the cup, wraps her right hand around the handle,
lifts it slowly, takes one small sip, and places it back in the same position.Abstrakte Anweisungen wie „natürlich sitzen“ geben Händen und Blick keinen physischen Zweck. Eine Tasse, ein Buch oder ein Telefon erzeugt dagegen ein Ziel und sichtbare Zustandsänderungen.


She looks across the lake. The phone vibrates, so her eyes move to the screen.
She pauses without picking it up, reaches for the cup instead, takes one small sip,
sets the cup down, and finally returns her attention to the phone.5. Den Prompt in vier Ebenen aufbauen
- Aufnahmestil – Gerät, Bildausschnitt, Licht und Kameraverhalten.
- Referenzzuweisung – die eindeutige Aufgabe jedes Inputs.
- Zeitachse – Aktionen in durchgehenden Zeitbereichen.
- Kontinuität und Grenzen – was stabil bleiben und was vermieden werden muss.

STYLE & CAMERA
Naturalistic lifestyle documentary look, authentic handheld smartphone footage, vertical 9:16,
bright late-afternoon window light, subtle handheld breathing, and a slow gentle push-in.
REFERENCE ASSIGNMENTS
Image 1 defines character identity only. Preserve the same facial features, hairstyle, pearl earrings,
white sleeveless top, dark skirt, and body proportions.
Image 2 defines the lakeside café location only. Preserve the warm wood interior, window seat, lake view,
and direction of the natural light. Ignore any annotations or visible text in the reference.
Image 3 defines the silver smartphone and Image 4 defines the ceramic coffee cup.
Image 5 defines the friend's eye-level composition across the round table.
Images 6–9 define room details, camera height, lighting continuity, and action continuity only.
Do not copy unrelated people, objects, or text from any reference.
SHOT TIMELINE
0–3 seconds: She sits relaxed beside the window and looks across the lake with a faint smile.
She gradually notices the friend filming, turns naturally toward the camera, makes brief eye contact,
and gives a warm, unforced smile.
3–6 seconds: The silver phone on the table vibrates softly. She looks down at the screen and pauses
without picking it up, then glances back toward the camera as if acknowledging the friend behind it.
6–10 seconds: She reaches for the coffee-cup handle, lifts the cup, takes one small sip, and places it
gently back on the table. While setting it down, she briefly looks at the camera before returning her gaze
to the phone.
CONTINUITY, SOUND & CONSTRAINTS
Throughout the shot, preserve natural blinking, breathing, tiny head tilts, subtle posture adjustments,
and believable eye-line changes. Keep the same person, wardrobe, café, phone, cup, shot direction,
framing, light, and overall duration. Use only quiet café room tone and a light lakeside breeze.
No extra people, subtitles, visible text, logos, music, or watermark.
6. Qualitätsbegriffe zuletzt ergänzen
Behebe zuerst die richtige Ebene: Identitätsdrift über Referenzen, falsche Raumwirkung über Licht und Perspektive, mechanische Bewegungen über Handlungsketten und Sprünge über eine kontinuierliche Zeitachse.
Ergänze Qualitätsangaben erst, nachdem Identität, Szene, Handlung und Timing stabil sind:
Natural skin texture, stable facial identity, physically believable hand contact,
consistent window-light direction, coherent object state, subtle handheld motion,
clean motion blur, no beauty-filter skin, no sudden camera jump, no extra fingers,
no disappearing props, no subtitles, no logos, no watermark.| Problem | Zuerst ändern |
|---|---|
| Das Gesicht verändert sich | Identitätsreferenzen und Labels |
| Die Figur wirkt in den Raum geklebt | Licht, Kamerahöhe und Perspektive |
| Die Bewegung wirkt mechanisch | Ursache, Reaktion und Mikropausen |
| Requisiten springen oder verschwinden | Zuständigkeit und Zustandsänderung |
| Der Shot wirkt gehetzt | Weniger Aktionen, längere Zeitbereiche |
Checkliste vor der Generierung
- Lässt sich jede Person eindeutig auf eine Identitätsreferenz zurückführen?
- Hat jede Referenz genau eine klar benannte Aufgabe?
- Sind Hauptlicht und Kamerahöhe physisch plausibel?
- Hat jede wichtige Handlung einen Auslöser und ein sichtbares Ergebnis?
- Deckt die Zeitleiste die gesamte Dauer ohne Lücken ab?
- Sind Identität, Kleidung, Requisiten, Ort, Bildausschnitt und Ton geschützt?
Frage am Ende nicht nur „Ist es filmisch?“, sondern: Wer ist diese Person? Wo ist sie? Warum bewegt sie sich? Was geschieht als Nächstes?
So bleibt das Gesicht bei KI-Bildbearbeitung konsistent
Warum KI-Bearbeitungen das Gesicht Ihres Models unbemerkt verändern und sieben praktische Techniken, mit denen die Identität über alle Generierungen hinweg erhalten bleibt.
Ein konsistentes KI-Video-Assetsystem aufbauen
Wiederverwendbare Assets für Identität, Kleidung, Mimik, Bewegung, Schauplatz, Requisiten und Farbe erstellen.