Cómo hacer que las personas generadas por IA parezcan reales
Un flujo práctico para mantener la identidad, elegir escenarios creíbles y dirigir utilería, cadenas de acciones y tiempos.
Un rostro nítido y una piel detallada no hacen que un video de IA sea automáticamente creíble. El realismo depende de la relación entre identidad, entorno, motivación y tiempo.
1. El realismo es una relación, no una resolución
Palabras como 8K, ultradetallado o textura de piel cinematográfica mejoran la superficie, pero no dirigen una actuación. También importan la continuidad del rostro, la dirección de la luz, la mirada, la causa de cada gesto y el estado de los objetos.
2. Bloquea la identidad antes de dirigir la actuación
Para un personaje recurrente, prepara una ficha compacta: primer plano frontal y de perfil, ropa de frente y vista completa posterior. Peinado, accesorios, vestuario y proporciones deben coincidir.

Numera cada recurso y dale una sola responsabilidad:
Image 1 = Character A identity and wardrobe only.
Image 2 = Lakeside café location and window-light direction only.
Image 3 = Coffee cup prop only.
Image 4 = Phone prop only.
Image 5 = Eye-level composition only.
Preserve Character A's face, hairstyle, outfit and body proportions.
Do not inherit unrelated people, objects or text from any reference.
Este prompt convierte la ficha de identidad en material de producción, no en un póster:
Create a photorealistic four-view character reference board on a clean gray studio background.
Use a precise, orderly layout with no story action, decorative typography, or complex environment.
Top section: two detailed face close-ups, front view on the left and profile view on the right.
Show the same facial structure, skin texture, hairstyle, makeup, earrings, and expression in both views.
Keep pores, fine lines, eyelashes, flyaway hair, and natural light response. No beauty retouching or plastic skin.
Bottom section: two wardrobe references.
Left: a close front A-pose wardrobe view from below the neck to the shoes. Keep both arms, hands, torso,
legs, and shoes visible; crop the head completely so the clothing occupies more of the frame.
Right: a normal full rear A-pose from head to shoes, showing the back of the hairstyle and garment construction.
The four panels must show the same fictional adult, the same outfit, accessories, body proportions,
materials, and styling. Use clean three-point studio lighting and a seamless gray background.
No poster layout, no text, no mismatched clothing, no warped proportions, no glossy CG skin.3. Elige escenarios con una lógica de cámara creíble
Una fotografía arquitectónica perfecta no siempre es la mejor referencia para una escena cotidiana. Una ventana un poco sobreexpuesta, cables, libros abiertos o telas arrugadas aportan señales de vida real.


Comprueba la fuente de luz, la altura de cámara y la perspectiva. La iluminación del sujeto debe coincidir con la ventana, y la altura de ojos o pecho suele funcionar mejor para video casual.


Evidencia de uso
Un espacio creíble muestra que alguien lo utiliza: una taza a medio llenar, un cable de carga, un libro abierto o una tela ligeramente arrugada.

4. Usa utilería y cadenas de acciones causales
Los objetos cotidianos y fáciles de manipular dan a las manos y a la mirada un objetivo físico.

Weak: The woman sits naturally at the table.
Stronger: She looks down at the cup, wraps her right hand around the handle,
lifts it slowly, takes one small sip, and places it back in the same position.“Sentarse con naturalidad” no ofrece un objetivo físico a las manos ni a la mirada. Una taza, un libro o un teléfono crea un motivo y cambios de estado visibles.


She looks across the lake. The phone vibrates, so her eyes move to the screen.
She pauses without picking it up, reaches for the cup instead, takes one small sip,
sets the cup down, and finally returns her attention to the phone.5. Construye el prompt en cuatro capas
- Estilo de captura — dispositivo, encuadre, luz y movimiento de cámara.
- Asignación de referencias — la función exacta de cada entrada.
- Línea de tiempo — acciones en intervalos continuos.
- Continuidad y límites — qué debe permanecer fijo y qué no debe aparecer.

STYLE & CAMERA
Naturalistic lifestyle documentary look, authentic handheld smartphone footage, vertical 9:16,
bright late-afternoon window light, subtle handheld breathing, and a slow gentle push-in.
REFERENCE ASSIGNMENTS
Image 1 defines character identity only. Preserve the same facial features, hairstyle, pearl earrings,
white sleeveless top, dark skirt, and body proportions.
Image 2 defines the lakeside café location only. Preserve the warm wood interior, window seat, lake view,
and direction of the natural light. Ignore any annotations or visible text in the reference.
Image 3 defines the silver smartphone and Image 4 defines the ceramic coffee cup.
Image 5 defines the friend's eye-level composition across the round table.
Images 6–9 define room details, camera height, lighting continuity, and action continuity only.
Do not copy unrelated people, objects, or text from any reference.
SHOT TIMELINE
0–3 seconds: She sits relaxed beside the window and looks across the lake with a faint smile.
She gradually notices the friend filming, turns naturally toward the camera, makes brief eye contact,
and gives a warm, unforced smile.
3–6 seconds: The silver phone on the table vibrates softly. She looks down at the screen and pauses
without picking it up, then glances back toward the camera as if acknowledging the friend behind it.
6–10 seconds: She reaches for the coffee-cup handle, lifts the cup, takes one small sip, and places it
gently back on the table. While setting it down, she briefly looks at the camera before returning her gaze
to the phone.
CONTINUITY, SOUND & CONSTRAINTS
Throughout the shot, preserve natural blinking, breathing, tiny head tilts, subtle posture adjustments,
and believable eye-line changes. Keep the same person, wardrobe, café, phone, cup, shot direction,
framing, light, and overall duration. Use only quiet café room tone and a light lakeside breeze.
No extra people, subtitles, visible text, logos, music, or watermark.
6. Añade el lenguaje de calidad al final
Corrige primero la capa adecuada: identidad con referencias, integración con luz y perspectiva, actuación con causalidad y saltos con una línea de tiempo continua.
Añade el lenguaje de calidad solo cuando identidad, escena, acción y ritmo ya sean estables:
Natural skin texture, stable facial identity, physically believable hand contact,
consistent window-light direction, coherent object state, subtle handheld motion,
clean motion blur, no beauty-filter skin, no sudden camera jump, no extra fingers,
no disappearing props, no subtitles, no logos, no watermark.| Síntoma | Corrige primero |
|---|---|
| El rostro cambia | Referencias y etiquetas de identidad |
| El sujeto parece pegado al fondo | Luz, altura y perspectiva |
| La actuación parece mecánica | Causa, reacción y micropausas |
| La utilería salta o desaparece | Función y cambio de estado |
| El plano va demasiado rápido | Menos acciones e intervalos más largos |
Lista de comprobación antes de generar
- ¿Se puede vincular cada persona con una única referencia de identidad?
- ¿Tiene cada referencia una sola responsabilidad explícita?
- ¿Son físicamente plausibles la luz principal y la altura de cámara?
- ¿Tiene cada acción importante un detonante y un resultado visible?
- ¿Cubre la línea de tiempo toda la duración sin huecos?
- ¿Están protegidos identidad, vestuario, utilería, lugar, encuadre y sonido?
La pregunta final no es solo “¿Se ve cinematográfico?”, sino: ¿Quién es esta persona? ¿Dónde está? ¿Por qué se mueve? ¿Qué debe suceder después?
Cómo mantener el rostro consistente al editar imágenes con IA
Por qué las ediciones con IA cambian sutilmente la cara de tu modelo y siete técnicas prácticas para conservar la misma identidad en cada generación.
Crea un sistema de recursos coherente para vídeo con IA
Prepara referencias reutilizables de identidad, vestuario, expresión, acción, localización, utilería y color.