Prompting Guide
This workflow provides two image references: `<Picture 1>` and `<Picture 2>`.
Use the images as sources for reusable visual content in the target video. A person, object, outfit, environment, style, pose, or other visible element taken from a reference image should normally be defined as a `<Subject N>`.
For a two-person setup:
<Subject 1> comes from <Picture 1>
<Subject 2> comes from <Picture 2>
Keep the same subject number everywhere in the prompt.
Full Reference Prompt Structure
Use these six sections in this order:
subject_definitions:
...
summary:
...
retention_analysis:
...
detailed_description:
...
overall_soundscape:
...
non_diegetic_music:
...
Subject Definitions
`subject_definitions` explains what reusable content comes from each image.
Example:
subject_definitions:
<Subject 1> is the woman in <Picture 1>, with her recognizable facial identity, long dark hair, skin tone, and body proportions.
<Subject 2> is the woman in <Picture 2>, with her recognizable facial identity, blonde hair, skin tone, and body proportions.
Define only the traits that matter.
The two references can have different jobs:
<Subject 1> is the woman in <Picture 1>, preserving her identity and hairstyle.
<Subject 2> is the red evening dress in <Picture 2>, preserving its cut, fabric, and color.
Or:
<Subject 1> is the person in <Picture 1>, preserving identity and body proportions.
<Subject 2> is the hotel interior in <Picture 2>, preserving its marble floor, warm lighting, and modern architecture.
Use `<Subject N>` for visible content that will actually be reused in the generated video.
A standalone `<Picture N>` is mainly for cases where the image itself is intended to act as a concrete frame, keyframe, composition anchor, or storyboard reference. For ordinary identity, clothing, environment, or style reference, cite the picture inside the corresponding subject definition.
Summary
`summary` is one short paragraph describing the target video and the main reference relationship.
For this workflow, the normal task type is:
[reference generation]
Example:
summary:
[reference generation] The target video places <Subject 1> and <Subject 2> together as two different women in a newly generated luxury hotel lobby, where they walk side by side and interact naturally.
Use only labels already defined in `subject_definitions`.
Retention Analysis
`retention_analysis` specifies how each referenced subject should carry into the result.
Relationship markers:
`fully_preserved` — keep the defined referenced characteristics.
`partially_preserved` — keep the subject while intentionally changing some characteristics.
`attribute_transfer` — transfer referenced characteristics to another identifiable target.
`weak_reference` — retain only broad similarity such as style, category, composition, or atmosphere.
Example:
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - preserve her facial identity, hairstyle, skin tone, and body proportions.
<Subject 2> (appears in [Shot 1]): fully_preserved - preserve her facial identity, hairstyle, skin tone, and body proportions.
If something should change, do not list it as fully preserved.
Example:
<Subject 1> (appears in [Shot 1]): partially_preserved - preserve her facial identity, hairstyle, skin tone, and body proportions while changing her clothing to a black evening dress.
Detailed Description
`detailed_description` is the main video description.
Begin with one or two sentences establishing visual style, lighting, and atmosphere. Then describe the target video shot by shot in playback order.
At the first clear appearance of an important subject, establish its position, relevant referenced traits, and current action.
Example:
detailed_description:
The target video uses cinematic live-action photography with warm interior lighting and natural human motion.
[Shot 1] A medium-wide two-shot shows <Subject 1> and <Subject 2> standing side by side in a modern luxury hotel lobby. Both women are visible together. <Subject 1> turns toward <Subject 2> and smiles. <Subject 2> looks back and laughs softly. They begin walking toward the camera together. The camera tracks backward slowly, keeping both women in frame. Their hair and clothing move naturally as they walk. They end side by side near the center of the lobby with both faces clearly visible.
Describe visible actions and state changes instead of abstract intent.
Weak:
<Subject 1> acts confident.
Better:
<Subject 1> straightens her shoulders, lifts her chin slightly, and steps forward while maintaining eye contact with <Subject 2>.
Assigning the Two References
Give each reference a clear purpose.
Useful roles include:
identity
facial appearance
hairstyle
body proportions
clothing
props
environment
architecture
color palette
visual style
pose
expression
If the images depict two different people, keep them as separate subjects.
If both must appear together, establish that in the opening composition:
[Shot 1] A medium-wide two-shot shows <Subject 1> on the left and <Subject 2> on the right, standing together in the same lobby.
State which traits belong to which subject instead of using vague instructions such as "use both references."
Shots and Timing
`[Shot 1]` is the opening shot and has no timestamp.
Later shots use increasing cut times:
[Shot 2] At 00:04.500, the shot cuts to a close-up of <Subject 2>.
Use additional shots when they reveal a genuinely new angle, reaction, detail, location, or state.
Camera Movement
Write camera movement naturally inside the current shot.
`Push In` / `Pull Out` — camera moves forward/backward
`Zoom In` / `Zoom Out` — lens zoom
`Pan Left` / `Pan Right` — camera rotates horizontally
`Truck Left` / `Truck Right` — camera moves sideways
`Tilt Up` / `Tilt Down` — camera rotates vertically
`Pedestal Up` / `Pedestal Down` — camera moves vertically
`Arc Shot` — camera moves around the subject
`Tracking Shot` — camera follows the subject
`Static Shot` — no camera movement
`POV` — point-of-view shot
`Roll Clockwise` / `Roll Counterclockwise`
`Shake Slightly` / `Shake Strongly`
Optional modifiers:
with small amplitude
with large amplitude
at slow speed
at fast speed
Example:
The camera tracks backward slowly, keeping <Subject 1> and <Subject 2> together in the center of the frame.
Dialogue
Use stable speaker IDs such as `(S1)` and `(S2)`.
The subject label identifies the referenced visual subject. The speaker ID identifies the voice.
Example:
<Subject 1> (S1) turns toward <Subject 2> and says softly, <d>[English] I didn't expect to see you here.</d>
<Subject 2> (S2) smiles and replies, <d>[English] Neither did I.</d>
Inside `<d>...</d>`, include only:
[Language] exact spoken dialogue
Put speaker identity, emotion, volume, and delivery outside the dialogue tag.
Reuse the same speaker ID every time that person speaks.
Visible Text
Put exact on-screen text in double quotes:
The illuminated sign behind them reads "GRAND HOTEL".
Sound
`overall_soundscape` summarizes ambience and physical sounds across the full video.
Example:
overall_soundscape: Quiet hotel room tone, soft footsteps on marble, distant conversation, subtle clothing movement, and natural laughter.
Dialogue belongs in `detailed_description` and should not be repeated here.
`non_diegetic_music` describes background music audible to the audience but not to the characters.
Example:
non_diegetic_music: Soft piano with sustained low strings at a slow tempo.
If no background score is wanted:
non_diegetic_music: N/A
Complete Two-Image Example
subject_definitions:
<Subject 1> is the woman in <Picture 1>, preserving her recognizable facial identity, dark hair, skin tone, and body proportions.
<Subject 2> is the woman in <Picture 2>, preserving her recognizable facial identity, blonde hair, skin tone, and body proportions.
summary:
[reference generation] <Subject 1> and <Subject 2> appear together as two different women in a newly generated luxury hotel lobby.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - preserve her facial identity, hairstyle, skin tone, and body proportions.
<Subject 2> (appears in [Shot 1]): fully_preserved - preserve her facial identity, hairstyle, skin tone, and body proportions.
detailed_description:
The target video uses cinematic live-action photography with warm hotel lighting and natural human motion.
[Shot 1] A medium-wide two-shot shows <Subject 1> and <Subject 2> standing side by side in a modern luxury hotel lobby. Both women are visible together. <Subject 1> turns toward <Subject 2> and smiles. <Subject 2> looks back and laughs softly. They step forward and begin walking toward the camera together. The camera tracks backward slowly, keeping both women in frame. Their hair and clothing move naturally while warm light reflects across the marble floor. They finish side by side with both faces clearly visible.
overall_soundscape: Quiet hotel ambience, soft footsteps on marble, distant conversation, and natural laughter.
non_diegetic_music: N/A
Best practice: define exactly what each image supplies, keep reference labels stable, state which traits should be retained, and describe the target video in explicit playback order with clear composition, physical action, camera movement, sound, and final visible state.



