What should you establish first to make different scenes look like the same model in the same shoot?

The model's face and outfit often changed from scene to scene, and adding motion could distort even her walk and expressions. Creating several still images and connecting those scenes to look like one shoot were two different challenges.

Rather than work through a platform's features in sequence, I approached this project by first separating out the things that must not change, then checking them at every stage.

Source1Full-body reference photo
Identity sheet6Angles, expressions, and beauty close-ups
Selected frames9Images that passed two rounds of review
Final19.945s2160 × 3644 master

It began with hands-on experience.

My starting point was a single front-facing, full-body photo of a model wearing a white tank top and denim. I first upscaled it in Magnific, then expanded it into a model sheet that let me examine the shape of her eyes, nose, lips, jawline, and hair silhouette from several angles. I was creating a reference standard for the next rounds of generation.

Front-facing close-up of the model with a neutral expression
Front-facing close-up of the model smiling
Close-up of the model with a bright smile
Beauty close-up of the model at a 45-degree angle
Side-profile close-up of the model
Beauty-angle close-up of the model with her chin slightly raised

IDENTITY SHEET — I used six close-ups with the full face in frame to check whether she still looked like the same person across different angles and expressions.

Next, I defined the outfit separately. I broke down the lime sleeveless top, lavender cardigan, terracotta wrap skirt, red flip-flops, and black woven bag into color, material, layering, and accessories. Rather than follow the reference photo's composition or pose, I transferred only the relationships that made the styling work.

Full fashion reference showing a lime top, lavender cardigan, terracotta skirt, and red flip-flops
REFERENCE LOOK
Full-body keyframe of the reference model wearing the reference outfit
GENERATED LOOK

OUTFIT REFERENCE & RESULT — I compared one reference image showing the complete look with one generated image applying the same outfit combination to the reference model.

Once the reference keyframe was stable, I expanded the scenes to include a low-angle shot, waiting at a traffic signal, a high-angle shot, a side view, a selfie, and a bird's-eye view. Changing just one element at a time, whether camera height, distance, or body orientation, made it easier to identify what was causing a result to drift.

I did not use every generated result in the actual edit. First, I rejected still-image candidates in which the face, hands, feet, outfit, or bag had changed. Even when an image looked natural, I removed the shot from the timeline if animating it distorted the walk, expression, hair, or movement of the clothing.

What this process confirmed was that the number of generated results does not equal quality. The final film was not a collection of everything I had made. It consisted only of scenes that passed both the image and motion reviews.

The quality of an AI lookbook comes not from generating more images, but from a process of selecting and regenerating them against the same standards.

— Park Siha, Creative Director

Key perspective

Making the scenes look like one shoot was not just a matter of facial resemblance. The shots connected into a shared time and place when the person, styling, scene, movement, and sound each stayed consistent with their own reference standards.

I did not ask one reference image to do everything, either. The model sheet was my reference for the person, the fashion reference for the outfit, and the first keyframe for the lighting and space. When each reference had a clear role, I could quickly identify what needed to be regenerated whenever a result drifted.

  • 01
    Identity anchor
    Establish consistent eyes, nose, lips, jawline, and hairline across multiple angles.
  • 02
    Style anchor
    Manage the color combination, materials, layering, and accessories as a separate reference.
  • 03
    Scene anchor
    Define the light, space, lens, and camera height for each scene.
  • 04
    Motion rule
    Design the person's subtle movements and the response of the clothing first.
  • 05
    Sound continuity
    Connect the sound so that the scenes feel like the same time and place, even across cuts.

I also judged the success of a still image separately from the success of a video scene. A good keyframe could still produce distortions in the person once animated, and individually strong scenes could lose their rhythm when joined together.

Finally, I rebuilt the music and ambient sound as one continuous span of time rather than dividing them by scene. Even when the image changed, continuous footsteps, city atmosphere, and fabric rustle made the breaks between shots less noticeable.

The key was using the first keyframe as the benchmark for comparison. With every new scene, I checked how far the face, outfit, light, and accessories had drifted from that reference.

Six steps to creating an AI lookbook film

What mattered more than any service's features was which reference standards I established, and in what order. These are the six steps I followed in the actual project.

  1. 01
    Source & Upscale

    Decide what must not change before upscaling.

    Even increasing resolution can change how a person looks.

    In Magnific, I restored skin texture, hair edges, the construction of the clothing, and the denim texture. I kept the settings low, however, so that the facial proportions, body shape, pose, and hair silhouette would remain unchanged.

    KEEPFacial proportions, body shape, pose, hair, and background structure

    AVOIDExcessive skin retouching, enlarged eyes, and changes to the hair silhouette

    Upscale prompt
    High-fidelity portrait and full-body upscale. Preserve the exact facial identity, body proportions, pose, hairstyle silhouette, clothing construction and background geometry. Recover natural skin texture, individual hair strands, rib-knit fabric and denim weave. Clean studio detail, realistic pores, no beauty filter, no facial redesign, no body reshaping.
  2. 02
    Identity Sheet

    Build facial references across multiple angles and expressions.

    One front-facing photo was not enough to define the profile and changes in expression.

    I created a neutral front view, a soft smile, a bright smile, a 45-degree view, a side profile, and a beauty angle, all with the same lighting and outfit. This sheet was not simply a collection of profile photos. It was a reference for checking the features that needed to stay fixed in subsequent scenes.

    Identity prompt
    Create a professional six-image fashion model identity sheet of the exact same young Korean woman from the reference images. Preserve her eye shape and spacing, nose bridge and tip, lip shape, jawline, cheek volume, skin tone, natural asymmetry, long layered black hair and small silver hoop earrings. Minimal clean makeup, realistic skin texture, neutral off-white studio, soft frontal daylight. Produce consistent close-up views: neutral front, soft smile, open smile, three-quarter, side profile and elevated beauty angle. Keep the full head, jawline and shoulders inside every frame. No face redesign, no age change, no glam retouching.
  3. 03
    Styling Transfer

    Transfer the outfit element by element, not the whole photograph.

    Rather than duplicate the reference photo, I isolated the elements that made up its styling.

    I defined the colors, materials, silhouettes, layering, and accessories individually, then applied them to the reference model. Afterward, I checked not only the outfit's colors and materials but also the size and position of the bag and the shape of the shoes.

    Styling prompt
    Dress the same identity-preserved model in a lime green asymmetric camisole, a soft lavender cardigan worn loose around the shoulders, a translucent terracotta wrap midi skirt, minimal red flip-flops, and a small black woven shoulder bag. Preserve the exact face, body proportions, hairstyle and natural skin. Contemporary Seoul street-fashion editorial, sunlit concrete architecture, realistic fabric drape and material texture.
  4. 04
    Scene Expansion & Selection

    Expand the scenes by changing one variable at a time.

    I kept the model and outfit fixed while changing camera height, distance, and body orientation one at a time.

    To avoid repetitive full-body walking shots, I generated separate low-angle, high-angle, side-view, selfie, and bird's-eye-view scenes. I checked that each scene served a different purpose while still looking as though it belonged to the same day and location.

    Scene variation prompt
    Generate a new shot from the same fashion film while preserving the exact model identity and complete outfit. Change only one cinematic variable at a time: camera height, lens distance, body orientation, gesture, or framing. Maintain the same bright midday Seoul concrete location, direct sunlight, natural hard-edged shadows, realistic skin and fabric. Editorial spontaneity, not a repeated catalog pose.

    1ST GATE · IMAGEI checked facial resemblance, hands and feet, and consistency in the outfit, bag, and shadows.

    2ND GATE · MOTIONI rejected shots with sliding feet, distorted expressions, unnatural speed, or warping in the clothing and background.

    I animated only the scenes that passed the still-image review, and removed any shots that broke down again in motion from the edit. Rather than briefly hide an unnatural moment, I replaced the scene with a more stable shot.

    Scene of the model walking toward the camera, viewed on a camera screen
    Key walking scene
    Low-angle walking scene
    Scene of the model waiting in the city
    Motion-blur scene
    High-angle scene
    Side-view scene with a drink
    Selfie scene
    Bird's-eye-view scene

    SELECTED IMAGE BOARD — From the walking shot on the camera screen to the final scene, these nine images were checked for consistency in the person, outfit, light, and space.

    Editing decisionRather than use every generated result, I kept only the scenes that would not strike the viewer as unnatural.

  5. 05
    Motion Design

    Review still images and video results separately.

    Even when a still frame looked natural, excessive motion could easily distort the face, hands, and clothing.

    I began by specifying the shift in weight while walking, brief changes in gaze, the response of the hair and skirt to the breeze, and the bag's gentle sway. I allowed the camera to track subtly, but limited larger movements that could cause the person's face to be reconstructed.

    Motion prompt
    The model takes two natural, confident steps toward camera. Her weight shifts realistically through hips and shoulders; the terracotta wrap skirt and loose lavender cardigan respond to a light summer breeze; the small woven bag sways gently. Subtle blinking and breathing, calm direct gaze. Smooth handheld tracking with minimal parallax. Preserve her exact face, hands, outfit and body proportions. No morphing, no sudden acceleration, no extra limbs, no camera orbit.
  6. 06
    Sound & 4K Finish

    Create a shared sense of time and place before focusing on music.

    Different sound in each scene exposed the edit points, even when the visuals looked natural.

    I removed all the music remaining in the early edit and created a new 20-second instrumental in Higgsfield to match the film's length. Over it, I layered quiet footsteps, distant traffic, city atmosphere, and fabric rustle to connect the scenes.

    Higgsfield music prompt
    20-second seamless instrumental for a sunlit Seoul street-fashion film, contemporary Korean indie electronic with restrained UK garage rhythm, warm bass, crisp two-step drums, airy glass synth plucks and subtle organic percussion; youthful, stylish, editorial, confident; immediate hook, gentle mid-section lift, clean resolved ending; no vocals, no speech, no lyrics, no cinematic trailer effects, no abrupt cuts, no corporate stock-music feeling.
    Final soundtrackHiggsfield music + city ambience + footsteps

    Generating the music in Higgsfield used 1.25 credits. I upscaled the final film from a 1080 × 1822 edit to 2160 × 3644.

How to apply this to your next project

This approach is not tied to any particular tool. After generating a scene or turning it into video, work through the following questions in order to quickly filter out problematic scenes before adding them to the timeline.

  • 01
    Do the relationships between the shape of the eyes, nose, lips, and jawline remain the same in front and side views?
  • 02
    Are any apparent changes in the hairline and hair length due only to the camera angle?
  • 03
    Are the colors and materials of the top, cardigan, skirt, shoes, and bag the same in every scene?
  • 04
    Do the direction of direct sunlight and the edges of the shadows suggest the same time of day on the same day?
  • 05
    Does each scene have a clear role as a low-angle, eye-level, high-angle, or close-up shot?
  • 06
    Does the movement preserve the shape of the face, hands, shoes, and clothing?
  • 07
    Is the music uninterrupted, and do the footsteps and ambience help conceal the cuts?
  • 08
    Have you checked the AI-generation disclosure and music licensing appropriate to the intended use?

Frequently asked questions

What is the most important way to keep the same face across multiple scenes?
Rather than repeatedly referencing a single front-facing photo, it is better to first create an identity sheet with front, 45-degree, and side views, along with close-ups of different expressions. Then, for each new scene, change only one of the camera, pose, or background, not the face.
Can I start with just one face reference photo?
You can, but rather than using that single image directly for video, a more stable approach is to build a sheet of multiple angles and then use the most consistent images as your references.
Did you use every generated image and video in the final film?
No. I first reviewed the still images for identity, hands and feet, clothing, and accessories. After animating them, I made another pass to reject shots in which the walk, expression, clothing physics, or background broke down. Rather than hide an unnatural moment, I deleted the entire scene and replaced it with another shot.
Which should I create first: the scene image or the motion?
The scene image comes first. Motion extends the composition and spatial relationships established in the still frame, so the face, outfit, hands, and shoes need to be stable in the keyframe before you design the movement.
Does the music need to be generated to match the exact length of the video?
Yes, if you do not plan to edit it again. Including the video's exact length, a hook in the first 1–2 seconds, a lift in the middle, and how the ending should resolve in your prompt makes any trimming less noticeable.
What should I be careful about when publishing an AI lookbook on a brand channel?
You need to check the rights relating to the person, clothing, background, and music, as well as each generation tool's commercial-use terms. If it is a styling experiment rather than an actual product, it is best to clearly label it as AI-generated and an unofficial concept.