Talking Head ScreenScreen RecordingStatic LockedMeasured Pace

Luma AI: Shot-by-Shot Breakdown

All 18 cuts of Luma AI's 40-second Facebook ad on the editing timeline — framing, angle, lighting, mood and the persuasion job of every shot, plus the classification intel behind the creative.

Updated 2026-03-2818 min read

18 cuts in 40 seconds, and 13 of them move the camera. This breakdown puts every shot of the Luma AI ad on the timeline — what is in the frame, how it is framed, and the persuasion job each cut performs while it owns the screen. Follow the playhead: all 7 beats and 18 shots below are pinned to their exact timeframe.

SRC00:00.0
Luma AI · Facebook · 0:40

Master Timeline · 7 beats · 18 cuts

00:00.0
00:00.000:39.8

Now on screen

Beat 01 · OPENING
01

Talking Head Classic

IN 00:00.0 · OUT 00:05.0 · DUR 5.0s

Are you still making images, video, audio with different AI tools? You gotta try Luma Agents.Full analysis

Cut list · click a shot to jump the playhead

BEAT 01

Direct Question Hook

OPENING00:00.000:04.5 · 1 shots

It opens with a direct question: “Are you still making images, video, audio with different AI tools?” This immediately frames the viewer as the problem-holder and sets up an answer the video will provide, pulling them into self-check mode right away.

Why this works:This leverages Direct Question Hook—because the question is clear and answerable, viewers mentally respond instantly (“yes/no”), which increases attention and continuation. It also uses Self-Relevance/Personalization: by asking about “you” and your current workflow, it creates immediate personal stakes, making it harder to scroll past.

Visual Direction

The opening is a static graphic featuring a casual man speaking directly, overlaid with bold, highlighted text. The neutral palette keeps focus on the text and speaker's expression. Typography clearly frames the problem and introduces the solution, 'Luma Agents,' with emphasis.

01

Talking Head Classic

IN 00:00.0 · OUT 00:05.0 · DUR 5.0s

Are you still making images, video, audio with different AI tools? You gotta try Luma Agents.
00:00.000:00.000:05.000:39.8

A man in a room speaking directly to the camera, gesturing with his hands. He wears a casual patterned shirt and an orange baseball cap.

Why it works: This opening shot directly addresses the viewer with a question, immediately engaging them and setting up the problem that the product will solve. The casual look and direct address build relatability and a personal connection.

Angle

Eye Level Straight

Camera

Static Locked

Lighting

Soft Window Light

Color

Neutral Palette

Mood

Excited And Fast-paced

BEAT 02

Quick Fix Instruction

DELIVERY00:04.500:10.2 · 4 shots

The speaker issues a direct, urgent recommendation: “You gotta try Luma Agents. You gotta check this out.” This functions as a quick instruction to take immediate action rather than explaining features or benefits yet.

Why this works:This leverages Commitment Bias and Action Bias: the repeated “You gotta” frames trying it as the default next step, making inaction feel like a missed opportunity. It also uses Social Proof/Popularity Cues (“check this out”) to reduce hesitation by implying others are already doing it.

Visual Direction

The beat rapidly transitions from a flash graphic to a dynamic picture-in-picture presentation. The main screen shifts from a polished product shot to an interactive software interface, always anchored by the enthusiastic speaker in a smaller frame. Close-ups on UI elements illustrate the tool's functionality.

02

Text Card

IN 00:05.0 · OUT 00:05.9 · DUR 1.0s

You gotta
00:00.000:05.000:05.900:39.8

A brief flash of white with horizontal blurred lines across the screen, a partial phrase appears.

Why it works: This ultra-short flash cut with a partial phrase acts as an abrupt visual break, re-emphasizing the previous statement and preparing the viewer for a significant reveal or demonstration, creating a sense of urgency.

Angle

Symmetrical Center

Camera

Static Locked

Lighting

Bright And Blown Out

Color

Clean White Minimal

Mood

Silent And Visual Only

03

Picture In Picture

IN 00:05.9 · OUT 00:07.9 · DUR 2.0s

check this out.
00:00.000:05.900:07.900:39.8

Man in a room speaking directly, positioned in a smaller frame, while a professional studio shot of a beige dirt bike on a white background occupies the main screen.

Why it works: This cut immediately transitions from abstract text to a concrete visual of a product, using a picture-in-picture to maintain the speaker's direct address. This visually demonstrates the product while the speaker provides context, creating a strong 'show and tell' effect.

Angle

Eye Level Straight

Camera

Static Locked

Lighting

Soft Window Light

Color

Clean White Minimal

Mood

Excited And Fast-paced

Product

Beige dirt bike centrally displayed on a clean white studio background.

04

Screen Recording

IN 00:07.9 · OUT 00:08.9 · DUR 1.0s

So all I had
00:00.000:07.900:08.900:39.8

A digital interface of a design tool showing two dirt bikes, one green and one beige, and a section for 'LUMA DIRT BIKE - PHOTOGRAPHY GUIDELINES'. The man speaks from a smaller frame at the top right.

Why it works: This cut shifts the focus to the actual AI tool's interface, showing the 'behind the scenes' of how images are generated. The smaller talking head keeps the personal connection while the primary visual demonstrates the tool's capabilities, bridging explanation and demonstration.

Angle

Screen Capture Angle

Camera

Scroll Capture

Lighting

Screen Glow

Color

Cool Clinical White

Mood

Excited And Fast-paced

Product

Digital images of two dirt bikes (green and beige) displayed in a design tool interface.

05

Screen Recording

IN 00:08.9 · OUT 00:10.9 · DUR 2.0s

add these
00:00.000:08.900:10.900:39.8

The digital interface, zoomed in on the two dirt bike images and the 'LUMA DIRT BIKE – PHOTOGRAPHY GUIDELINES' section. A mouse cursor points at the green bike.

Why it works: By zooming in and highlighting the initial inputs (the two dirt bike images), this shot clearly illustrates the simple starting point for the AI generation process. The cursor reinforces the user interaction, making the process feel accessible and straightforward.

Angle

Screen Capture Angle

Camera

Slow Zoom In

Lighting

Screen Glow

Color

Clean White Minimal

Mood

Instructional And Patient

Product

Digital images of a green and beige dirt bike displayed on a white background within the interface.

BEAT 03

Process Setup

CONTEXT00:10.200:18.8 · 4 shots

The speaker lays out the exact workflow: “all I had to do was add these two images… and I gave it some photography guidelines.” This turns the moment into a procedural recipe, telling the viewer what actions to take in what order.

Why this works:This leverages Process Setup by making the task feel step-by-step and doable—“add… and… give it” reduces mental effort and uncertainty. It also uses Specificity Bias through concrete inputs (“two images,” “two different dirt bikes,” “photography guidelines”), which makes the method feel more replicable, so viewers stay engaged to see the full procedure.

Visual Direction

The beat features a continuous screen recording, initially zoomed in on specific input parameters like 'PHOTOGRAPHY STYLE' and 'LIGHTING.' It gradually zooms out, progressively revealing a vast grid of diverse, high-quality AI-generated images of dirt bikes, showcasing the output's quantity and variety.

06

Screen Recording

IN 00:10.9 · OUT 00:11.9 · DUR 1.0s

dirt bikes.
00:00.000:10.900:11.900:39.8

The digital interface, zoomed in further on the 'PHOTOGRAPHY STYLE' and 'LIGHTING' guidelines sections of the Luma Dirt Bike tool.

Why it works: This shot emphasizes the level of control and detail users can input, showcasing the AI's ability to understand specific photography styles and lighting conditions. This suggests advanced capabilities and customizable outputs, adding to the tool's perceived value.

Angle

Screen Capture Angle

Camera

Slow Zoom In

Lighting

Screen Glow

Color

Clean White Minimal

Mood

Instructional And Patient

07

Screen Recording

IN 00:11.9 · OUT 00:14.8 · DUR 3.0s

photography
00:00.000:11.900:14.800:39.8

The digital interface, zoomed out slightly to show the photography guidelines and the beginning of a grid of generated images of dirt bikes in various outdoor settings. A mouse cursor moves.

Why it works: This reveals the initial results of the AI's generation, creating a 'before and after' effect by showing the guidelines and then the diverse images produced. The visible grid conveys the quantity and variety of outputs available from a simple input.

Angle

Screen Capture Angle

Camera

Slow Zoom Out

Lighting

Screen Glow

Color

Natural Green

Mood

Instructional And Patient

Product

Grid of small generated images showing dirt bikes in desert and forest environments.

08

Screen Recording

IN 00:14.8 · OUT 00:16.8 · DUR 2.0s

and just
00:00.000:14.800:16.800:39.8

The digital interface, zoomed out further to show more of the generated image grid, which depicts dirt bikes in desert and forest settings. The man's picture-in-picture is still visible.

Why it works: Continuing to zoom out expands the viewer's perception of the sheer volume and diversity of generated images, reinforcing the AI's efficiency in producing multiple creative options from minimal input. The speaker's continued presence provides a human anchor.

Angle

Screen Capture Angle

Camera

Slow Zoom Out

Lighting

Screen Glow

Color

Warm Earth Tones

Mood

Excited And Fast-paced

Product

Grid of generated images featuring dirt bikes, with riders, in different outdoor environments.

09

Screen Recording

IN 00:16.8 · OUT 00:18.8 · DUR 2.0s

to get all
00:00.000:16.800:18.800:39.8

The digital interface, zoomed out almost completely to show a large grid of generated dirt bike images across multiple categories (Desert + Forest, Mountain + Canyon). The man in the top right.

Why it works: The final wide view of the image grid delivers a strong impact, visually communicating the extensive output of the AI. This 'wow' moment showcases the tool's ability to generate a complete suite of assets from a single prompt, highlighting its productivity.

Angle

Screen Capture Angle

Camera

Slow Zoom Out

Lighting

Screen Glow

Color

Warm Earth Tones

Mood

Excited And Fast-paced

Product

Extensive grid of generated dirt bike images, featuring different models and riders in diverse outdoor settings.

BEAT 04

Feature Cascade

DELIVERY00:18.800:26.9 · 4 shots

The speaker delivers a rapid-fire feature cascade: “all of these image assets… four locations, four aspect ratios for two bikes with one prompt.” This stacks concrete outputs in quick succession, making the result feel like a high-yield payoff from a single input (“one prompt”).

Why this works:This leverages Feature Cascade to create value density—each new number (“four locations,” “four aspect ratios,” “two bikes”) adds another reason to keep watching. It also uses Specificity Bias: the exact counts make the claim feel measurable and credible, so the viewer mentally recalculates what they could generate too.

Visual Direction

The visual strategy employs continuous screen recording with vertical scrolling through extensive grids of AI-generated images, then transitions to playing video assets. The picture-in-picture speaker directs attention as the scroll highlights diverse locations, aspect ratios, and multiple bike models, culminating in a dynamic video reveal.

10

Screen Recording

IN 00:18.8 · OUT 00:21.8 · DUR 3.0s

I was able
00:00.000:18.800:21.800:39.8

The digital interface showing a grid of generated dirt bike images. The screen is scrolling down, highlighting the variety of locations and aspect ratios.

Why it works: This scrolling action emphasizes the customization capabilities of the AI, specifically showcasing the different locations and aspect ratios generated. It reinforces the tool's versatility for various marketing needs and platform requirements.

Angle

Screen Capture Angle

Camera

Scroll Capture

Lighting

Screen Glow

Color

Warm Earth Tones

Mood

Instructional And Patient

Product

Grid of generated images, showing dirt bikes in four distinct locations and four aspect ratios.

11

Screen Recording

IN 00:21.8 · OUT 00:24.8 · DUR 3.0s

for two bikes with one prompt and after that,
00:00.000:21.800:24.800:39.8

The digital interface, continuing to scroll down through generated image assets, featuring two different bike models across various settings. The man in the top right.

Why it works: Highlighting 'two bikes with one prompt' showcases the AI's efficiency in generating multiple product variations simultaneously. This emphasizes time-saving and scalability, crucial selling points for content creation.

Angle

Screen Capture Angle

Camera

Scroll Capture

Lighting

Screen Glow

Color

Natural Green

Mood

Excited And Fast-paced

Product

Grid of generated images showing two distinct dirt bike models in various settings.

12

Screen Recording

IN 00:24.8 · OUT 00:26.7 · DUR 2.0s

some of
00:00.000:24.800:26.700:39.8

The digital interface, scrolling down through more generated image assets. Some images have green checkmarks, indicating selection. The man in the top right.

Why it works: The visual of selected images with checkmarks reinforces the idea of choice and curation within the tool, suggesting that users can easily pick their favorite outputs. This highlights user control and the refinement process.

Angle

Screen Capture Angle

Camera

Scroll Capture

Lighting

Screen Glow

Color

Warm Earth Tones

Mood

Excited And Fast-paced

Product

Grid of generated dirt bike images, some with green checkmarks indicating selection.

13

Screen Recording

IN 00:26.7 · OUT 00:28.7 · DUR 2.0s

and was able to turn it
00:00.000:26.700:28.700:39.8

The digital interface, scrolling down to reveal generated video assets. The first video asset is playing, showing a dirt bike moving through a desert landscape. The man in the top right.

Why it works: This cut introduces the video generation capability, showcasing the tool's versatility beyond static images. Seeing the generated video playing immediately demonstrates a more dynamic and engaging output, expanding the product's utility.

Angle

Screen Capture Angle

Camera

Scroll Capture

Lighting

Screen Glow

Color

Warm Earth Tones

Mood

Excited And Fast-paced

Product

A generated video asset of a dirt bike moving through a desert environment, playing within the interface.

BEAT 05

Overwhelm → Control

SHIFT00:26.900:33.9 · 3 shots

The speaker reframes the process from “messy trial-and-error” into a controlled workflow: after picking favorites, they “turn it into video” and “test multiple models at the same time.” This shifts the viewer’s mental model from randomness to an intentional system for running parallel experiments.

Why this works:This leverages Overwhelm → Control by replacing the perceived chaos of testing with a clear, manageable method (pick favorites → convert → run simultaneous tests). The “multiple models at the same time” detail creates a sense of command and efficiency, making the next step feel doable rather than daunting.

Visual Direction

This beat rapidly transitions between close-ups of two distinct, high-quality AI-generated video assets playing within the interface, before cutting to a static graphic. The graphic features the speaker, partially obscured by a text box emphasizing key claims, signaling a shift to reinforcing benefits.

14

Screen Recording

IN 00:28.7 · OUT 00:30.7 · DUR 2.0s

multiple models
00:00.000:28.700:30.700:39.8

The digital interface, zoomed in on a generated video asset, showing a rider on a dirt bike kicking up dust in a desert landscape during golden hour. The man in the top right.

Why it works: Zooming into a single generated video allows the viewer to appreciate the quality and dynamic nature of the AI's video output. The text overlay emphasizes the ability to generate 'multiple models' within this video format, highlighting further versatility.

Angle

Screen Capture Angle

Camera

Slow Zoom In

Lighting

Screen Glow

Color

Warm Golden

Mood

Excited And Fast-paced

Product

A generated video of a dirt bike with a rider, prominently displayed and playing within the interface.

15

Screen Recording

IN 00:30.7 · OUT 00:31.7 · DUR 1.0s

at the same time.
00:00.000:30.700:31.700:39.8

Another generated video asset playing within the digital interface, showing a different rider and dirt bike model in a desert environment with cacti.

Why it works: This rapid cut to a second distinct video further reinforces the AI's speed and capability to generate diverse video content. It visually proves the claim of producing 'multiple models at the same time,' demonstrating efficiency.

Angle

Screen Capture Angle

Camera

Static Locked

Lighting

Screen Glow

Color

Warm Golden

Mood

Excited And Fast-paced

Product

A generated video of a dirt bike with a rider in a desert landscape with cacti, playing prominently.

16

Talking Head Classic

IN 00:31.7 · OUT 00:34.6 · DUR 3.0s

And all this was without any notes, all I had
00:00.000:31.700:34.600:39.8

Back to the man speaking directly to the camera in a room. A large white rectangular box appears below him, partially obscuring the background.

Why it works: Returning to the talking head shot, the speaker directly addresses the ease of use, emphasizing that no complex notes or instructions were needed. This builds confidence in the product's accessibility and user-friendliness.

Angle

Eye Level Straight

Camera

Static Locked

Lighting

Soft Window Light

Color

Neutral Palette

Mood

Matter Of Fact

BEAT 06

Before/After Proof

VALIDATION00:33.900:38.2 · 1 shots

The speaker validates the method by emphasizing a constraint-free result: “And all this was without any nodes.” This frames the outcome as achieved without a commonly assumed requirement, so the viewer mentally compares “with nodes” vs “without nodes” and registers the method as more capable than expected.

Why this works:This leverages Before/After Proof by contrasting the implied baseline (“with nodes”) against the stated reality (“without any nodes”). That contrast creates validation through expectation violation: if the usual dependency isn’t needed, the viewer’s trust increases because the result feels harder to fake and more impressive.

Visual Direction

A single, deliberate screen recording shot begins with a close-up of a dynamic AI-generated video playing. It then slowly pulls back, progressively revealing the vast array of other generated assets, including images, creating a wide-angle overview of the tool's comprehensive output.

17

Screen Recording

IN 00:34.6 · OUT 00:39.6 · DUR 5.0s

was just talk tell it what I wanted I was able all of these assets.
00:00.000:34.600:39.600:39.8

The digital interface, showing a playing generated video of a rider on a dirt bike in the desert, then zooming out to reveal more generated assets. The man's picture-in-picture is in the top right.

Why it works: This final shot effectively brings the narrative full circle, starting with a live demonstration and then pulling back to show the vast output generated from simple commands. It's a powerful visual summary of the tool's efficiency and comprehensive asset generation capabilities.

Angle

Screen Capture Angle

Camera

Pull Back

Lighting

Screen Glow

Color

Warm Earth Tones

Mood

Excited And Fast-paced

Product

A playing generated video, then a grid of multiple generated images and videos of dirt bikes.

BEAT 07

Lesson

CLOSE00:38.200:39.8 · 1 shots

It lands a simple “how it works” takeaway: “All I had to do was just talk to the agent, tell it what I wanted and I was able to get all of these assets.” The beat compresses the process into one effortless sequence (talk → tell → get assets), turning the story into a reusable rule.

Why this works:This leverages **Specificity Bias** (the exact steps “talk… tell it… get all of these assets” feel actionable and credible) and **Cognitive Fluency** (the process sounds easy and low-effort). Because the viewer can mentally rehearse the steps immediately, **Actionability Heuristic** kicks in: it feels like something they can replicate right away, so they stay engaged with the method rather than just the story.

Visual Direction

This sustained screen recording shows a playing AI-generated video and then deliberately pulls back to reveal a comprehensive grid of diverse image and video assets, showcasing the tool’s extensive output capacity. The speaker remains visible in a picture-in-picture frame.

18

Screen Recording

IN 00:34.6 · OUT 00:39.6 · DUR 5.0s

was just talk tell it what I wanted I was able all of these assets.
00:00.000:34.600:39.600:39.8

The digital interface, showing a playing generated video of a rider on a dirt bike in the desert, then zooming out to reveal more generated assets. The man's picture-in-picture is in the top right.

Why it works: This final shot effectively brings the narrative full circle, starting with a live demonstration and then pulling back to show the vast output generated from simple commands. It's a powerful visual summary of the tool's efficiency and comprehensive asset generation capabilities.

Angle

Screen Capture Angle

Camera

Pull Back

Lighting

Screen Glow

Color

Warm Earth Tones

Mood

Excited And Fast-paced

Product

A playing generated video, then a grid of multiple generated images and videos of dirt bikes.

Ad Intel

The intelligence layer behind the shots — how this creative classifies, the role it plays in Luma AI’s media mix, and the structural signature that makes it identifiable at a glance.

Classification DNA

  • VerticalSaaS & Software90%
  • Creative FormatTalking Head Screen95%
  • Marketing AngleHow To Tutorial72%

Intelligence Profile

Video Type

Product Demo

Platform Mode

FB Product First

Behavioral Role

Demonstrate Method

Offer State

Offer Present

0:40Duration
7Beats
18Cuts
2.6Cuts / Beat
6/10Energy

What the intel says:A Product Demo creative running in FB Product First mode with the offer present puts this at the consideration stage — it is not introducing the brand, it is closing the gap between “I've seen this” and “I'd try this.” The 2.6 cuts-per-beat rate is the tell: fast enough to stay alive on feed, slow enough to stay believable.

Ready to put this into practice?

Create your first AI UGC video ad in minutes — no filming, no actors, no editing.

Try Klip Kanvas free

More in this section

Ready to make ads like these?

Paste a product link and Klip Kanvas writes the script, casts the creator and renders the ad — no filming, no actors, no editing.