Blog

How Many Scenes Does a 12 Minute Video Have? 720 Seconds, 72 to 120 Cuts, and the Rhythm That Decides Retention in 2026

Not a matter of taste. The arithmetic of 720 seconds, the band that works for narrated video, the real cost of each scene type, and the one mistake most ai video generators still make.

Ricardo AlmeidaFounder13 min read
A long filmstrip broken into uneven segments, tightly packed at the start and stretching into wider calm blocks further along.

The arithmetic almost nobody does: 720 seconds divided by your cut length

Twelve minutes is 720 seconds, and that single number answers the question. Divide it by how long one visual stays on screen and you have your scene count. A change every 5 seconds is 144 scenes. Every 8 seconds is 90. Every 12 seconds is 60. Every 20 seconds is 36. Same script, same narration, same runtime, four completely different production jobs with four different price tags. Those same 720 seconds are what a clip generator cannot fill on its own, and the arithmetic between a prompt and a full video counts the generations and the six other jobs the count leaves out.

The band that works for a narrated faceless channel is 6 to 10 seconds per scene, which lands between 72 and 120 scenes. Under 4 seconds the episode turns into a music video: the viewer feels rushed, tires around minute three and leaves without being able to explain why. Over 15 seconds the eye gives up before the ear does, because the narration is still moving while the screen has stopped.

Ninety scenes is the honest middle of that band and the number to plan against. It is also the number that makes people underestimate the work, because 90 scenes means 90 decisions about what the viewer sees, 90 searches, 90 trims and 90 sync points against the narration. Most creators who abandon a long form channel did not run out of ideas. They ran out of patience around scene 40 of video number six.

  • 12 minutes is 720 seconds, the base of every scene calculation.
  • A visual every 5 seconds is 144 scenes, every 8 seconds is 90.
  • Every 12 seconds is 60 scenes, every 20 seconds only 36.
  • The working band for narrated video: 6 to 10 seconds, 72 to 120 scenes.

The rule almost nobody applies: the cut rhythm is not constant

Here is where most channels lose. They pick a number, 8 seconds, and cut every 8 seconds for twelve minutes. It is tidy, it is trivial to automate and it is wrong. The first 30 seconds want cuts of 2 to 4 seconds, because that is the hook: roughly 10 scenes before the viewer has decided anything.

The explanatory body is a different animal. Once the promise is made and the viewer has committed, 8 to 12 seconds per scene is comfortable and the visuals stop competing with the narration. Then comes the part almost everyone gets wrong: the turning point of the story wants a long shot, 15 to 20 seconds, because visual silence is what gives a sentence weight. Cut through the revelation and you flatten it.

A blind cut every X seconds is the signature of lazy automation, and it appears in the retention graph as a sawtooth: small drops at regular intervals, nothing dramatic, just a slow bleed. The map that works on a twelve minute episode is about 10 scenes in the first 30 seconds, 60 to 70 in the body at 8 to 12 seconds, 3 or 4 anchor moments held for 15 to 20, and a closing block back at 6 to 8. In FalconVid the cut lands on the sentence and on the narration beat instead of on a timer, so the rhythm breathes instead of ticking.

  • First 30 seconds: 2 to 4 seconds per scene, around 10 scenes of hook.
  • Explanatory body: 8 to 12 seconds, the comfortable cruising speed.
  • Turning points: 15 to 20 seconds, because visual silence adds weight.
  • A fixed interval reads as a sawtooth of small regular drops.
An editing timeline whose blocks change width, ten narrow glowing blocks at the start and three wide calm blocks near the end.

A scene is not always a clip: five types and what each one costs

The word scene tricks people into thinking clip, and clips are the most expensive item on the list. A scene is any visual unit that holds the screen for one beat: a still image with a slow pan or zoom, an animated clip, a chart or data card, a map, or a text card. A twelve minute documentary made of 100 percent animated clips is not better, it is more expensive and usually more restless.

The honest split for narrated long form is roughly 60 to 70 percent stills with slow movement, 15 to 25 percent animated clips reserved for moments that need motion, and 10 to 15 percent charts, maps and text cards. Animating everything is waste: most b roll does its job with a 4 percent zoom over 8 seconds and nobody notices the difference. Motion should mean something happened, not that the budget existed.

That distribution is exactly what the quality mode buys in FalconVid, and it is why the same twelve minute video costs 1,008 credits in economy, 8,676 in balanced and 26,760 in premium. Spread over 90 scenes that is about 11 credits per scene in economy, roughly 96 in balanced and near 297 in premium, which in dollars is about $0.04, about $0.30 and about $0.93 per scene. The runtime never changes. What changes is how many scenes are generated and animated, and by which engine.

  • Still with slow pan or zoom: cheapest scene and the backbone of b roll.
  • Animated clip: reserve it for motion that actually means something.
  • Chart or data card: cheap, and it buys attention on every number.
  • Map: the strongest scene type for history, geopolitics and travel.
  • Same video: 1,008 credits economy, 8,676 balanced, 26,760 premium.

The script decides the cut: 1,800 words, one sentence per scene

Narration runs at about 150 words per minute in a comfortable documentary register, so a twelve minute video is roughly 1,800 words of script. Divide 1,800 words by 90 scenes and you get 20 words per scene, which at 150 words per minute is exactly 8 seconds. The band is not a stylistic preference. It falls out of the way people speak.

That is why one sentence per scene is the natural cut. A sentence is how the script already segments an idea, and a visual change on a sentence boundary feels invisible while a change mid clause feels like a glitch. A 30 word sentence wants a 12 second scene. A punchy 8 word line wants 3 seconds and works beautifully in the hook.

It also gives you a diagnostic worth more than any editing trick. If you cannot find a different visual for three consecutive sentences, the script is repeating itself and the fix belongs in the writing, not in the edit. FalconVid writes the script and plans the scenes in the same pass, so the visual plan is born from the sentence structure instead of being retrofitted onto it, and the narration never waits on a picture with nothing new to say.

  • About 150 words per minute means roughly 1,800 words in 12 minutes.
  • 1,800 words over 90 scenes is 20 words, about 8 seconds each.
  • One sentence per scene is the natural, invisible cut point.
  • Three sentences with no new visual means the script is repeating.

Who assembles 90 scenes: 6 to 10 hours of an editor, or 30 minutes of a pipeline

This is the part the tutorials skip. For each of those 90 scenes somebody decides what appears, finds the footage, downloads it, trims it, places it under the right sentence, adjusts the movement, checks the license and moves on. A competent youtube video editor spends 6 to 10 hours on a twelve minute narrated episode, about 5 minutes per scene with no breaks. At $80 to $400 per long video, publishing twice a week is a second salary.

The FalconVid pipeline does that job in parallel instead of in sequence. A researcher, a scriptwriter, a narrator, an editor and a sound designer work at the same time rather than one after another, which is why the finished video lands in up to 30 minutes instead of days. The music and sound effects library is built in, so no scene waits on licensing. Karaoke subtitles in 15 plus styles are produced with the cut, not after it, and you pick the engine per project: economy for volume, premium for the episodes that deserve it.

The objection is always the same: what if scene 47 is wrong? You watch the first version in the Studio and swap that one media, shorten the intro, change the track, or open the timeline and do it yourself. You are not rebuilding a video, you are correcting one scene out of 90. That is the whole difference between reviewing ai generated videos and editing them, and it is why the calendar is the only thing you actually approve.

  • Parallel specialists deliver the finished video in up to 30 minutes.
  • Built in music and sound effects library, no licence chase per scene.
  • Karaoke subtitles in 15 plus styles produced along with the cut.
  • The Studio replaces one scene instead of redoing the whole video.

Where an AI avatar fits in the scene map, and why it is the priciest scene

If your format puts a presenter on screen the map changes, because lip sync is billed per second of face rather than per scene. Kling Avatar Std is 32 credits per second, about $6.02 per minute. Kling Avatar Pro is 64, HeyGen is 80 and OmniHuman 1.5 is 112 credits per second, around $21.05 per minute. Face time is the most expensive real estate in the entire video, by a wide margin.

Run it against the scene budget and the decision writes itself. A single 8 second avatar scene on the cheapest engine costs 256 credits, a quarter of what the entire twelve minute video costs to produce in economy mode at 1,008 credits. Eight seconds of face buys as much as roughly three minutes of finished video. A full twelve minute talking head is 24,048 credits, around $75.34, while 90 seconds of face on the same episode totals 3,888 credits, around $12.18.

So the ai avatar belongs exactly where the rhythm already told you to slow down: the hook, the two or three turning points and the closing. The other 80 scenes stay b roll, charts and maps, which is also better television. A face that appears at the anchor moments reads as authority. A face that never leaves the screen for twelve minutes reads as a webinar, and costs six times more on the same episode.

  • Lip sync is billed per second of face, not per scene.
  • One 8 second avatar scene, 256 credits, a quarter of the whole economy video at 1,008.
  • 90 seconds of face on a 12 minute episode: 3,888 credits, around $12.18.
  • Put the face on the hook and the turning points, not on all 90 scenes.

144 cuts or 60 good scenes: the verdict for 2026

Forced to give one answer: fewer, better chosen scenes beat 144 mechanical cuts every time. A video with 60 scenes where each one says something the narration does not say holds better than a video with 144 where the screen changes for the sake of changing. Cut rhythm is information, and information you repeat is noise.

The catch is that choosing well takes time, and time is exactly what a manual workflow does not have. That is the real ceiling, and it is a ceiling of the manual route: one person can curate 90 scenes carefully for one video a week, maybe two. The moment you want four videos a week across two channels, you are choosing between careful and possible, and careful loses.

That ceiling disappears the moment production runs in parallel and the quality mode becomes a dial instead of a wall. Economy at 1,008 credits per video is the volume setting, balanced at 8,676 is the workhorse, premium at 26,760 is for the episodes that carry the channel, and you decide per video rather than per channel. Starter runs 2 generations at the same time, Scale runs 50 generations across 50 channels. Careful and many stop being opposites the second 90 scenes stop being 8 hours of your evening.

  • 60 well chosen scenes beat 144 mechanical cuts on retention.
  • The quality mode is the dial: economy 1,008, balanced 8,676, premium 26,760.
  • Choose per video, not per channel, once generations run in parallel.
  • 2 simultaneous generations on Starter, 50 on Scale.

FAQ

Got questions? We've got answers.

How many scenes does a 12 minute video have?

Between 72 and 120 in the band that works for narrated video, with 90 as the practical default. Twelve minutes is 720 seconds, so a visual change every 8 seconds gives 90 scenes, every 5 seconds gives 144 and every 12 seconds gives 60. Under 4 seconds per scene the episode feels like a music video, and over 15 seconds the eye leaves before the ear does.

How long should each scene be?

Between 6 and 10 seconds on average, but never at a fixed interval. The first 30 seconds want 2 to 4 second cuts because that is the hook, the explanatory body is comfortable at 8 to 12 seconds, and the turning points want a long shot of 15 to 20 seconds. A constant interval shows up in the retention graph as a sawtooth of small regular drops.

Do I need to animate every scene?

No, and animating everything is the most common way to burn a budget. Around 60 to 70 percent of a narrated long form video works with still images and a slow pan or zoom, roughly 15 to 25 percent deserves animated clips, and the rest is charts, maps and text cards. That distribution is also what separates 1,008 credits in economy from 26,760 in premium for the same twelve minutes.

Will I have to edit the video myself?

No. The FalconVid pipeline writes, narrates, cuts the 90 scenes, adds music, sound effects and karaoke subtitles and publishes to 5 networks, with the finished video arriving in up to 30 minutes. If scene 47 is wrong you open the FalconVid Studio, watch the first version and swap that single media, shorten the intro or change the track, or open the timeline and do it yourself. You approve the calendar, not each cut.

How many scenes should the first 30 seconds have?

About 8 to 12, which means cuts of 2 to 4 seconds. The hook is the only part of the video where visual density is worth more than visual meaning, because you are buying attention before the viewer has decided anything. Once the promise is delivered, slow down to 8 to 12 seconds per scene or the audience burns out around minute three.

How much does each scene cost?

Divide the video cost by the scene count. A twelve minute video is 1,008 credits in economy, 8,676 in balanced and 26,760 in premium, so across 90 scenes that is roughly 11, 96 and 297 credits per scene, about $0.04, about $0.30 and about $0.93. The runtime never changes, only which engine generates and animates each scene, which is why the mode is a per project decision and not a plan upgrade.

Do I need to appear on camera or use an AI avatar in the scenes?

Neither is required. A faceless channel runs entirely on b roll, stills, charts and maps with narration on top, which is the standard format for documentary, explainer and top 10, and it is what FalconVid builds by default. If you do want a presenter, in FalconVid lip sync is billed per second of face: 32 credits per second on Kling Avatar Std, so an 8 second avatar scene alone is 256 credits against 1,008 for the whole twelve minute video in economy, and a 90 second appearance totals 3,888 credits, around $12.18, against 24,048 credits for a full talking head. Put the ai avatar on the hook and the turning points.

Ninety scenes, cut on the sentence, finished in 30 minutes

In FalconVid the pipeline writes the script, breaks it into scenes on the sentence instead of a timer, narrates in up to 63 languages, animates only what deserves motion, adds music, sound effects and karaoke subtitles and publishes to 5 networks, in parallel, while you only approve the calendar. Starter is $47 with 15,000 credits, about 10 to 12 videos a month mixing economy with one in balanced, or 14 in pure economy at 1,008 credits each, with balanced at 8,676 and premium at 26,760 for the episodes that carry the channel. Pro is $97 with 30,000 credits and 5 channels, up to Scale at $997 with 50 channels and 50 simultaneous generations. Every creation feature on every plan, 7 day trial with 2,000 credits and a 7 day guarantee.

Create my channel now

Charged today · 7-day guarantee · Cancel anytime

Keep reading