Blog

Why your ai avatar sounds like a badly read teleprompter, and the six rules that fix the script

The script is not the problem because the model is weak. The script is the problem because it was written for a voice and is being said by a face, and a face is billed by the second while a voice is billed by the minute. That single difference changes every sentence you write.

Ricardo AlmeidaFounder20 min read
A translucent golden facial mask floating above a strip of film with glowing second markers fading along it.

Narration is billed per minute, the face is billed per second, and that is the whole problem

Start with the meter, because everything else in this article comes out of it. Narration is billed per minute of finished video, not per voice: 3 credits a minute in economy mode and 24 in premium, which is 36 and 288 credits across a twelve minute video. A second character with a completely different voice costs zero extra, because the meter counts minutes of video and does not care how many people are talking in them. A narration script can afford to breathe, to repeat itself, to take the long way around.

The avatar does not work like that at all. Lip sync is billed per second of face on screen, and the rate depends on the engine: Kling Avatar Std costs 32 credits a second, Kling Avatar Pro 64, HeyGen 80 and OmniHuman 1.5 112. Nothing in that meter asks whether the second was useful. A second of silence, a second of breathing in, a second of smiling at the end and a second carrying the sharpest sentence in the video all cost exactly the same.

Put numbers on it and the argument stops being abstract. Ninety seconds of face on Kling Avatar Std cost 2,880 credits, almost three times the entire twelve minute video in economy mode, which costs 1,008. The finished piece, twelve minutes of video with ninety seconds of face inside it, closes at 3,888 credits. The face is a minute and a half of the runtime and roughly three quarters of the bill.

The version people ask for first is the talking head from start to finish, and that one deserves to be priced out loud. Twelve minutes is 720 seconds of face, which on Kling Avatar Std is 23,040 credits of lip sync, plus the 1,008 the video itself costs, for 24,048 credits. That is twenty four times the same video with no face on screen at all. On the heavier engines it climbs further, because OmniHuman 1.5 at 112 credits a second puts those same twelve minutes in a different order of magnitude.

So here is the sentence that organizes everything below. The avatar text is the most expensive text in your video per second, which means it also has to be the densest. Before you commit anything, generate a 15 second test for 480 credits and approve the tone, the framing and the voice there, not on the finished piece. In FalconVid the engine is a choice per project, so the same script can be proved cheaply and then delivered on a heavier engine once it has earned it.

  • Narration: 3 credits a minute in economy, 24 in premium, so 36 or 288 credits for twelve minutes.
  • Lip sync: Kling Avatar Std 32 credits a second, Kling Avatar Pro 64, HeyGen 80, OmniHuman 1.5 112.
  • Ninety seconds of face on Kling Std: 2,880 credits, against 1,008 for the whole economy video.
  • Economy video with ninety seconds of face inside it: 3,888 credits.
  • Talking head across the full twelve minutes on Kling Std: 24,048 credits.
  • A 15 second test costs 480 credits, and that is where tone, framing and voice get approved.

Your script is not broken, it was written for a voice and is being said by a face

The sequence is always the same. The narration script is finished, it reads well out loud, somebody drops it into the avatar and the result looks like a person reading a teleprompter badly. The first instinct is to blame the model, then the engine, then the voice. The text is what is wrong, and it is wrong in a very specific way: it was written to be heard while something else was on screen, and now it is being said by the thing on screen.

A narration describes. It says the channel needs three things, the market changed, most creators fail right here. That works because the viewer is looking at footage, at a chart, at a cut of b-roll, and the voice is a layer over the picture. The moment a face appears, the same sentence stops being a layer and becomes a person making a general observation to nobody in particular, which is exactly the sensation everyone ends up calling robotic.

There is also nowhere to hide. In narrated video every awkward clause can be covered by a cut, a zoom, a new image landing on the beat. The avatar segment has one shot, one face and no cutaway, so every subordinate clause plays out on the mouth of somebody who is now visibly working through a sentence. The camera never blinks, which means the writing has to be clean before it is generated, not repaired afterwards.

The economics make that correction non negotiable. When narration costs 36 credits for twelve minutes, a clumsy paragraph costs nothing but a little patience. When the face costs 32 to 112 credits a second, a clumsy paragraph is a line item on the invoice. Forty wasted seconds of face on Kling Std are 1,280 credits, more than the entire economy video that surrounds them. That is why the avatar script should be a separate document from the start, with the same research and a completely different grammar.

Six rules for writing the avatar segment, and what each one saves

Write in direct address, in the second person. The narration says the channel needs a hook in the first ten seconds. The face says you have ten seconds before they leave. A face on screen delivering an impersonal sentence is the single loudest source of the teleprompter feeling, because people do not look into a camera to make general observations, they look into a camera to talk to somebody. Rewrite every sentence of the avatar segment until it has a you in it, or until it obviously could.

Then cut the sentences down to one idea each. The face has nowhere to rest a subordinate clause, so anything carrying more than two commas gets split or deleted. And a long list is not the job of the face at all: the avatar announces that there are three, and the screen shows the three. A face reciting seven items burns 40 seconds, which is 1,280 credits on Kling Std, to deliver what on screen text delivers in four seconds while the face says one line.

Spoken numbers are where the robot shows up, every single time. Twenty three thousand and forty eight credits, said out loud by a synthetic mouth, is where the illusion collapses, because the model has to hold the rhythm of a long numeral and no engine does that convincingly yet. The rule is to round in the mouth and be exact on the screen. The face says twenty four thousand, the caption writes 24,048, the viewer gets both and neither one sounds wrong.

Never reference the screen or the timeline. As I showed at minute three, in the chart beside me, in the next section, all of these weld the clip to a position it may not keep. Avatar segments get re-cut, moved, reused in a Short or dropped into a different edit of the same video, and each of those references becomes a small lie that costs 32 to 112 credits a second to regenerate. Write the segment so it stays true in any position.

Start and end on speech. The meter counts a second of silence exactly like a second of speech, so the little breath at the start and the polite smile at the end cost 32 to 112 credits each, and they are the two seconds nobody notices they are paying for. Cut the run up, cut the tail, let the first frame land on the first word and the last frame land on the last one. Across a four dose structure that habit alone pays for a second test generation on every video. In FalconVid the script comes out of an AI writer working in parallel with the narrator and the editor, and the channel DNA locks the face and the voice, so these six rules are applied to a presenter who is already the same person in every video.

  • Direct address: the narration describes, the face talks to somebody. Every sentence gets a you.
  • One idea per sentence: anything with more than two commas gets split or deleted.
  • Long lists belong to the screen: the face says there are three, the screen shows the three.
  • Round in the mouth, exact on screen: the face says twenty four thousand, the caption writes 24,048.
  • No references to screen or time, because the clip gets re-cut and may end up somewhere else.
  • Start and end on speech: silence costs the same as speech, 32 to 112 credits a second.

The four doses of face: ninety seconds that carry a twelve minute video

The structure that survives contact with the meter is not a face that comes and goes at random. It is four planned doses, placed where a face does something an image cannot, and everything between them is narration over scenes. Written this way the avatar stops being a presenter reading a video and becomes the person who shows up at the four moments that decide whether anyone stays.

The first dose is the hook, 15 to 20 seconds, and it is the one that pays the most. The promise said by somebody looking into the camera is a different object from the same promise read over a wide shot, because a face makes it a commitment instead of a caption. This is where the second person rule earns its keep, and where a 15 second test at 480 credits should be spent before anything else in the video is generated.

The second dose is the transition in the middle, around 15 seconds, and it exists to reanchor the viewer at the point where retention drops. Something happens when a video that has been narrated over footage for four minutes suddenly puts a face back on screen: the viewer looks up. That is the entire function, and it is why the dose is short. Fifteen seconds of face on Kling Std is 480 credits, which is roughly half the price of the whole economy video, spent on the single spot where you were losing people.

The third dose is the verdict, 25 to 30 seconds, and it is the one where the face is worth more than any image you could put there. An opinion, a judgement, a recommendation has an author, and a face is the cheapest way to show that the piece has one. Stock footage cannot take a position. This is the longest dose on purpose, 800 to 960 credits on Kling Std, and it is the part of the script that should be rewritten the most times before generation.

The fourth dose is the closing call, 15 to 20 seconds, a direct request delivered to the camera. Add the four together and you get 70 to 90 seconds of face, which is 2,240 to 2,880 credits on Kling Avatar Std. Everything else in the twelve minutes is narration over scenes, and that costs 36 credits in economy mode. Four doses and a narration is the shape that lets a face appear in every video you publish instead of once a quarter. In FalconVid the Studio adjusts one of those doses without rebuilding the video, 5 to 320 credits against the 1,008 to 26,760 a full remake costs, which is what makes a four dose structure safe to run every week.

  • Hook, 15 to 20 seconds: the promise said to camera. 480 to 640 credits on Kling Std.
  • Mid transition, 15 seconds: 480 credits, spent exactly where retention drops.
  • Verdict, 25 to 30 seconds: 800 to 960 credits, the judgement no image can make.
  • Closing call, 15 to 20 seconds: 480 to 640 credits, a direct request to camera.
  • Total: 70 to 90 seconds of face, 2,240 to 2,880 credits on Kling Avatar Std.
  • Everything between the doses is narration over scenes: 36 credits for the whole twelve minutes.
A horizontal golden timeline with four luminous blocks of different sizes spaced along it and the rest of the line glowing faintly.

Where FalconVid fits, and why the same face has to come back next week

An avatar only becomes a presenter through repetition, and repetition is what most people lose first. A face generated fresh each week drifts, and a viewer who cannot recognise the person on screen is watching a stock model, not a host. The channel DNA in FalconVid locks the same face, the same voice and the same identity across every video of that channel, so the ai avatar accumulates recognition instead of restarting from zero every Monday.

Around that sit the choices that decide the bill. You pick the engine per project, economy or premium, you can clone a voice so the face speaks in a voice that belongs to the channel, and narration runs in up to 63 languages. The 15 second test at 480 credits belongs in this workflow rather than beside it, because approving tone and framing on a cheap clip is the difference between one generation and three.

Then there is the day the segment comes out wrong, which happens to everyone. The Studio adjusts the avatar segment without rebuilding the video, and that is a 5 to 320 credit operation against the 1,008 to 26,760 that a full remake costs depending on the quality mode. Shortening a dose, swapping the media around it or replacing the music is a repair, not a rerun, and on a piece carrying 2,880 credits of face the difference between those two words is the entire month.

Languages behave differently once an avatar is in the video, and it is worth knowing before you plan a catalog. Duplicating the project into another language repays only the narration, which is 36 credits in economy mode and 288 in premium for twelve minutes. In a video with an avatar, though, what rules the price of the copy is the seconds of face, not the language, which is precisely why the four dose structure is what makes a multilingual catalog affordable at all.

The plan arithmetic follows from there. Starter is $47 with 15,000 credits, 1 channel and 2 simultaneous generations: three videos with the full four doses at 3,888 each plus three narrated videos at 1,008 each closes at 14,688 credits, and credit packs that never expire cover an avatar heavy month, Boost 2,000 for $9, Power 6,000 for $24 and Mega 14,000 for $49. Pro is $97 with 30,000 credits, 5 channels and 5 simultaneous generations, which doubles that arithmetic, and it runs up to Scale at $997 with 320,000 credits, 50 channels and 50 simultaneous generations.

The same writing rules carry into the two formats where a face is the format. A UGC style product testimonial is four doses with the proportions changed, and a VSL modelled for any language is the verdict dose stretched into the whole piece with the numbers on screen instead of in the mouth. Every creation feature is on every plan, so what changes between plans is volume, channels, simultaneous generations, the dedicated server from Pro, the Senior Analyst from Pro with 7 free days on Starter, and support.

Two avatars, one cover, and the cheapest place to put a face

The question that comes right after the first video works is whether you can have two of them. You can, and the price depends entirely on one detail. Two avatars alternating, one talking while the other is off screen, add up to the same 90 seconds and the same 2,880 credits. Two avatars on screen at the same time double it to 5,760, because each face runs its own clock and the meter counts faces, not people in the story.

That gives you the rule for cast. An ensemble does not get more expensive for having more characters, it gets more expensive for having more simultaneous face. In narration the same principle runs in your favour, since a character with a distinct voice costs zero extra and a two voice dialogue over scenes is free. So write the dialogue as alternating shots rather than as a two shot, and the same scene costs half.

The cheapest application of a face is not in the video at all, it is on the cover. In FalconVid a face on the thumbnail is billed as an image, 479 credits in economy mode and 214 in premium, while the same face speaking for 90 seconds costs 2,880. On a faceless channel the cover is where a face buys the most attention per credit, because it is doing the one job that decides whether the video is opened, and it never spends a second of lip sync doing it.

One last thing that belongs in the plan and not in the panic. YouTube asks for the altered or synthetic content disclosure when a video shows a realistic person who does not exist or an event that did not happen, which is precisely what a talking avatar is. Marking it is a checkbox at upload, it is cheap, and it does not sink reach by itself. Deciding this once for the channel is worth more than deciding it in a hurry on the video that finally worked.

What to do with this on your next video

The whole article reduces to one habit. Stop extracting the avatar segment from the narration script and start writing it as its own document: second person, one idea per sentence, lists on the screen, numbers rounded in the mouth, no references to position, first frame on the first word. Ninety seconds written that way are worth more than four minutes of face written the other way, and they cost a third as much.

Doing it by hand is where the arithmetic gets grim. Writing to camera, lighting a room, recording, discarding the takes where you stumbled and cutting the rest into 90 usable seconds is most of an afternoon, per video, before anything is edited. At four videos a month that is four afternoons that exist only so a face appears, and it is the reason most channels that plan a presenter end up publishing a slideshow with a voice on top.

On the automated side those 90 seconds come out of a calendar you approve once. The AI specialists work in parallel, the video is ready in up to 30 minutes, the channel DNA keeps the same face and voice from one video to the next and the Studio repairs a dose for 5 to 320 credits instead of the 1,008 to 26,760 a remake costs. A twelve minute video costs 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium, and the four doses add 2,880 on top in economy.

So the ninety seconds is not a limitation you accept, it is a budget decision that happens to also be the better edit. Write the four doses, spend the 480 credits on a 15 second test before you commit, put the face on the cover of everything, and let narration over scenes carry the other ten and a half minutes. Then run that same structure across as many channels as you want, which is the part that stopped being a question of hours a long time ago.

FAQ

Got questions? We've got answers.

How do I write a script for an ai avatar so it does not sound like a teleprompter?

Write it in the second person, with one idea per sentence, and cut anything carrying more than two commas. A narration describes and a face talks to somebody, so an impersonal sentence delivered by a face is the loudest source of that robotic feeling. Move long lists to on screen text, round long numbers in the mouth while the caption stays exact, and let the clip start on the first word instead of on a breath. In FalconVid the avatar segment and the narration come out of the same project, so rewriting one of them is a writing decision and not a second production.

How many seconds of ai avatar should a twelve minute video have?

Around 70 to 90 seconds, split into four doses: a hook of 15 to 20 seconds, a mid transition of 15, a verdict of 25 to 30 and a closing call of 15 to 20. On Kling Avatar Std that is 2,240 to 2,880 credits. Everything between the doses is narration over scenes, which costs 36 credits across the whole twelve minutes in economy mode, so the face lands only where it does something an image cannot.

How much does a minute of ai avatar cost in credits?

Lip sync is billed per second of face on screen, so a minute is 1,920 credits on Kling Avatar Std at 32 a second, 3,840 on Kling Avatar Pro at 64, 4,800 on HeyGen at 80 and 6,720 on OmniHuman 1.5 at 112. Narration is billed the other way, per minute of video rather than per voice, at 3 credits a minute in economy and 24 in premium. That difference is why a face is planned in doses and a voice is not.

Should the ai avatar read the whole video?

Almost never, and the price says why. Twelve minutes of face is 720 seconds, which on Kling Avatar Std is 23,040 credits of lip sync plus the 1,008 the video costs, for 24,048 credits, twenty four times the same video with no face on screen. A full talking head also gives the editor nothing to cut to for twelve minutes. Four doses of 15 to 30 seconds deliver the presence and leave the budget for the rest of the catalog.

Do I have to appear on camera or record anything to use an ai avatar?

No. FalconVid researches, writes, narrates, edits, captions and publishes, and what you approve is the calendar, not each script. The face and the voice are locked by the channel DNA so the same presenter comes back in every video, voice cloning is available and narration runs in up to 63 languages. A twelve minute video costs 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium, and ninety seconds of face adds 2,880 on Kling Avatar Std.

What if the avatar segment comes out wrong?

In FalconVid you fix that segment instead of rebuilding the video. The Studio adjusts the avatar dose, the media around it or the music for 5 to 320 credits, against the 1,008 to 26,760 a full remake costs depending on the quality mode. The cheaper habit is to approve a 15 second test for 480 credits before committing, because tone, framing and voice are decided there, and after that the finished piece is a formality rather than a gamble.

Can I have two ai avatars talking on screen at the same time?

Yes, but the price depends on how you stage it. Two avatars alternating, one on screen while the other is not, add up to the same 90 seconds and the same 2,880 credits on Kling Avatar Std. Two faces visible at once double it to 5,760, because each face runs its own clock. A cast does not get expensive for having more characters, it gets expensive for having more simultaneous face, so write the dialogue as alternating shots.

Do I have to label a video made with an ai avatar on YouTube?

YouTube asks for the altered or synthetic content disclosure when a video shows a realistic person who does not exist or an event that did not happen, and a talking avatar is exactly that. It is a checkbox at upload, it costs nothing and it does not sink reach on its own. Decide it once for the channel rather than in a hurry on the video that finally worked, and keep the same answer across the catalog.

Ninety seconds of face, and the engine carries the other ten and a half minutes

FalconVid researches, writes, narrates, edits, captions and publishes to YouTube, Instagram, TikTok, Rumble and Facebook in up to 63 languages, from a calendar you approve once, with AI specialists working in parallel and a video ready in up to 30 minutes. The channel DNA keeps the same face and the same voice in every video, so the ai avatar accumulates recognition instead of restarting every week, and the Studio repairs an avatar dose for 5 to 320 credits instead of the 1,008 to 26,760 a remake costs. A twelve minute video costs 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium. Starter at $47 with 15,000 credits, 1 channel and 2 simultaneous generations is around 10 to 12 videos a month mixing economy with one in balanced. Pro is $97 with 30,000 credits, 5 channels and 5 simultaneous generations, up to Scale at $997 with 320,000 credits, 50 channels and 50 simultaneous generations. Every creation feature on every plan, with volume, channels, simultaneous generations, the Senior Analyst (from Pro) and support changing by plan, a 7 day trial with 2,000 credits and a 7 day guarantee.

Create my channel now

Charged today · 7-day guarantee · Cancel anytime

Keep reading