Two billing rules, and they point in opposite directions
Everyone arrives at this question with the same instinct: a cast must be expensive, because more characters mean more work and more bill. Inside the engine the opposite happens, for one very specific reason. Narration is charged per minute of finished video. Lip sync is charged per second of face on screen. Those two units decide the format before any argument about taste even starts.
Start with narration. On a 12 minute video, the premium voice costs 24 credits per minute, which is 288 credits for the whole video, about $0.90 at a credit value of $0.003133. The economy voice costs 3 credits per minute, which is 36 credits for the same 12 minutes, about $0.11. Now look at what is missing from that formula: the number of voices. One narrator reading straight through or six characters trading lines, it is still 12 minutes of video, so it is still 288 credits in the premium voice.
Now lip sync, which is what an AI avatar actually is. It is charged by the second of face on screen, because every one of those seconds is a generation. Kling Avatar Standard costs 32 credits per second, about $6.02 per minute of face. Kling Avatar Pro costs 64 per second, about $12.03 per minute. HeyGen costs 80 per second, about $15.04 per minute. OmniHuman 1.5 costs 112 per second, about $21.05 per minute. There is no license per character and no monthly fee per avatar. What you are buying is screen time.
Here is the equivalence that makes the asymmetry concrete. Nine seconds of a second face on screen at Kling Avatar Standard costs 288 credits. That is exactly what full premium narration of all 12 minutes costs, for every character in the script, however many there are. Nine seconds of one extra visible face equals twelve minutes of every voice you could want. Put the other way around: a little more than one second of that same face costs 36 credits, the price of narrating the entire video with the economy voice.
So the honest sentence, the one that should drive the format decision, is this. Dialogue is nearly free and a visible cast is not. Writing four characters into a script adds zero credits to the narration line. Showing four faces adds 32 credits per second, per face, for as long as they stay in frame. Everything else in this article is a consequence of that single asymmetry, including the parts that look counterintuitive at first.
The cast arithmetic: it is not the characters, it is the simultaneous faces
Start from the baseline. A 12 minute video in economy mode costs 1,008 credits, about $3.16. Add 90 seconds of face with Kling Avatar Standard and you add 90 times 32, which is 2,880 credits, about $9.02. The video totals 3,888 credits, about $12.18. That is the reference number for a single host format with a sane amount of face time, and every comparison below is measured against it.
Now put two avatars in the same video, alternating, never together in the frame. An anchor takes 45 seconds and a reporter takes 45 seconds. Count the face seconds actually on screen: still 90, because there is only ever one face at a time. The lip sync line is still 2,880 credits and the video still totals 3,888. The second character, with its own name, its own voice and its own visual identity, added exactly zero credits to the bill.
Now put the same two avatars in the same frame for those same 90 seconds, split screen or a two shot. Two meters run at the same time. The engine is generating 180 seconds of face, so the lip sync line becomes 5,760 credits and the video totals 6,768, about $21.20. Same script, same length, same two characters, same 90 seconds of running time, and the bill went up 74 percent. Nothing changed except how many faces shared the frame.
For the outer edge of the range, take the full talking head: 12 minutes of face, which is 720 seconds. At Kling Avatar Standard that is 23,040 credits, so the video totals 24,048, about $75.34. That is 6.2 times the alternating cast version of exactly the same content, which is why the format decision is worth more than any tweak you will make to the script.
The rule that falls out of this is short and it is the whole article. A cast does not get more expensive by having more characters. It gets more expensive by having more faces on screen at the same time. Reaction cuts and alternating shots are the cheap format. Split screen is the expensive one. That happens to be the same editing grammar every well made interview show on television has used for decades, so the cheap choice is also the one that looks the most professional.
- 12 minute video in economy mode: 1,008 credits, about $3.16.
- One host, 90 seconds of face at Kling Avatar Standard: 2,880 credits, total 3,888.
- Two hosts alternating, same 90 seconds on screen: same 3,888 credits.
- Two hosts sharing the frame, 90 seconds: 180 face seconds, 5,760 credits, total 6,768.
- Full 12 minute talking head: 720 face seconds, 24,048 credits, about $75.34.
When the single host wins: recognition is built by repetition
The strongest argument for one host has nothing to do with credits. It is recognition. A viewer scrolling a feed makes the click decision in well under a second, and inside that window a familiar face is a stronger signal than any headline. After 30 or 40 videos with the same presenter, that face stops being a design choice and becomes the logo of the channel. People recognize it on the thumbnail before they read a single word of the title.
A cast dilutes exactly that. Three presenters split the recognition three ways and each one accumulates a third of the exposure. A channel with 50 published videos and one host has delivered 50 impressions of the same face to the same audience. The same channel with three hosts has delivered around 17 impressions each, which is below the point where a casual viewer starts recognizing anybody. The audience does not remember three faces from a channel it barely knows, it remembers zero.
This matters most in the first 90 days, when the channel has no subscribers, no history and no authority. A new channel needs one memorable thing, not three interesting ones. That is why the standard advice for a new faceless channel is a single voice and, if you use an avatar at all, a single face, held rigid until the audience proves it recognizes it. A cast is a format upgrade you buy after recognition exists, not a tool for creating it.
On a manual channel this is not even a decision, it is a constraint. A recurring host means the same human being available on the same afternoon, every week, forever. Two hosts means two schedules, two fees and two people who both have to show up. And even then the identity drifts: a haircut, different lighting, a new background, a bad night of sleep, and the face on the thumbnail stops matching the face people remember.
In FalconVid this drift is a project setting rather than a battle. The channel DNA holds the presenter fixed, the same face, the same voice, the same visual style and the same pacing, video after video, so those 50 impressions land on one identical face instead of 50 slight variations of it. The presenter is never unavailable, never gets a haircut and never negotiates, which is precisely why an automated channel can commit to a single recognizable host for a year without a single scheduling problem.
When the cast wins: some formats simply do not exist with one voice
There are formats that cannot be produced by a single narrator without collapsing into the flattest kind of explainer. An interview needs someone asking and someone answering. A debate needs two positions defended by two people. News needs an anchor who frames and a reporter who details. Teaching often needs a student who asks the question the viewer has in their head. Question and answer needs a question that does not come from the same mouth as the answer.
The mechanism underneath is retention, and it is simple. A voice change is an attention reset. In a 12 minute explainer the viewer's attention decays continuously, and every time the speaker changes, the brain re-orients to work out who is talking and why. That is why conversation formats hold audiences for hours while a monologue on the same subject loses them in six minutes. A dialogue exchange every 90 seconds gives you 8 resets across a 12 minute video, spread exactly where attention usually leaks.
The important part, and the reason this section is not about avatars at all: this works with voice alone. Zero faces. Two distinct voices trading lines over generated scenes and b roll cost the same 288 credits of premium narration as one narrator reading the same 12 minutes, or 36 credits in economy. It is the cheapest retention upgrade available to a faceless channel, because it does not cost anything. The cast can be entirely invisible and still do most of the work people credit to the faces.
A caveat that saves you from the worst version of this. A second character has to carry a real function, otherwise it becomes noise. If the second voice exists only to say that is amazing, tell me more, you have added a distraction, not a reset. The sharpest use is to give the second voice the viewer's own skepticism: the narrator makes a claim, the second character asks how do you know, and the answer lands harder because somebody in the video demanded it.
- Interview: a host asks, a specialist answers. Works on voice only, zero face seconds.
- Debate: two defensible positions on the same fact, ideal for comparison and controversy topics.
- Anchor and reporter: the anchor frames the story, the reporter delivers the detail.
- Teacher and student: the student asks the exact question the viewer is thinking.
- Question and answer: 8 questions in 12 minutes means 8 resets of attention.
- Narrator plus quoted character: a second voice reads testimonies, quotes and dialogue.
How FalconVid runs a cast without opening a second budget
Multi character with distinct voices is on in every plan. Every creation feature in FalconVid is on in every plan: what changes between plans is volume, how many channels you run, how many generations run at the same time, the Senior Analyst (from Pro, with 7 free days on Starter) and support. The script gets characters, each character gets its own voice from the catalogue, including ultra realistic premium Cartesia voices and voice cloning if you want a specific one, and the video is still billed as a 12 minute video. That means 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium, exactly the same as if a single narrator had read the whole thing.
The identity of each character is held by the channel DNA, so the anchor of video 12 is still the anchor of video 40, with the same face, the same voice and the same visual signature. Face drift between videos is the most common failure of AI casts, because a face regenerated from a prompt each week is a different person each week. Here it is a persistent project setting, which is what allows a cast to accumulate recognition instead of resetting it every publication.
There is one place where a face is cheap, and it is worth knowing before you plan any of this. The face on the thumbnail is charged as a cover, 479 credits in economy and 214 in premium, not per second. The cover has no meter running, so it is the cheapest possible application of an AI avatar in the entire channel. A channel can run a visible recognizable presenter on every single cover and still have zero avatar seconds inside the video.
Karaoke captions with CTAs come with the render, in over 15 styles, on every plan, and they have no credit line of their own. That matters more with a cast than with a single narrator, because dialogue formats are consumed with sound off far more often than monologues, and captions are what keeps a two voice exchange readable when there is no audio. It is also part of why the dialogue upgrade genuinely costs nothing: neither the extra voice nor the captions add a line to the bill.
Around all of that, the operation stays the same as any other video. You approve a calendar once and AI specialists work in parallel on each video, researcher, scriptwriter, narrator, editor and sound design, with a video ready in up to 30 minutes and simultaneous generations running from 2 at once on Starter up to 50 on Scale. Before anything publishes you can open Studio, watch version 1 and adjust it, shorten the intro, swap a scene, change the music, or open the timeline if you want to work frame by frame.
The layout that costs little and reads as expensive production
The design that gets the most out of this asymmetry is straightforward. Keep the cast in the voices, where it is free, and keep the faces to four measured doses: the hook, the chapter transitions, the verdict and the close. Those doses add up to roughly 90 seconds in a 12 minute video, and the remaining 11 minutes or so are carried by generated scenes, b roll, graphics and the narration that the characters are already doing. Face when the video makes a claim or a request, scenes when the video delivers information.
With a cast you distribute those doses instead of stretching them. The host takes the hook and the close, the second character takes a transition and the verdict, or whatever the script actually calls for. Total face seconds on screen stays at 90 and the lip sync line stays at 2,880 credits, because the doses were split, not added. The viewer sees two presenters in the video and the bill sees one presenter's worth of screen time.
The single rule that protects the budget is this: never let two faces occupy the frame at the same moment during those doses. Alternate them. The reaction cut, which is the default answer of every interview program ever made, shows one face at a time and therefore runs one meter at a time. It also reads better, because a full frame face carries more presence than two half faces competing for the same attention.
Run the month and the difference becomes a business decision instead of a preference. On Pro at $97 with 30,000 credits, videos with an alternating cast at 3,888 credits each fit 7 times. The same videos with the two hosts sharing the frame at 6,768 credits fit 4 times. Turned into full talking heads at 24,048 credits, the same budget fits 1. Nobody actually runs a month in a single configuration, so the realistic plan is mixing plain economy videos at 1,008 credits with a handful carrying the 90 second dose, and reserving the heavier modes for the topics that pay for them.
- Hook 8 to 15 seconds, transitions 5 to 8 seconds each, verdict 15 to 25, close 10 to 15.
- Split the doses between characters instead of adding doses: still 90 face seconds, 2,880 credits.
- Same 90 seconds with both faces in frame: 5,760 credits, and the video goes from 3,888 to 6,768.
- Rest of the video: bank scene 5 credits, economy image 63, research plus script 307, base pipeline fee 16.
- Karaoke captions and CTAs: part of the render, no credit line, on every plan.
The same decision, on a manual channel and on an automated one
On a manual channel, a cast means more people in front of a camera. Two presenters means two schedules to align, two fees to pay, two shooting days or one day where both have to be available, two sets of retakes and two people whose absence stops the publication. Add a third character and the coordination problem grows faster than the content does. This is why almost no small channel has a cast: it is not a creative decision, it is a payroll and logistics decision, and the answer is usually no.
On an automated channel the same decision costs almost nothing to make. A character is a script decision. The extra voice is free, because narration is billed per minute of video, so 288 credits of premium narration or 36 of economy narration cover as many characters as the script needs. The face is a dose measured in seconds, 32 credits each at Kling Avatar Standard, and you decide how many of those seconds you buy and which character gets them.
And the same cast runs in as many channels and as many languages as you want. FalconVid writes, narrates, edits, captions and publishes to YouTube, Instagram, TikTok, Rumble and Facebook in up to 63 languages, and duplicating a project into another language costs only the difference, so an anchor and reporter pair built once can front the English channel, the Spanish one and the Portuguese one without being rebuilt. Starter runs 1 channel, Pro runs 5, Business 10, Agency 25 and Scale has no channel limit at all.
So the practical answer splits by stage rather than by taste. A new channel should hold a single recognizable host, use a cast of voices for the retention benefit and spend its face seconds on the hook and the close. An established channel can add a visible second character wherever the format asks for it, keeping the faces alternating rather than sharing the frame. Either way, the constraint that used to settle this argument, whether a second human being was available and affordable, is not the constraint anymore.
