Why the voice gets the name wrong: it reads spelling, not meaning
The model never sees a name. It sees a string of characters and picks a pronunciation from the spelling plus the rules of the language it was told to speak. It does not know that FBI is spelled out, that Nguyen is Vietnamese, or that 1.2M is money. Everything arrives as text and gets read by one set of rules.
That is why the mistakes look arbitrary. The same voice carries a 12 minute script beautifully and then destroys the one word your video is about. Viewers forgive a voice they know is synthetic. They do not forgive a channel that mispronounces the company it spent nine minutes analysing, because that reads as a creator who did not do the homework.
Put a number on it. A finance or history channel publishing three times a week ships around 36 videos a quarter. Four mangled names or figures per video is 144 moments where an informed viewer quietly downgrades you. One audible slip every 90 seconds moves a comment section from questions about the topic to jokes about the narrator.
And in a manual workflow the repair is brutally out of proportion to the mistake. One wrong word means a new take, matching it to the old timing, re-syncing cuts built around the old audio and rendering again: one to three hours to fix two seconds. Most creators do that maths, publish the error and hope nobody notices.
The 6 families of pronunciation error, with a written example of each
Almost every mispronunciation belongs to one of six families. Naming them matters because each has its own fix, and because once you recognise the pattern you catch the error while writing instead of while watching the finished cut.
Read the list with your own script open. A 12 minute video usually holds between eight and twenty candidates, and roughly a quarter of them actually come out wrong. Marking them takes about four minutes.
- Acronyms. NASA is a word, FBI is three letters, SaaS is neither until you decide. Write F B I with spaces or F.B.I. with periods to force the letters, and write sass to force the word.
- Foreign proper names inside an English sentence: Nguyen, Huawei, Bouygues, Xiaomi, Nietzsche. The voice applies English rules to a name that does not obey them and produces something your audience has never heard.
- Large numbers, money and percentages. 1.2M comes out as one point two em and 1,250 as one comma two five zero. Write one point two million, four point seven billion dollars, twelve fifty.
- Dates and times. 3/4/2026 is ambiguous even to a human, 1980s can become one thousand nine hundred eighty s, 9:30 can become nine colon thirty. Write March fourth twenty twenty six, the nineteen eighties, nine thirty.
- Units and abbreviations. kg, km/h, St. and Dr. all depend on context, and St. Louis and 5th St. are the same two letters read two different ways. Write kilograms, kilometres per hour, Saint Louis, Fifth Street.
- Homographs that change with meaning. The lead engineer will lead the tour. I read it yesterday. The sentence decides, and the model does not always parse it the way you wrote it in your head.
The golden rule: the fix goes in the script, not in the voice
When a name comes out wrong the instinct is to change voices. It almost never works. The new voice reads the same string with the same rules and gives you a different flavour of the same error, and now your channel has also lost the vocal identity it was building.
Write for the ear instead. A narration script is not a document, it is a set of instructions for a mouth. Numbers written the way you would say them, currency in words, acronyms broken with spaces when they must be spelled and written as they sound when they must be spoken. If a stranger reading it aloud would produce the right sound, so will the model.
Then comes phonetic respelling, the trick that saves the hard names. Take the word that fails and rewrite it in ordinary letters that produce the correct sound in the language being narrated: Huawei becomes wah way, Nguyen becomes nwin, Bouygues becomes bweeg. You are not writing IPA, you are writing what a native reader would say seeing it for the first time.
One honest caveat. Captions are generated from the narrated text, so a respelling that looks broken on screen is a real trade. Prefer respellings that still read as the name, or keep the correct spelling and fix that specific stretch afterwards.
This is also where a script tool plus a separate voice tool loses. You write in one place for a reader, paste into another for a speaker, and the seam between them is where every name breaks. In FalconVid the scriptwriter and the narrator are two AI specialists inside the same pipeline working on the same video at the same time, so the text arrives already shaped for the voice that will speak it. There is no handoff to lose the name in.

How FalconVid keeps names, numbers and dates on the rails
Pronunciation stays broken on most channels not out of ignorance but out of repair cost. When a fix takes one to three hours you stop fixing, and by video forty the channel has a permanent accent problem nobody planned.
In FalconVid the pipeline runs in parallel: researcher, scriptwriter, narrator, editor and sound designer work on the same video at once, so a long video returns as a first cut in up to 30 minutes. You approve a content calendar and the videos arrive on their own. When a word still lands wrong you open the Studio, adjust that stretch and remount, or shorten the intro, swap a piece of media, change the music. The Studio is on every plan, like every other creation feature: a bigger plan buys volume, channels, simultaneous generations, the AI senior analyst from Pro, with a 7 day trial on Starter, and support.
The voice is not where you compromise. Narration runs in 63 languages with ultra realistic premium voices, and you can clone your own voice so the channel keeps a single identity. Fixing a word never means abandoning a voice your audience already recognises.
The bill is credits and the quality mode decides it. A 12 minute video costs 1,008 credits in economy, 8,676 in balanced and 26,760 in premium. Starter is $47 a month with 15,000 credits: 14 videos if the whole month stays in economy, or the version people actually run, around 10 to 12 a month with the one that matters in balanced. Pro is $97 for 30,000 credits and 5 channels, and Scale is $997 with 50 channels and 50 concurrent generations.
- Scriptwriter and narrator in the same pipeline, so the text is written to be spoken
- 63 narration languages with ultra realistic premium voices, plus voice cloning
- Studio on every plan: adjust the stretch and remount instead of rebuilding the video
- Parallel production: first cut of a long video in up to 30 minutes
- 12 minute video: 1,008 credits in economy, 8,676 in balanced, 26,760 in premium
- Starter $47 with 15,000 credits: around 10 to 12 videos a month mixing modes, or 14 in pure economy
The 30 second test that saves the 12 minute render
Before committing 12 minutes of narration, generate only the paragraph carrying the risky words and listen to it. Thirty seconds of audio is a small fraction of a video that costs 1,008 credits in economy mode, and it catches the error while being wrong is still cheap.
Build the test paragraph properly. Put every risky term inside it, in the sentence shapes you will actually use, because context is what decides a homograph. A test that reads names as a shopping list proves nothing. Six sentences shaped like your real script prove everything.
Listen on a phone speaker, not on the headphones you edit with. Your audience watches on a phone at half volume in a noisy room, and half the pronunciation problems that survive good headphones become obvious once the mids get crushed.
In a manual pipeline nobody runs this test, for an entirely rational reason: if failing means re-recording, re-syncing and re-rendering, you would rather not know. When a first cut comes back in up to 30 minutes and several videos are produced at the same time, 2 concurrent generations on Starter and 50 on Scale, testing stops being a sacrifice.
Every language trips on a different word
With one channel this is a detail. With the same format running in three languages it is the difference between a second market and a second embarrassment, because each language fails somewhere different and the fix does not travel.
Portuguese stumbles on English acronyms and on currency: CEO, KPI and SaaS get read with Portuguese letter names, and R$ 1.250,00 comes out as a row of digits. It also hides homographs behind a missing accent, so sabia, sábia and sabiá are three words inside one typo.
Spanish trips on English names and anglicisms: Nike read as written, Huawei spelled letter by letter, marketing with full Spanish vowels. Its cruellest trap is the missing accent, because publico, publicó and público are three different words. English, in turn, flattens Latin and German names: João, Beauchamp, Nietzsche, Porsche.
So when you take a proven video into a new language, proper names need new treatment, not translation: a respelling that works in Portuguese produces a different sound in English. In FalconVid you duplicate the project into another language paying only the difference, and it returns narrated by a voice for that language, on its own channel with its own calendar and identity. The only work left is a short list of names.
The 8 item checklist to run before you publish
All of it collapses into a pass that takes about five minutes on a 12 minute script. Run it once and you will spot your own pattern: most creators break two or three families constantly and never touch the other three.
Done by hand this list is pure discipline, and discipline is the first thing to vanish in a bad week. In an automated pipeline most of it happens before the video ever reaches you, which is why automated channels still sound consistent at video 200 while manual ones hold until roughly video 20.
- Scan for every acronym and decide, one by one, spell it out or say it as a word
- Rewrite every number, currency amount and percentage the way you would say it
- Rewrite every date and time in spoken form, never in digits and slashes
- Expand units and abbreviations: kg, km/h, m2, St., Dr.
- Respell foreign proper names phonetically for the language being narrated
- Check homographs by reading the whole sentence, never the isolated word
- Generate the 30 second test paragraph and listen to it on a phone speaker
- When duplicating into another language, rebuild the name list instead of translating it

