Three separate jobs wearing the same name
The word captions covers three things that behave completely differently, and confusing them is why most channels handle them badly. The first is the closed caption track, the one a viewer toggles on with the C key, which YouTube stores as text alongside the video. The second is burned in text, the styled words drawn onto the frames themselves, which cannot be turned off. The third is the transcript, which is the same text again but used for search and for translation.
The closed caption track is what accessibility law and YouTube's own tooling care about. Burned in text is what actually holds a scrolling viewer who never touches the caption button. The transcript is what a search engine reads. A channel that ships only one of the three is leaving two of the jobs undone, and most channels ship only the automatic version of the first one.
On long form this matters more than on Shorts, and for an unintuitive reason. A Short is watched in a feed with the sound often on and the thumb ready to swipe, so burned text wins by a wide margin. A 12 minute video is frequently watched in a second tab, on a phone in a quiet room, or by someone who speaks your language as a second language, and those three viewers all behave differently around text on screen.
- Closed captions: a text track the viewer toggles, stored by YouTube.
- Burned in text: styled words drawn on the frames, always visible.
- Transcript: the same words, used for search and translation.
- Most channels ship only automatic closed captions and call it done.
The silent audience is larger than the accessibility audience
The population that benefits from captions on a long video is not primarily deaf or hard of hearing viewers, though those viewers absolutely matter and are the reason the feature exists. The larger group is people watching with the sound off by circumstance: on public transport, in an office, next to a sleeping child, or in a second browser tab while doing something else. Industry surveys of social video consistently put sound off viewing in the double digits as a share of total watch time.
The second large group is non native speakers. A faceless channel published in English is watched by people all over the world whose English is good enough to read but not always good enough to follow a fast narration with an unfamiliar accent. Text on screen turns a video they would have abandoned at 40 seconds into one they finish.
Neither group tells you they exist. They do not comment about it and they do not appear as a segment in Analytics. They only appear as retention, which is the metric that decides whether the algorithm keeps showing your video. This is the honest reason to care about captions on long form: not compliance, retention.
- Sound off viewing is a double digit share of social video watch time.
- Non native speakers read faster than they listen to an unfamiliar accent.
- Neither group shows up in Analytics as a segment, only as retention.
- Retention is the metric that decides distribution, which makes this a revenue question.
Why the automatic caption gets your niche wrong
YouTube generates automatic captions with speech recognition, and on clean studio audio in a common accent it is genuinely good. Where it breaks is exactly where a specialist channel lives: proper nouns, brand names, acronyms, product model numbers, foreign words, technical vocabulary and figures. A finance channel saying a ticker, a tech channel saying a model number, a history channel saying a place name in another language: those are the words that carry the meaning, and those are the words that come back wrong.
The error also compounds. When speech recognition mishears one word it often mangles the phrase around it, so a single wrong proper noun can turn a sentence into nonsense. The viewer reading captions gets a paragraph that does not parse, decides the channel is sloppy, and leaves. And because the caption track is also the transcript, the wrong words are the ones YouTube has on file as the content of your video.
There is a second, quieter problem specific to AI narration. Speech recognition is trained on human speech, and a synthetic voice reading a number as a string of digits or an acronym letter by letter is transcribed inconsistently. The narration itself may be perfectly clear to a human ear and still produce a messy automatic transcript.
There is a way around the whole problem, and it is to never transcribe in the first place. When the script already exists, the caption can be written from it instead of being guessed from the audio. That is how FalconVid produces the caption layer, which is why a ticker, a model number or a place name arrives in the captions exactly as it was written rather than as speech recognition heard it.
- Automatic captions fail hardest on names, acronyms, numbers and foreign words.
- One misheard proper noun frequently corrupts the sentence around it.
- The caption track doubles as the transcript, so the errors become your record.
- Synthetic narration transcribes inconsistently even when it sounds clean.
What captions do and do not do for search
YouTube has the full text of your video, either from your uploaded track or from its own transcription, and it uses that text to understand the topic. What it does not do is rank the video on caption text the way a web page ranks on body copy. Nobody has ever won a competitive video keyword by stuffing the transcript, and treating captions as a keyword field is a waste of effort.
Where the transcript genuinely helps is in two narrower places. It clarifies ambiguous topics, so a video whose title could belong to three different niches gets classified correctly. And it powers automatic translation of captions into other languages, which is how viewers outside your language find and tolerate your video at all.
The practical conclusion is unglamorous. Write for the human, get the words right, and let the transcript be an accurate record rather than a marketing surface. The title, the thumbnail and the first 30 seconds still decide whether anyone clicks.
- YouTube reads the transcript to understand the topic, not to rank on keywords.
- An accurate transcript helps classify a video whose title is ambiguous.
- The transcript powers automatic caption translation into other languages.
- Keyword stuffing a transcript has never won a competitive video term.
Burned in text on a 12 minute video: how much is too much
Word by word karaoke styling is the default on Shorts because a Short is 30 seconds of pure attention capture. On a 12 minute video the same treatment is exhausting: constant motion in the lower third competes with the visual you spent credits generating, and by minute four it reads as noise. The version that works on long form is calmer, showing a full line or two at a time, sitting still long enough to be read, and holding a fixed position so the eye stops hunting for it.
There is also a placement problem specific to long form. YouTube draws its own progress bar and controls across the bottom of the frame when the viewer moves the mouse, and burned text sitting too low disappears behind them. Keeping the text above the lower eighth of the frame solves it, which is a detail nobody notices until they watch their own video on a laptop.
The strongest use of burned text on long form is selective rather than constant: on numbers, on names, on the two or three sentences per video that carry the argument, and on the call to action. That gives the sound off viewer anchors to follow without turning the whole video into a wall of moving words.
Consistency is the part that decays fastest when this is done by hand, because the style gets chosen again on every upload. FalconVid generates the caption layer with the video in more than 15 styles and keeps it tied to the channel identity, so the font, the position and the behaviour stay identical across a hundred videos without anyone deciding them again. It is on every plan from Starter at $47 upward, so it is not a setting you upgrade into.
- Word by word karaoke suits Shorts, not a 12 minute watch.
- Show a full line, hold it still, keep the position fixed.
- Stay above the lower eighth of the frame or the player controls cover it.
- Burn selectively: numbers, names, the argument and the call to action.

How FalconVid handles this without adding a step to your week
The reason most channels never fix their captions is that fixing them is a per video chore in a workflow that already has too many. FalconVid produces the caption layer as part of the pipeline rather than as a task afterwards: the karaoke captions and the CTAs are generated with the video, in more than 15 styles, and they are on every plan from Starter at $47 upward rather than locked behind a higher tier. There is no separate credit line for captions, because they are part of the render.
It also sidesteps the transcription problem entirely. The system already holds the script it wrote and narrated, so the caption text is the script rather than a machine's guess at what the narration said. Names, acronyms and figures are correct because they were never transcribed in the first place, which is precisely the failure mode described three sections above.
For an idea of what the surrounding budget looks like, narration on a 12 minute video is 288 credits with premium voices and 36 in economy mode, against 1,008 credits for the entire video in economy mode and 26,760 in premium. Captions add nothing to that, and video SEO, the title, description and tags are written by the same pipeline, in the video's own language.
And when the same video goes to another language, duplicating a finished project repays only the difference, which is the narration line, not the research, the script or the scenes. The captions in the new language come from the translated script, so the second language starts with the same accuracy as the first instead of inheriting a machine transcription of a machine voice.
- Karaoke captions and CTAs are generated with the video, 15 or more styles.
- Included on every plan from Starter at $47, with no separate credit line.
- The caption text is the script, so names and numbers are never transcribed.
- Narration on a 12 minute video: 288 credits premium, 36 in economy mode.
A 10 minute checklist that covers the whole thing
Start by watching two minutes of your own most recent video on a phone with the sound off. Not skimming it, watching it. If you cannot follow the argument, no amount of caption theory matters, because that is what a real share of your audience is experiencing right now.
Then read the automatic transcript YouTube generated for your last upload and look specifically at the proper nouns, the acronyms and the numbers. If they are wrong, upload a corrected track, because that text is both what accessibility users read and what YouTube has on file about your topic.
Finally, decide a fixed style and stop revisiting it. One font, one position above the lower eighth of the frame, one behaviour for burned text, applied to every video. Consistency in the caption layer is part of channel identity in exactly the way the thumbnail is, and a channel that changes it every upload looks assembled rather than produced.
- Watch two minutes of your last video on a phone with the sound off.
- Read the automatic transcript and check names, acronyms and figures.
- Upload a corrected track when the transcription is wrong.
- Fix one style and one position, then never change them per video.

