Blog

Captions on a long video: the silent audience, the search signal, and why YouTube gets your names wrong

Captions are treated as an accessibility checkbox and then ignored. On a long form faceless channel they are three different things at once, and only one of them is accessibility.

Ricardo AlmeidaFounder13 min read
A wide video frame outline in gold with a thick glowing bar across its lower third and a muted speaker shape.

Three separate jobs wearing the same name

The word captions covers three things that behave completely differently, and confusing them is why most channels handle them badly. The first is the closed caption track, the one a viewer toggles on with the C key, which YouTube stores as text alongside the video. The second is burned in text, the styled words drawn onto the frames themselves, which cannot be turned off. The third is the transcript, which is the same text again but used for search and for translation.

The closed caption track is what accessibility law and YouTube's own tooling care about. Burned in text is what actually holds a scrolling viewer who never touches the caption button. The transcript is what a search engine reads. A channel that ships only one of the three is leaving two of the jobs undone, and most channels ship only the automatic version of the first one.

On long form this matters more than on Shorts, and for an unintuitive reason. A Short is watched in a feed with the sound often on and the thumb ready to swipe, so burned text wins by a wide margin. A 12 minute video is frequently watched in a second tab, on a phone in a quiet room, or by someone who speaks your language as a second language, and those three viewers all behave differently around text on screen.

  • Closed captions: a text track the viewer toggles, stored by YouTube.
  • Burned in text: styled words drawn on the frames, always visible.
  • Transcript: the same words, used for search and translation.
  • Most channels ship only automatic closed captions and call it done.

The silent audience is larger than the accessibility audience

The population that benefits from captions on a long video is not primarily deaf or hard of hearing viewers, though those viewers absolutely matter and are the reason the feature exists. The larger group is people watching with the sound off by circumstance: on public transport, in an office, next to a sleeping child, or in a second browser tab while doing something else. Industry surveys of social video consistently put sound off viewing in the double digits as a share of total watch time.

The second large group is non native speakers. A faceless channel published in English is watched by people all over the world whose English is good enough to read but not always good enough to follow a fast narration with an unfamiliar accent. Text on screen turns a video they would have abandoned at 40 seconds into one they finish.

Neither group tells you they exist. They do not comment about it and they do not appear as a segment in Analytics. They only appear as retention, which is the metric that decides whether the algorithm keeps showing your video. This is the honest reason to care about captions on long form: not compliance, retention.

  • Sound off viewing is a double digit share of social video watch time.
  • Non native speakers read faster than they listen to an unfamiliar accent.
  • Neither group shows up in Analytics as a segment, only as retention.
  • Retention is the metric that decides distribution, which makes this a revenue question.

Why the automatic caption gets your niche wrong

YouTube generates automatic captions with speech recognition, and on clean studio audio in a common accent it is genuinely good. Where it breaks is exactly where a specialist channel lives: proper nouns, brand names, acronyms, product model numbers, foreign words, technical vocabulary and figures. A finance channel saying a ticker, a tech channel saying a model number, a history channel saying a place name in another language: those are the words that carry the meaning, and those are the words that come back wrong.

The error also compounds. When speech recognition mishears one word it often mangles the phrase around it, so a single wrong proper noun can turn a sentence into nonsense. The viewer reading captions gets a paragraph that does not parse, decides the channel is sloppy, and leaves. And because the caption track is also the transcript, the wrong words are the ones YouTube has on file as the content of your video.

There is a second, quieter problem specific to AI narration. Speech recognition is trained on human speech, and a synthetic voice reading a number as a string of digits or an acronym letter by letter is transcribed inconsistently. The narration itself may be perfectly clear to a human ear and still produce a messy automatic transcript.

There is a way around the whole problem, and it is to never transcribe in the first place. When the script already exists, the caption can be written from it instead of being guessed from the audio. That is how FalconVid produces the caption layer, which is why a ticker, a model number or a place name arrives in the captions exactly as it was written rather than as speech recognition heard it.

  • Automatic captions fail hardest on names, acronyms, numbers and foreign words.
  • One misheard proper noun frequently corrupts the sentence around it.
  • The caption track doubles as the transcript, so the errors become your record.
  • Synthetic narration transcribes inconsistently even when it sounds clean.

What captions do and do not do for search

YouTube has the full text of your video, either from your uploaded track or from its own transcription, and it uses that text to understand the topic. What it does not do is rank the video on caption text the way a web page ranks on body copy. Nobody has ever won a competitive video keyword by stuffing the transcript, and treating captions as a keyword field is a waste of effort.

Where the transcript genuinely helps is in two narrower places. It clarifies ambiguous topics, so a video whose title could belong to three different niches gets classified correctly. And it powers automatic translation of captions into other languages, which is how viewers outside your language find and tolerate your video at all.

The practical conclusion is unglamorous. Write for the human, get the words right, and let the transcript be an accurate record rather than a marketing surface. The title, the thumbnail and the first 30 seconds still decide whether anyone clicks.

  • YouTube reads the transcript to understand the topic, not to rank on keywords.
  • An accurate transcript helps classify a video whose title is ambiguous.
  • The transcript powers automatic caption translation into other languages.
  • Keyword stuffing a transcript has never won a competitive video term.

Burned in text on a 12 minute video: how much is too much

Word by word karaoke styling is the default on Shorts because a Short is 30 seconds of pure attention capture. On a 12 minute video the same treatment is exhausting: constant motion in the lower third competes with the visual you spent credits generating, and by minute four it reads as noise. The version that works on long form is calmer, showing a full line or two at a time, sitting still long enough to be read, and holding a fixed position so the eye stops hunting for it.

There is also a placement problem specific to long form. YouTube draws its own progress bar and controls across the bottom of the frame when the viewer moves the mouse, and burned text sitting too low disappears behind them. Keeping the text above the lower eighth of the frame solves it, which is a detail nobody notices until they watch their own video on a laptop.

The strongest use of burned text on long form is selective rather than constant: on numbers, on names, on the two or three sentences per video that carry the argument, and on the call to action. That gives the sound off viewer anchors to follow without turning the whole video into a wall of moving words.

Consistency is the part that decays fastest when this is done by hand, because the style gets chosen again on every upload. FalconVid generates the caption layer with the video in more than 15 styles and keeps it tied to the channel identity, so the font, the position and the behaviour stay identical across a hundred videos without anyone deciding them again. It is on every plan from Starter at $47 upward, so it is not a setting you upgrade into.

  • Word by word karaoke suits Shorts, not a 12 minute watch.
  • Show a full line, hold it still, keep the position fixed.
  • Stay above the lower eighth of the frame or the player controls cover it.
  • Burn selectively: numbers, names, the argument and the call to action.
A golden audio waveform with a row of glowing rectangles beneath it, aligned to the peaks by faint guide lines.

How FalconVid handles this without adding a step to your week

The reason most channels never fix their captions is that fixing them is a per video chore in a workflow that already has too many. FalconVid produces the caption layer as part of the pipeline rather than as a task afterwards: the karaoke captions and the CTAs are generated with the video, in more than 15 styles, and they are on every plan from Starter at $47 upward rather than locked behind a higher tier. There is no separate credit line for captions, because they are part of the render.

It also sidesteps the transcription problem entirely. The system already holds the script it wrote and narrated, so the caption text is the script rather than a machine's guess at what the narration said. Names, acronyms and figures are correct because they were never transcribed in the first place, which is precisely the failure mode described three sections above.

For an idea of what the surrounding budget looks like, narration on a 12 minute video is 288 credits with premium voices and 36 in economy mode, against 1,008 credits for the entire video in economy mode and 26,760 in premium. Captions add nothing to that, and video SEO, the title, description and tags are written by the same pipeline, in the video's own language.

And when the same video goes to another language, duplicating a finished project repays only the difference, which is the narration line, not the research, the script or the scenes. The captions in the new language come from the translated script, so the second language starts with the same accuracy as the first instead of inheriting a machine transcription of a machine voice.

  • Karaoke captions and CTAs are generated with the video, 15 or more styles.
  • Included on every plan from Starter at $47, with no separate credit line.
  • The caption text is the script, so names and numbers are never transcribed.
  • Narration on a 12 minute video: 288 credits premium, 36 in economy mode.

A 10 minute checklist that covers the whole thing

Start by watching two minutes of your own most recent video on a phone with the sound off. Not skimming it, watching it. If you cannot follow the argument, no amount of caption theory matters, because that is what a real share of your audience is experiencing right now.

Then read the automatic transcript YouTube generated for your last upload and look specifically at the proper nouns, the acronyms and the numbers. If they are wrong, upload a corrected track, because that text is both what accessibility users read and what YouTube has on file about your topic.

Finally, decide a fixed style and stop revisiting it. One font, one position above the lower eighth of the frame, one behaviour for burned text, applied to every video. Consistency in the caption layer is part of channel identity in exactly the way the thumbnail is, and a channel that changes it every upload looks assembled rather than produced.

  • Watch two minutes of your last video on a phone with the sound off.
  • Read the automatic transcript and check names, acronyms and figures.
  • Upload a corrected track when the transcription is wrong.
  • Fix one style and one position, then never change them per video.

FAQ

Got questions? We've got answers.

Do captions actually improve YouTube ranking?

Not directly. YouTube reads the transcript to understand what your video is about, which helps it classify an ambiguous topic and powers automatic caption translation, but nobody wins a competitive keyword by stuffing a transcript. The real gain is retention, because sound off viewers and non native speakers finish videos they would otherwise abandon, and retention is what drives distribution.

Are YouTube automatic captions good enough to just leave alone?

On clean audio in a common accent they are decent for ordinary speech, and bad exactly where it counts: proper nouns, brand names, acronyms, model numbers, foreign words and figures. Those are the words that carry your meaning, and a single wrong one often corrupts the sentence around it. On a specialist channel, leaving them alone means publishing a record of your video that gets your own vocabulary wrong.

Should I use word by word karaoke captions on a 12 minute video?

Usually not for the whole thing. Karaoke styling was designed for a 30 second Short where the job is to hold a thumb. Across 12 minutes the constant motion competes with your visuals and reads as noise by minute four. Show a line or two at a time, hold it still, and reserve heavy styling for numbers, names, the core argument and the call to action.

Where on the frame should burned in text sit?

Above the lower eighth. YouTube draws its progress bar and controls across the bottom of the frame whenever the viewer moves the mouse or taps, and text placed too low disappears behind them at exactly the moment someone is interacting with your video. Keep the position identical across every upload so the viewer's eye stops searching for it.

Do I have to caption every video by hand, or pay for it separately?

Not with FalconVid. The karaoke captions and CTAs are produced inside the pipeline along with the video, in more than 15 styles, and they are included on every plan from Starter at $47 upward with no separate credit line. Captions stop being a per video chore, which is the actual reason most channels never fix theirs.

How do captions work when I publish the same video in another language?

The safe way is to caption from the translated script rather than transcribing the translated narration, because transcription errors compound in a language you may not speak well enough to proofread. In FalconVid, duplicating a finished project into another language repays only the difference, which is the narration line and not the research, the script or the scenes, and the captions come from that translated script.

Does adding captions cost me anything in credits?

In FalconVid, no. Karaoke captions and CTAs are part of the render rather than a priced piece, and they are on every plan rather than gated by tier. For scale, narration on a 12 minute video is 288 credits with premium voices and 36 in economy mode, against 1,008 credits for the whole video in economy mode, 8,676 in balanced and 26,760 in premium.

My narration is AI. Does that make automatic transcription worse?

Often yes, in a way that surprises people. Speech recognition is trained on human speech, so a synthetic voice reading a string of digits or spelling out an acronym transcribes inconsistently even when it sounds perfectly clear to your ear. The reliable fix is to caption from the script you already have rather than asking a machine to listen to another machine.

Let the captions come out of the script, not out of a guess

FalconVid researches, writes, narrates, edits, captions and publishes to YouTube, Instagram, TikTok, Rumble and Facebook in up to 63 languages, from a calendar you approve once, with AI specialists working in parallel and a video ready in up to 30 minutes. Karaoke captions in over 15 styles, CTAs, 9:16 Shorts and video SEO are on every plan, not gated by tier. A 12 minute video costs 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium, and you pick the mode per video. Starter at $47 with 15,000 credits, 1 channel and 2 simultaneous generations is around 10 to 12 videos a month mixing economy with one in balanced, or 14 in pure economy, which is about 3 hours of video. Pro is $97 with 30,000 credits, 5 channels and 5 simultaneous generations, 29 economy videos or 6 hours, up to Scale at $997 with 320,000 credits, 50 channels and 50 simultaneous generations. Every creation feature on every plan, 7 day trial with 2,000 credits and a 7 day guarantee.

Create my channel now

Charged today · 7-day guarantee · Cancel anytime

Keep reading