Muted playback is the default state, not the exception
The format gets watched in conditions no editor ever tests in: a phone at arm's length, in a queue or in bed, the volume low or off, and a thumb already resting on the glass. Whatever your first two seconds say out loud, a large share of the audience never hears it. Treat muted playback as the default and sound as the bonus.
Without captions, those two seconds are a blank information screen. Something moves, a voice nobody can hear explains the premise, and nothing on screen says what the viewer gets for staying. The thumb goes up, the early test batch comes back negative, and the loss has nothing to do with the idea you wrote.
Captions are the cheapest fix available, because they are the only element that survives silence, daylight and a five inch screen without asking the viewer to do anything. In short form the caption is not decoration on top of the hook, it is how the hook gets delivered. Size, position and timing deserve the attention you give the script.
The safe zone: where a 9:16 caption is allowed to live
A vertical frame is not a clean canvas. The top loses room to the status bar and the app header, and the bottom loses a much bigger strip to the title, the channel handle, the description line and the column of action buttons. Your visible working area is smaller than the file you export, and the file is what your editor keeps showing you.
Work with roughly 12 to 15 percent of clear margin at the top and 20 to 25 percent at the bottom. On a 1080 by 1920 export that is about 230 to 290 pixels free above and 380 to 480 below. Anchor the block in the lower middle, around 60 to 70 percent down the frame, close to the action and never glued to the bottom edge.
When a caption falls behind the interface the result is not ugly, it is invisible. In the editor the words look perfect; in the feed the title covers the first line and the buttons clip the last word, so the viewer never learns the video was captioned at all. FalconVid renders the 9:16 output with captions already inside that safe area, so the same file survives the YouTube, Instagram and TikTok layouts.
- Top of the frame: leave 12 to 15 percent clear for the header
- Bottom of the frame: leave 20 to 25 percent clear for title and buttons
- On 1080 by 1920 that is about 230 to 290 px above and 380 to 480 below
- Anchor the block around 60 to 70 percent down the frame
- A caption under the interface is invisible, not just untidy

Two to four words per block, never the full sentence
This is the change with the best effort to result ratio in the whole format: show 2 to 4 words at a time in a Short, against the 6 to 9 that work fine in a long video. A short block is caught in one glance, so the eye never leaves the image. A full line forces real reading, and reading pulls attention off the scene.
The arithmetic backs it up. At natural narration speed, roughly 150 to 170 words per minute, a 3 word block lives on screen for about 1.1 seconds, long enough to land and short enough to keep a pulse. A 9 word line sits there for over 3 seconds, which on a 30 second Short is a tenth of the video frozen as a paragraph.
Break blocks by meaning and not by width. The words and then the make a technically valid block and a useless one. Keep the number with its unit and the name with its verb, and give the punchline a block of its own so the last beat lands alone.
Size, weight and outline: readable at arm's length
For size, work around 7 to 9 percent of frame height for the caption line, roughly 135 to 170 pixels on a 1080 by 1920 export. It looks absurd in a desktop preview and correct on a phone. The honest test is to shrink the preview until the video is about the size of a credit card: if you cannot read it there, nobody reads it in the feed.
Weight is not optional. Bold or heavy, never regular, and always an outline or a shadow, because a bright background eats white text alive: a caption over a sky, a beach or a white kitchen simply vanishes for two seconds. A stroke of 2 to 3 pixels scaled with the font, or a shadow at 40 to 60 percent opacity, holds up over any footage.
Cap the whole channel at two font families. One display face for captions and one for on screen labels is plenty, because recognition is built by repetition and not by variety. Channel DNA in FalconVid holds font, size, color and position identical across every video and every Short, so new videos inherit the identity instead of rebuilding it from memory.
- Caption height around 7 to 9 percent of frame height, about 135 to 170 px
- Bold or heavy weight, never regular or light
- Outline of 2 to 3 px or a shadow at 40 to 60 percent, always on
- Credit card test: shrink the preview and try to read it there
- Two font families maximum across the entire channel
Where FalconVid fits: captions that arrive finished
Everything above is a set of decisions, and decisions are exactly what a template should carry for you. FalconVid burns karaoke captions with more than 15 styles, already placed inside the 9:16 safe zone and already synced to the narration it generated, so there is no editor to open and no keyframe to nudge. You choose a style for the channel and every video inherits it.
The same long video becomes 9:16 Shorts automatically, with the same voice and the same caption identity, which is the cheapest way to pull a batch of Shorts out of work you already paid for. The Studio lets you change the caption style after watching the first cut, along with trimming the intro or swapping the music, and the video re renders and goes out on the calendar you approved.
The contrast is worth measuring instead of feeling. Captioning one Short by hand with word level karaoke costs 20 to 40 minutes between transcription, timing correction, styling and safe zone checks. A channel posting 2 Shorts a day spends 20 to 40 hours a month on that single task, typing words it already wrote.
The volume behind that has numbers tied to the mode, not to a slogan. On a 12 minute long video the engine spends 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium. Starter is $47 a month with 15,000 credits, which is 14 videos in pure economy or a realistic 10 to 12 a month mixing economy with one in balanced, and generations run side by side, 2 at a time on Starter and 50 on Scale.
- Karaoke captions in 15+ styles, burned in and synced to the narration
- Already inside the 9:16 safe zone, with no repositioning per cut
- Long video turned into Shorts automatically, same voice and same identity
- Channel DNA locks font, color and position across the whole channel
- Studio: change the style after watching the first cut, then it re renders
- By hand: 20 to 40 minutes per Short, 20 to 40 hours a month at 2 a day
Karaoke highlighting without the fatigue
Word by word highlighting works for one specific reason: it gives the caption a beat. The viewer always knows which word is being said, the rhythm of the narration becomes visible, and part of what the audio would have carried is restored on a muted screen.
It stops working when it turns into a light show. A color pulsing on every single word for 60 straight seconds is exhausting, and the exhaustion reads as cheap. Keep the base color constant and animate one property at a time, either the highlight color or a scale bump of 4 to 6 percent, never three effects on the same word.
Take the highlight color from the channel identity instead of defaulting to yellow. Yellow is the factory setting of every captioning app, so it announces template before the first word is read. Pick a tone from your own palette that still contrasts with your most common footage, and avoid saturated red or green over skin tones and nature shots.
Sync to the syllable, and why auto captions are not a style
Timing is judged more harshly than design. The caption has to change with the syllable and not with the sentence, and a drift of 200 to 300 milliseconds already reads as amateur to a viewer who could never name what is wrong. Whenever the edit allows it, align the block change with the visual cut, so the cut and the new words land as one beat.
This is also why the automatic captions on YouTube do not solve the problem. They arrive outside your style, they sit in the area the interface covers, they carry no channel identity, and the viewer has to turn them on. They are an accessibility feature, and a good one, but accessibility and retention are different jobs.
Doing all of this by hand is possible, and at one Short a week it is even reasonable. It stops being reasonable at 2 Shorts a day, which is 20 to 40 hours a month of transcription and timing before anyone touches the idea. That is a limit of the manual workflow, not of the format.
Generated instead, the same discipline becomes a channel setting: the style is chosen once, the karaoke is synced to a narration the system already has aligned, the 9:16 output lands inside the safe zone, and the videos go out on the calendar you approved. What stays with you is the part worth your time, deciding what the Short says, on one channel or on as many as you want to run.

