The same headphones, two videos, two different volumes
A viewer opens your twelve minute video on a phone, pushes the volume to about seventy percent to hear the narration comfortably, and settles in. Three minutes later the next video starts, from a channel in your niche, and it arrives so much louder that the first reflex is to pull the volume back down. That is not the headphones and it is not an impression. It is a measurable difference in delivery level, decided long before anybody pressed play.
The damage is not the discomfort, it is the comparison. Quiet audio reads as amateur before a single sentence gets judged, and the opening thirty seconds are exactly when a viewer decides whether this channel sounds like the ones that already earned their attention. Nobody thinks the integrated loudness is low here. They think it sounds cheap, or they think nothing at all and click the next thumbnail, which is the version that lands in your retention graph with no label attached to it.
The instinct is to blame the voice. Creators re-record the narration, buy a different text to speech voice, stack an equaliser on the track, and nothing moves, because timbre was never the problem. Two exports of the same narration, from the same voice, can sit six decibels apart, and six decibels is the whole distance between sounding professional and sounding like a bedroom recording. On a faceless channel this hurts twice, because the audio is the entire performance and there is no face on screen to carry the slack.
Four decisions settle the question, and all four happen at export: the loudness target, the peak ceiling above it, the speaker that will play the file back, and the compression that keeps the quiet sentence alive. Set them once and they never change again. Skip them and every upload gambles with a playback level you did not choose.
How YouTube normalizes: loud gets pulled down, quiet never gets pushed up
Every upload is measured. The platform reads the integrated loudness of the whole file, compares it against a house target that sits around -14 LUFS, and applies playback gain to bring loud material closer to it. Anything that arrives meaningfully louder is attenuated on playback. Anything that arrives quieter is left exactly where it was. That asymmetry is the whole article, and it is the part almost nobody explains out loud.
LUFS measures perceived loudness, not peak. A peak meter tells you the tallest single sample in the file, which says nothing about how loud twelve minutes of video feel. Integrated LUFS averages the entire programme with a weighting that follows human hearing, so a dense mix and a sparse one that peak in the same place can sit eight decibels apart in perceived level. Loudness is the number the platform reads. Peak is the number that protects you from distortion. They are not interchangeable and they never were.
Now put the asymmetry to work. Deliver around -14 LUFS and you play in the same field as everyone else in the sidebar. Deliver at -20 and you sit six decibels under that field on every single playback, permanently, because no gain is ever added on your behalf. Ten decibels is roughly the point where the ear calls something half as loud, so six is not a subtlety. The viewer does not diagnose it and does not raise the volume. They leave.
You can read your own verdict in thirty seconds. Right click the player, open Stats for nerds and find the content loudness line. A negative number means the platform pulled your audio down by that much, which is a good sign and means you delivered above the target. Zero or a positive number means nothing was pulled down and nothing was added, so the level you exported is exactly the level a stranger hears, quiet included. Do it on your last three videos and on the competitor you keep losing to.
This is one of the reasons the mix inside FalconVid is not a decision you make per video. The engine writes the script, generates the narration, lays the music and the effects and renders the mix already balanced, so the same standard leaves the render on the first video of a channel and on the three hundredth. Nobody is measuring a file at one in the morning and hoping they remembered the routine from last week.
- Integrated loudness is the whole video averaged, not its loudest moment.
- YouTube attenuates uploads that arrive meaningfully above roughly -14 LUFS.
- Uploads under the target stay quiet, no gain is ever added for you.
- A file delivered at -20 LUFS plays about 6 dB under the field, on every view.
- Stats for nerds, content loudness: zero or positive means you left loudness on the table.

True peak under -1 dBTP: the distortion that only shows up after the upload
Loudness is one number and peak is the other, and the second one is where good narration turns into crackle. Your editor shows the file peaking at -0.1 dBFS and calls it safe. It is not safe, because the file the viewer receives is not the file you exported. The platform re-encodes everything into a lossy format, and the waveform reconstructed on the other side does not land on the same points yours did.
A sample peak meter only reads the samples that exist. Between two samples the real curve can rise higher than either of them, and lossy encoding exaggerates that overshoot. A track sitting at -0.1 dBFS can come out of the decoder above zero, where there is no headroom left, and the result is clipping on exactly the loudest consonants: the letter s, the letter t, an excited word at the end of a sentence. It sounds like a cheap microphone. It was a good microphone and a bad ceiling.
The fix costs nothing at all. Set the limiter ceiling to -1 dBTP with the meter in true peak mode, which is the standard recommendation for lossy delivery, and drop to -1.5 or -2 dBTP when the material is heavily limited or full of hard consonants. You are not giving up loudness, because loudness comes from the integrated level and not from the last decibel of headroom. Switch the meter to true peak as well, otherwise it keeps reporting the reading that hid the problem in the first place.
Run one export with both meters open and it becomes muscle memory: integrated near the target, true peak under the ceiling, no exceptions, no video published without the check. Inside FalconVid there is no export step to remember, because the render produces the finished mix and the publishing goes straight out to YouTube, Instagram, TikTok, Rumble and Facebook without a file ever passing through your hands.
The speaker decides more than the mixer: phones, TVs and the low end that never arrives
You mixed in headphones, at night, in a quiet room. Your audience is on a phone speaker at half volume with a fan running, or on the built in speakers of a television on the other side of the room. Those drivers are physically incapable of moving low frequencies. Below roughly 200 Hz almost nothing survives on a phone, and on the smaller ones the roll off starts higher, somewhere between 300 and 500 Hz.
So the warm bottom end you loved in the headphones is inaudible on the device where most of your audience lives, and it is not free either. That energy still exists in the file, it still counts when the loudness is measured, and it drags the whole mix down when the platform normalises it. You end up spending loudness on frequencies nobody can hear, while the part that carries meaning gets whatever change is left over.
High pass the narration between 80 and 100 Hz. Nothing you need in a human or a synthetic voice lives under 80 Hz, and what does live there is rumble, air conditioning, desk thumps and the low blast of a plosive. Cutting it changes nothing on good speakers, cleans the mix everywhere else, and hands the loudness meter back the room it was wasting on air.
The band that decides whether a sentence is understood is 1 kHz to 4 kHz. That is where consonants live, and consonants are the difference between a word and a smear of sound. A gentle lift of two to three decibels around 3 kHz reads as clearer rather than louder, which is the exact trade you want, because that is also the band a phone speaker reproduces best. Clarity there buys you more than volume does anywhere else.
Music fights you in the same place, and the short answer is to keep the bed well under the voice and let it duck while words are happening. That is a whole subject of its own and it is not this one. Finish the loudness work first, then check the result by listening on a phone at half volume with real background noise in the room, because that is the room your video is actually playing in.
- Phone speakers give up under about 200 Hz, and small ones under 300 to 500 Hz.
- High pass the narration at 80 to 100 Hz, nothing you need lives below it.
- 1 kHz to 4 kHz carries the consonants that decide whether the sentence lands.
- Two to three decibels around 3 kHz reads as clearer, not louder.
- Final check on a phone, half volume, with real noise in the room.
Narration dynamics: 3:1 to 4:1 so the quiet sentence survives the kitchen
Integrated loudness is an average, and an average hides its extremes. A file that measures -14 LUFS overall can still have a thoughtful sentence sitting at -24 and an excited one at -8. In headphones that range reads as expressive. On a phone in a kitchen the quiet sentence falls under the noise floor of the room, and the viewer loses a line, then loses the thread, then loses interest, in that order.
A gentle compressor fixes it. Ratio between 3:1 and 4:1, threshold set so the peaks take three to six decibels of gain reduction while normal speech barely moves, attack of five to ten milliseconds, release of sixty to a hundred and fifty milliseconds. Then make up the gain and normalise to your target. For spoken content a loudness range of roughly five to eight LU is comfortable: still alive, but nothing falls off the bottom of the room.
Synthetic narration needs this as much as a human does, for a different reason. A long video is rendered in segments, and levels drift between them, often by two or three decibels, so a twelve minute narration assembled from dozens of blocks can arrive with audible steps in it. Emphasis, questions and exclamations push peaks in ways that vary from one voice model to another. The compressor is what makes the whole thing feel like one person talking for twelve minutes instead of a playlist of sentences.
Do not overdo it. Eight to one on a voice flattens the emphasis that carries attention, and a narration with no dynamics left is exactly the sound people describe as robotic. You are building a floor, not a wall. Everything should still rise and fall with the meaning, it simply is not allowed to disappear under a fan, a car or a conversation in the next room.
Where FalconVid hands you the mix already balanced
Everything above is per video work, and per video work is the part that never scales. In FalconVid the mix comes out of the render finished: narration, music and effects already balanced against each other, with the track dropping under the voice by itself. The sound design specialist runs in parallel with the researcher, the scriptwriter, the narrator and the editor, and it is not guessing where the speech is, because the same engine generated that narration a moment earlier.
The narration is a line you can price before you start. It costs 24 credits per minute on the premium voices and 3 credits per minute on the economy voice, so a twelve minute video carries 288 or 36 credits of narration. The finished twelve minute video costs 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium, with a base pipeline fee of 16 credits, and a credit is worth $0.003133, so you always know the price of the audio before you approve anything.
Karaoke captions are switched on in every plan, in more than fifteen styles, and they carry no credit line of their own because they are part of the render and not an extra. That matters for this article in particular: a large part of the audience on phones watches with the sound off entirely, so the caption is what holds them in the moments when loudness cannot reach them at all.
When you want to hear it before it goes out, the Studio plays the V1. You can swap the music, shorten the intro or replace a piece of media without regenerating the video, and open the timeline when you want to be precise. If the video carries an ai avatar the mixing rules do not change at all, only the bill does, since lip sync is charged per second of face on screen, starting at 32 credits a second on Kling Avatar Std.
Volume comes from the plan and not from your evening. Starter is $47 with 15,000 credits, 1 channel and 2 simultaneous generations, which realistically means around 10 to 12 videos a month mixing the economy mode with one in balanced, because nobody runs an entire month in a single mode. Pro is $97 with 30,000 credits, 5 channels and 5 simultaneous generations, and every creation feature is on every plan: what changes by plan is volume, channels, simultaneous generations, the Senior Analyst (from Pro; Starter gets 7 free days) and support.
Two ceilings: the manual one and the automated one
Do this by hand and it is a real routine. Open the meter, check the integrated loudness against the target, check the true peak against the ceiling, export, and confirm the file before it goes up. Once you know the moves it costs about ten minutes per video, which is five hours a month on a channel publishing once a day, all of it spent on a step the viewer will never notice when it is done right.
And it is not the skill that breaks first, it is the repetition. Video forty, one in the morning, after a render that failed and a thumbnail you redid twice, you skip the check just this once. That video plays five decibels under the rest of the channel forever, and of course it is the one the algorithm decides to show to strangers. The ceiling of the manual track was never your ear. It is your evening.
The automated track moves that ceiling somewhere else. The mix leaves the render with the same standard on video one and on video three hundred, across every channel you run, in up to 63 languages, from a calendar you approve once. Nobody has to remember the routine, because the routine is the render itself. What limits you stops being your patience and becomes the credits on the plan you chose.
So keep both numbers in view. By hand, a good audio pass costs about ten minutes per video and it is the first thing to go when the week gets heavy. Automatically, the standard is identical on every upload, in as many channels as you want, and the only thing left for you to decide is how much volume you want to buy and which quality mode you want it in.
- Manual: a meter, an export and a check on every single publication.
- Manual: about 10 minutes per video, five hours a month at one a day.
- Manual: the check is the first thing skipped when you are tired.
- Automated: the same mix standard on video 1 and on video 300.
- Automated: identical across every channel and up to 63 languages.

