Blog

Your video sounds quieter than the competitor: how YouTube normalizes loudness, and why quiet audio never gets turned up

The platform pulls loud uploads down and leaves quiet ones exactly where they are. Deliver at the wrong level and your channel plays under everybody else on every single view, and the viewer solves it by clicking the next video instead of touching the volume.

Ricardo AlmeidaFounder17 min read
A glowing golden loudness meter arc with one needle pinned high on the scale and a second needle resting far below it.

The same headphones, two videos, two different volumes

A viewer opens your twelve minute video on a phone, pushes the volume to about seventy percent to hear the narration comfortably, and settles in. Three minutes later the next video starts, from a channel in your niche, and it arrives so much louder that the first reflex is to pull the volume back down. That is not the headphones and it is not an impression. It is a measurable difference in delivery level, decided long before anybody pressed play.

The damage is not the discomfort, it is the comparison. Quiet audio reads as amateur before a single sentence gets judged, and the opening thirty seconds are exactly when a viewer decides whether this channel sounds like the ones that already earned their attention. Nobody thinks the integrated loudness is low here. They think it sounds cheap, or they think nothing at all and click the next thumbnail, which is the version that lands in your retention graph with no label attached to it.

The instinct is to blame the voice. Creators re-record the narration, buy a different text to speech voice, stack an equaliser on the track, and nothing moves, because timbre was never the problem. Two exports of the same narration, from the same voice, can sit six decibels apart, and six decibels is the whole distance between sounding professional and sounding like a bedroom recording. On a faceless channel this hurts twice, because the audio is the entire performance and there is no face on screen to carry the slack.

Four decisions settle the question, and all four happen at export: the loudness target, the peak ceiling above it, the speaker that will play the file back, and the compression that keeps the quiet sentence alive. Set them once and they never change again. Skip them and every upload gambles with a playback level you did not choose.

How YouTube normalizes: loud gets pulled down, quiet never gets pushed up

Every upload is measured. The platform reads the integrated loudness of the whole file, compares it against a house target that sits around -14 LUFS, and applies playback gain to bring loud material closer to it. Anything that arrives meaningfully louder is attenuated on playback. Anything that arrives quieter is left exactly where it was. That asymmetry is the whole article, and it is the part almost nobody explains out loud.

LUFS measures perceived loudness, not peak. A peak meter tells you the tallest single sample in the file, which says nothing about how loud twelve minutes of video feel. Integrated LUFS averages the entire programme with a weighting that follows human hearing, so a dense mix and a sparse one that peak in the same place can sit eight decibels apart in perceived level. Loudness is the number the platform reads. Peak is the number that protects you from distortion. They are not interchangeable and they never were.

Now put the asymmetry to work. Deliver around -14 LUFS and you play in the same field as everyone else in the sidebar. Deliver at -20 and you sit six decibels under that field on every single playback, permanently, because no gain is ever added on your behalf. Ten decibels is roughly the point where the ear calls something half as loud, so six is not a subtlety. The viewer does not diagnose it and does not raise the volume. They leave.

You can read your own verdict in thirty seconds. Right click the player, open Stats for nerds and find the content loudness line. A negative number means the platform pulled your audio down by that much, which is a good sign and means you delivered above the target. Zero or a positive number means nothing was pulled down and nothing was added, so the level you exported is exactly the level a stranger hears, quiet included. Do it on your last three videos and on the competitor you keep losing to.

This is one of the reasons the mix inside FalconVid is not a decision you make per video. The engine writes the script, generates the narration, lays the music and the effects and renders the mix already balanced, so the same standard leaves the render on the first video of a channel and on the three hundredth. Nobody is measuring a file at one in the morning and hoping they remembered the routine from last week.

  • Integrated loudness is the whole video averaged, not its loudest moment.
  • YouTube attenuates uploads that arrive meaningfully above roughly -14 LUFS.
  • Uploads under the target stay quiet, no gain is ever added for you.
  • A file delivered at -20 LUFS plays about 6 dB under the field, on every view.
  • Stats for nerds, content loudness: zero or positive means you left loudness on the table.
Two golden sound waves of different heights meeting one horizontal ceiling line, the taller one squashed flat against it and the shorter one far below.

True peak under -1 dBTP: the distortion that only shows up after the upload

Loudness is one number and peak is the other, and the second one is where good narration turns into crackle. Your editor shows the file peaking at -0.1 dBFS and calls it safe. It is not safe, because the file the viewer receives is not the file you exported. The platform re-encodes everything into a lossy format, and the waveform reconstructed on the other side does not land on the same points yours did.

A sample peak meter only reads the samples that exist. Between two samples the real curve can rise higher than either of them, and lossy encoding exaggerates that overshoot. A track sitting at -0.1 dBFS can come out of the decoder above zero, where there is no headroom left, and the result is clipping on exactly the loudest consonants: the letter s, the letter t, an excited word at the end of a sentence. It sounds like a cheap microphone. It was a good microphone and a bad ceiling.

The fix costs nothing at all. Set the limiter ceiling to -1 dBTP with the meter in true peak mode, which is the standard recommendation for lossy delivery, and drop to -1.5 or -2 dBTP when the material is heavily limited or full of hard consonants. You are not giving up loudness, because loudness comes from the integrated level and not from the last decibel of headroom. Switch the meter to true peak as well, otherwise it keeps reporting the reading that hid the problem in the first place.

Run one export with both meters open and it becomes muscle memory: integrated near the target, true peak under the ceiling, no exceptions, no video published without the check. Inside FalconVid there is no export step to remember, because the render produces the finished mix and the publishing goes straight out to YouTube, Instagram, TikTok, Rumble and Facebook without a file ever passing through your hands.

The speaker decides more than the mixer: phones, TVs and the low end that never arrives

You mixed in headphones, at night, in a quiet room. Your audience is on a phone speaker at half volume with a fan running, or on the built in speakers of a television on the other side of the room. Those drivers are physically incapable of moving low frequencies. Below roughly 200 Hz almost nothing survives on a phone, and on the smaller ones the roll off starts higher, somewhere between 300 and 500 Hz.

So the warm bottom end you loved in the headphones is inaudible on the device where most of your audience lives, and it is not free either. That energy still exists in the file, it still counts when the loudness is measured, and it drags the whole mix down when the platform normalises it. You end up spending loudness on frequencies nobody can hear, while the part that carries meaning gets whatever change is left over.

High pass the narration between 80 and 100 Hz. Nothing you need in a human or a synthetic voice lives under 80 Hz, and what does live there is rumble, air conditioning, desk thumps and the low blast of a plosive. Cutting it changes nothing on good speakers, cleans the mix everywhere else, and hands the loudness meter back the room it was wasting on air.

The band that decides whether a sentence is understood is 1 kHz to 4 kHz. That is where consonants live, and consonants are the difference between a word and a smear of sound. A gentle lift of two to three decibels around 3 kHz reads as clearer rather than louder, which is the exact trade you want, because that is also the band a phone speaker reproduces best. Clarity there buys you more than volume does anywhere else.

Music fights you in the same place, and the short answer is to keep the bed well under the voice and let it duck while words are happening. That is a whole subject of its own and it is not this one. Finish the loudness work first, then check the result by listening on a phone at half volume with real background noise in the room, because that is the room your video is actually playing in.

  • Phone speakers give up under about 200 Hz, and small ones under 300 to 500 Hz.
  • High pass the narration at 80 to 100 Hz, nothing you need lives below it.
  • 1 kHz to 4 kHz carries the consonants that decide whether the sentence lands.
  • Two to three decibels around 3 kHz reads as clearer, not louder.
  • Final check on a phone, half volume, with real noise in the room.

Narration dynamics: 3:1 to 4:1 so the quiet sentence survives the kitchen

Integrated loudness is an average, and an average hides its extremes. A file that measures -14 LUFS overall can still have a thoughtful sentence sitting at -24 and an excited one at -8. In headphones that range reads as expressive. On a phone in a kitchen the quiet sentence falls under the noise floor of the room, and the viewer loses a line, then loses the thread, then loses interest, in that order.

A gentle compressor fixes it. Ratio between 3:1 and 4:1, threshold set so the peaks take three to six decibels of gain reduction while normal speech barely moves, attack of five to ten milliseconds, release of sixty to a hundred and fifty milliseconds. Then make up the gain and normalise to your target. For spoken content a loudness range of roughly five to eight LU is comfortable: still alive, but nothing falls off the bottom of the room.

Synthetic narration needs this as much as a human does, for a different reason. A long video is rendered in segments, and levels drift between them, often by two or three decibels, so a twelve minute narration assembled from dozens of blocks can arrive with audible steps in it. Emphasis, questions and exclamations push peaks in ways that vary from one voice model to another. The compressor is what makes the whole thing feel like one person talking for twelve minutes instead of a playlist of sentences.

Do not overdo it. Eight to one on a voice flattens the emphasis that carries attention, and a narration with no dynamics left is exactly the sound people describe as robotic. You are building a floor, not a wall. Everything should still rise and fall with the meaning, it simply is not allowed to disappear under a fan, a car or a conversation in the next room.

Where FalconVid hands you the mix already balanced

Everything above is per video work, and per video work is the part that never scales. In FalconVid the mix comes out of the render finished: narration, music and effects already balanced against each other, with the track dropping under the voice by itself. The sound design specialist runs in parallel with the researcher, the scriptwriter, the narrator and the editor, and it is not guessing where the speech is, because the same engine generated that narration a moment earlier.

The narration is a line you can price before you start. It costs 24 credits per minute on the premium voices and 3 credits per minute on the economy voice, so a twelve minute video carries 288 or 36 credits of narration. The finished twelve minute video costs 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium, with a base pipeline fee of 16 credits, and a credit is worth $0.003133, so you always know the price of the audio before you approve anything.

Karaoke captions are switched on in every plan, in more than fifteen styles, and they carry no credit line of their own because they are part of the render and not an extra. That matters for this article in particular: a large part of the audience on phones watches with the sound off entirely, so the caption is what holds them in the moments when loudness cannot reach them at all.

When you want to hear it before it goes out, the Studio plays the V1. You can swap the music, shorten the intro or replace a piece of media without regenerating the video, and open the timeline when you want to be precise. If the video carries an ai avatar the mixing rules do not change at all, only the bill does, since lip sync is charged per second of face on screen, starting at 32 credits a second on Kling Avatar Std.

Volume comes from the plan and not from your evening. Starter is $47 with 15,000 credits, 1 channel and 2 simultaneous generations, which realistically means around 10 to 12 videos a month mixing the economy mode with one in balanced, because nobody runs an entire month in a single mode. Pro is $97 with 30,000 credits, 5 channels and 5 simultaneous generations, and every creation feature is on every plan: what changes by plan is volume, channels, simultaneous generations, the Senior Analyst (from Pro; Starter gets 7 free days) and support.

Two ceilings: the manual one and the automated one

Do this by hand and it is a real routine. Open the meter, check the integrated loudness against the target, check the true peak against the ceiling, export, and confirm the file before it goes up. Once you know the moves it costs about ten minutes per video, which is five hours a month on a channel publishing once a day, all of it spent on a step the viewer will never notice when it is done right.

And it is not the skill that breaks first, it is the repetition. Video forty, one in the morning, after a render that failed and a thumbnail you redid twice, you skip the check just this once. That video plays five decibels under the rest of the channel forever, and of course it is the one the algorithm decides to show to strangers. The ceiling of the manual track was never your ear. It is your evening.

The automated track moves that ceiling somewhere else. The mix leaves the render with the same standard on video one and on video three hundred, across every channel you run, in up to 63 languages, from a calendar you approve once. Nobody has to remember the routine, because the routine is the render itself. What limits you stops being your patience and becomes the credits on the plan you chose.

So keep both numbers in view. By hand, a good audio pass costs about ten minutes per video and it is the first thing to go when the week gets heavy. Automatically, the standard is identical on every upload, in as many channels as you want, and the only thing left for you to decide is how much volume you want to buy and which quality mode you want it in.

  • Manual: a meter, an export and a check on every single publication.
  • Manual: about 10 minutes per video, five hours a month at one a day.
  • Manual: the check is the first thing skipped when you are tired.
  • Automated: the same mix standard on video 1 and on video 300.
  • Automated: identical across every channel and up to 63 languages.

FAQ

Got questions? We've got answers.

What loudness should I export a YouTube video at?

Aim for integrated loudness around -14 LUFS with true peaks under -1 dBTP. That is the neighbourhood the platform normalises toward, so you play at the same level as the videos around you. Being a decibel over is harmless, because the loud side is only turned down. Being six under is permanent, because the quiet side is never turned up. If you only keep two numbers in your head, keep -14 and -1.

Does YouTube turn quiet videos up?

No, and that single fact is why the problem survives on so many channels. Normalisation applies gain reduction to material that arrives louder than the target and leaves everything quieter untouched. A video delivered at -20 LUFS plays six decibels under the field on every device, on every view, and no amount of promotion repairs it. The viewer will not raise the volume to compensate either, they will simply compare you with whatever plays next.

How do I check what happened to my audio after upload?

Right click the player, open Stats for nerds and read the content loudness line. A negative value means the platform pulled your audio down by that much, which means you delivered above the target and you are fine. Zero or a positive value means nothing was applied, so the level you exported is exactly what strangers hear. Check your own last three videos and then check the competitor whose videos always sound bigger than yours.

Is the target the same on Instagram, TikTok and Facebook?

Close, but not identical, and the platforms adjust it without announcing anything. In practice they normalise into a band somewhere between roughly -16 and -13 LUFS, so a master around -14 with true peaks under -1 dBTP travels well everywhere instead of being wrong in two places. If you need one file for five networks, that is the file. Building a separate master per network for a difference of one or two decibels is not worth the hour.

Why does my audio distort only after uploading?

Because the platform re-encodes it. Your export peaks at -0.1 dBFS and looks safe, but the lossy encoder reconstructs the waveform slightly differently and the peaks between samples can land above zero, where there is no headroom left to absorb them. It shows up on sibilants and hard consonants as crackle that was never in your file. Set the limiter to -1 dBTP with the meter in true peak mode and the distortion disappears without costing you any loudness.

Do I need audio software and plugins to get this right?

By hand, yes, plus the discipline to repeat the routine on every publication. In FalconVid the mix leaves the render already balanced between narration, music and effects, with the track dropping under the voice by itself, so there is no meter to open and no file to export. The narration costs 24 credits per minute on the premium voices and 3 credits per minute on the economy voice, which is 288 or 36 credits on a twelve minute video.

What if the audio comes out wrong on one video, do I regenerate everything?

No. The Studio plays the V1 before anything is published, and there you can change the music, shorten the intro or swap a piece of media without regenerating the video, or open the timeline when you want a precise cut. Regenerating a full twelve minute video costs 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium, so the adjustment in the Studio is exactly the thing that keeps those credits in your account.

Half of my audience watches with the sound off, does loudness still matter?

It matters for the half that listens, and it decides how the silent half meets you the day they turn the sound on. For that silent half, karaoke captions are switched on in every plan, in more than fifteen styles, with no credit line of their own because they are part of the render. Loudness keeps the listeners, captions keep the muted, and nobody should have to pick one of the two.

Set the level once, then let every upload arrive at it

FalconVid researches, writes, narrates, edits, captions and publishes to YouTube, Instagram, TikTok, Rumble and Facebook in up to 63 languages, from a calendar you approve once, with AI specialists working in parallel and a video ready in up to 30 minutes. The mix leaves the render already balanced between narration, music and effects, with the track dropping under the voice by itself and karaoke captions on in every plan, so the same audio standard ships on video 1 and on video 300. A twelve minute video costs 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium, and the narration inside it is 288 credits on the premium voices or 36 on the economy voice. Starter at $47 with 15,000 credits, 1 channel and 2 simultaneous generations is around 10 to 12 videos a month mixing economy with one in balanced. Pro is $97 with 30,000 credits, 5 channels and 5 simultaneous generations, up to Scale at $997 with 320,000 credits, 50 channels and 50 simultaneous generations. Every creation feature is on every plan, what changes is volume, channels, simultaneous generations, the Senior Analyst and support, plus a 7 day trial with 2,000 credits and a 7 day guarantee.

Create my channel now

Charged today · 7-day guarantee · Cancel anytime

Keep reading