The voice sets the floor: narration at about -16 LUFS
A mix under narration has exactly one anchor, and it is the voice. Set the narration to roughly -16 LUFS integrated with true peaks under -1 dBTP and every other decision in this article turns into a number you can measure instead of a feeling. LUFS is loudness the way a listener perceives it across time, not the tallest spike in the waveform, which is why a track can peak in the same place as your voice and still bury half of it.
The reason -16 LUFS is the useful target is that platforms normalize. YouTube turns down anything meaningfully louder than about -14 LUFS and it does not turn quiet material up, and Instagram, TikTok and Facebook land in the same neighbourhood. Normalization is one single fader applied to the whole program, so you cannot rescue a bad music to voice balance by exporting louder. If the track sits 8 dB under the narration in your project, it sits 8 dB under the narration on the phone of every viewer, at every volume setting, forever.
Two numbers to keep on screen while you work: short term loudness during speech hovering between -16 and -14 LUFS, and true peak never crossing -1 dBTP, which leaves the lossy encoder somewhere to breathe without clipping. On a faceless channel the voice carries the entire video, since there is no face and no camera performance doing any of that work, and the mix has to say so out loud.
How far under the voice the music belongs: 18 to 24 dB
Here is the number the whole article turns on. While the narrator is speaking, the music bus should sit 18 to 24 dB below the voice. In plain meter terms, if your narration peaks around -6 dBFS, the music peaks land between -26 and -30 dBFS. Solo the music at that level and it sounds almost embarrassingly quiet. Put the voice back on top and it sounds exactly right, which is the entire trick.
Most channels land at 8 to 12 dB under instead, and the reason is always the same room: you mix at night, in headphones, with no background noise, and a bed at -10 dB feels cinematic. Then the video plays on a phone at half volume in a kitchen with a fan running, and that bed is chewing consonants off the ends of words. Within roughly 10 dB of the voice, music masks the 2 to 6 kHz region where t, k, s and f live. The brain reconstructs the missing pieces, but it costs attention. Nobody thinks the mix is bad. They just get tired, and tired viewers leave without knowing why.
Adjust the target by genre rather than by taste. Informative, finance and how to content wants 22 to 24 dB of separation, because every sentence carries something the viewer cannot afford to miss. Story, true crime and documentary live at 16 to 20 dB, where the music is doing real emotional work. Motivational content can push to 14 to 18 dB, but only with a track that has no melodic movement in the voice band. Below 14 dB you are no longer scoring the video, you are competing with it.
Ducking: 6 to 10 dB, with the voice pulling the trigger
A fixed level is not enough, because narration is not fixed. It gets loud on an emphasis, drops on an aside, pauses to breathe. Ducking, also called sidechain compression, puts a compressor on the music and lets the voice trigger it: when the narrator speaks the music drops, when the narrator stops it comes back. Target 6 to 10 dB of gain reduction while speech is present.
The settings matter more than the plugin. Ratio of 4:1 to 8:1. Threshold set so normal speech triggers full reduction while breaths do not. Attack of 10 to 30 ms, fast enough to catch the first syllable, slow enough not to click. Hold of 80 to 150 ms so the gaps between words do not open the gate. Release of 250 to 400 ms, which is what makes the music return like a decision instead of a reflex. Every failure has a signature: pumping on every word means the release is under 100 ms, a swell after each sentence means it is over 800 ms, a swallowed first syllable means the attack is too slow, and music that dips when the narrator merely inhales means the threshold is too low.
Stack the two mechanisms rather than choosing between them. Set the static bed 18 dB under the voice, then let the ducker take another 6 to 10 dB while words are actually happening. During speech the music is effectively 24 to 28 dB down, and in the pauses it lifts back to that 18 dB bed, so the track breathes with the narration instead of fighting it. That breathing is what listeners register as professional audio, even though they could never name it.
This is the exact stage where FalconVid stops charging you an hour per video. The sound designer runs in parallel with the researcher, the scriptwriter, the narrator and the editor, and it does not have to guess where the speech is: the engine generated the narration track itself, so it knows every word boundary and every pause to the millisecond. The ducking curve is written from that, not detected after the fact, and the first version comes back mixed in up to 30 minutes.

Volume is half the problem: high-pass the music around 200 Hz
Two sounds only fight where they share frequencies. A speaking voice has its fundamental between roughly 85 and 180 Hz for most men and 165 to 255 Hz for most women, its intelligibility between 1 and 4 kHz, and its sibilance between 5 and 8 kHz. Music keeps most of its energy under 250 Hz, in the kick and the bass. That low end is not masking your consonants, but it is eating headroom, and headroom is what normalization takes back from you later.
So put a high-pass filter on the music between 150 and 250 Hz, with 200 Hz as the default that works nine times out of ten. Sweep it while the voice plays, never with the music soloed, because soloed it sounds thin and you will talk yourself out of it. Under narration it sounds correct, and it typically frees 3 to 6 dB of headroom that goes straight back to the voice. Then carve the presence band: a wide dip of 2 to 4 dB between 1.5 and 3 kHz opens the exact window where consonants work, and a dynamic EQ keyed to the voice keeps the track at full character in the gaps.
The device check makes all of this concrete. Phone speakers roll off almost everything under 300 to 500 Hz, so the bass you fought to keep is inaudible on the device where most of your audience is, while on a soundbar or in headphones that same bass is the thing making your mix muddy. Cutting it costs you nothing on the phone and buys clarity everywhere else.
Where the track is allowed to rise
The rule fits in one line: the music only goes up where there is no speech. Everything else is negotiable, this is not. A track that swells while a sentence is running does not sound cinematic, it sounds like a mistake, and the viewer resolves the confusion by leaving. The good news is that a 12 minute video has plenty of legal places to be loud.
Give the rise real numbers so it is a move and not a wobble. In those gaps let the music come up to 8 to 12 dB below the voice level, a lift of roughly 10 to 14 dB from the speaking bed, with fades of 1 to 2 seconds on the way up and 0.5 to 1 second on the way down so the narrator never talks over a decaying swell. Automate it on the timeline, because a ducker only knows how to get out of the way, it does not know how to make an entrance. One to two seconds of music forward audio between blocks works as punctuation: it says a chapter closed, and it buys the next block a moment of goodwill.
This is another place where FalconVid is playing with the answer sheet. The engine wrote the script and generated the narration, so it knows where the intro ends, where each block hands off to the next, and where the outro begins. The music rides that structure automatically: forward in the intro, forward on transitions, down under speech, forward again on the end card. You approve the calendar, and the videos come out already mixed to that shape.
- The intro, roughly 0:00 to 0:08, before the narration starts talking
- Transitions between blocks, 1 to 2 seconds of music forward audio
- Any pause longer than about 1.5 seconds, including deliberate dramatic beats
- Pure b-roll stretches where nothing is being narrated at all
- The reveal or payoff moment, where the lift is doing emotional work
- The outro and end card, from the last sentence to the final frame
Where music really costs retention: the minute 2 to 3 fatigue
Loud music almost never produces a cliff in the retention graph. It produces a ramp, and ramps are harder to notice and much more expensive. The typical shape is 2 to 3 points lost every 15 seconds from about 1:30 to 3:00, with no single second you can point at. Nothing bad happened at any timestamp. The listener simply spent ninety seconds working slightly harder than they wanted to, and their patience ran out somewhere in minute three.
Constant epic orchestral music under informative content is the classic version of this. The track makes an emotional promise every eight bars that the content never pays, the brain habituates to the drama in 60 to 90 seconds, and after that the music is just noise with a budget. For informative material use a minimal bed at 60 to 90 BPM, no lead melody in the 1 to 4 kHz voice band, no dynamic swings, then set it 22 dB down and forget it exists.
The opposite mistake is real too. Total silence under a story video also loses people, because silence exposes every edit and the viewer starts hearing the assembly instead of the story. A bed 24 to 30 dB below the voice makes cuts disappear, so save true silence for the 2 to 4 seconds before a reveal. And watch the loop: one track running 12 straight minutes gets noticed around its third repeat, usually between minutes 5 and 7. Change the cue every 3 to 4 minutes, at a block boundary and never mid idea, which on a 12 minute video means 3 or 4 tracks.
FalconVid handles that alignment because it built the blocks in the first place, so cue changes land on the structural seams rather than on a stopwatch. And when you disagree with a choice you do not go back to a timeline: you watch the first version in the Studio and swap the music there, or shorten the intro, or change a piece of media, and the video reassembles itself and goes out on the calendar you already approved.
Where FalconVid does the mixing for you
Everything above is roughly 30 to 50 minutes of careful work per video for someone who already knows how, and a two week learning curve for someone who does not. That is the honest price of a correct mix by hand, charged again on every upload, forever. The reason it barely registers inside FalconVid is architectural: sound design is a specialist running in parallel with the researcher, the scriptwriter, the narrator and the editor, so the mix is not a stage bolted onto the end of the pipeline. The whole video comes back ready in up to 30 minutes.
The track itself comes from the built in library of music and effects, with BGM, SFX and Freesound sources, and it arrives already mixed against the narration: level set, ducking applied, low end cleared. That answers the licensing question in the same move, because the sourcing problem for a faceless channel was never finding a nice song, it was finding one that will not collect a Content ID claim and quietly send your revenue to a rights holder in month four.
The voice matters here more than people expect. Ultra realistic premium narration from Cartesia has body and consistent level, which is precisely what survives a music bed: a thin, uneven voice needs the track pushed 28 dB down just to stay intelligible, while a full one holds its own at 18 to 20. You get 63 narration languages on top of that, plus cloning your own voice. And if the machine picks a track you hate, the Studio is one adjustment instead of a re export: by hand, redoing a bed on a finished 12 minute video is 1 to 2 hours of re importing, re balancing, re ducking and re rendering.
The economics are plain, and the mode weighs more than the plan. A 12 minute video costs 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium. Starter is $47 a month with 15,000 credits, which is 14 videos in pure economy or a realistic 10 to 12 a month mixing economy with one in balanced. Pro is $97 with 30,000 credits, 5 channels and 5 simultaneous generations, up to Scale at $997 with 50 channels and 50 simultaneous generations. Every creation feature ships on every plan; what changes by plan is volume, channels, simultaneous generations, the Senior Analyst (from Pro, with 7 free days on Starter) and support. There is a 7 day trial carrying 2,000 credits and a 7 day guarantee. And if music is the point rather than the bed, music channels are open on every plan too: hours long instrumental sessions for lofi, sleep, meditation, frequencies, yoga and law of attraction, or a sung artist channel that keeps the same face and the same voice with the lyrics captioned in sync.
The mixing checklist and the phone test
Keep this in a note next to your export button. Nine checks, every one a number rather than an opinion, and together they are the difference between audio that disappears and audio that costs you the minute 3 exit.
Then run the only listening test that matters. Play the finished video on a phone, speaker only, at medium volume, at arm's length, in a room with real noise in it. If you lose a single word, the music is 3 to 6 dB too loud, because that phone is where the majority of your audience lives. Then check the two extremes: closed headphones for low end mud and pumping ducking, both invisible on a phone speaker, and a TV or soundbar for the bass you left in. Three devices, five minutes, and it catches almost everything.
Now the honest arithmetic. Doing the whole list by hand is 30 to 50 minutes per video plus the ear training to know what you are hearing, which caps a solo creator at two or three properly mixed videos a week before quality starts slipping. That is the ceiling of the manual workflow, not the ceiling of the channel. With FalconVid the mix arrives already done on every video, cue changes land on the block seams, the library is already clear of Content ID trouble, and the only thing left for you is approving the calendar. The ceiling stops being your ears and becomes how many videos you decided to publish.
- Narration at about -16 LUFS integrated, true peak under -1 dBTP
- Music bus 18 to 24 dB below the voice during every spoken passage
- Ducking of 6 to 10 dB, attack 10 to 30 ms, release 250 to 400 ms
- High-pass filter on the music at roughly 200 Hz, swept against the voice
- A wide 2 to 4 dB dip in the music between 1.5 and 3 kHz
- Music forward only in the intro, transitions, pauses and outro
- Never a track with vocals or lyrics behind narration, instrumental only
- A new cue every 3 to 4 minutes, always on a block boundary
- Library music cleared for Content ID before the first upload, not after

