MANUAL ORIGINAL

This tutorial is an original manually prepared language version. It does not use the article live-translation service.

Quick answer

Start with a clean voice file, lower music under speech, use gentle fades and automation, reduce competing midrange frequencies when needed, and review the final mix on headphones and an ordinary phone speaker.

Use the relevant BestAI tool with one short test file or paragraph first. Confirm the result, then process a larger batch. This prevents a small setting mistake from being repeated across many clips or a long narration.

Why this tutorial matters

Music can support emotion, but the audience leaves when words are difficult to understand. A mix that sounds impressive on headphones may hide consonants on phone speakers. Voice intelligibility should remain the priority.

The BestAI workflow is designed around a simple rule: one active control should own one decision. When a mode overrides another setting, the conflicting control is disabled in the interface and ignored by the backend. This makes the plan shown on screen match the actual result.

Recommended starting settings

SettingPractical starting point
Voice sourceUse WAV when further editing is planned.
Music levelStart clearly below the narration and adjust by ear.
FadesUse short smooth fades at starts, ends and transitions.
AutomationLower music more during dense or quiet speech.
ReviewHeadphones, phone speaker and low playback volume.

These values are starting points, not universal rules. Source quality, speaking style, platform, audience and the purpose of the video can require a different choice.

Step-by-step tutorial

Step 1: Prepare a clean voice master

Correct pronunciation, clipping and unwanted silence before adding music. Mastering cannot repair a wrong line hidden under a louder track.

Step 2: Choose music that leaves space

Dense vocals, strong lead instruments and busy midrange compete with speech. Instrumental music with a stable arrangement is easier to place under narration.

Step 3: Set the voice first

Bring the narration to a clear, comfortable level without music. Then add the background quietly instead of pushing both tracks upward.

Step 4: Automate around important lines

Lower music during names, numbers, instructions and emotional reveals. Allow slightly more music in visual sections without speech.

Step 5: Use EQ and fades carefully

If the music masks the voice, reduce some competing midrange rather than only turning everything down. Add smooth fades so music changes do not distract.

Step 6: Test on real devices

Listen at low volume on a phone. If words disappear, the mix is not finished even if it sounds wide and powerful on headphones.

Quality-control checklist

  • The voice is clean before mixing.
  • Music contains no distracting vocal or lead conflict.
  • Important words remain clear at low volume.
  • Music changes use smooth automation and fades.
  • The final mix is checked on a phone speaker.

Common mistakes to avoid

  • Making music loud because the edit feels empty.
  • Using a compressed MP3 as the only voice master for heavy processing.
  • Applying one fixed music level to every sentence.
  • Trying to solve masking only by raising the voice.
  • Reviewing only on studio headphones.

Frequently asked questions

How loud should background music be?

There is no universal number. Set it relative to the voice, content and device, then verify that every word remains clear at low playback volume.

Should I use WAV for the voice?

WAV is preferable when mixing and mastering because it avoids additional compression before the final export.

Can music be louder between sentences?

Yes. Gentle automation can increase music during pauses and visual moments, then lower it before speech resumes.

Practical example

A sad story uses piano music that competes with the narrator's midrange. Lowering the music alone makes the scene feel empty. Gentle automation under speech and a small reduction in the competing frequency range preserve emotion while restoring word clarity.

Protect the audience's attention

Music should guide mood without asking the listener to work harder to understand the script. When in doubt, make the narration clearer rather than the soundtrack more impressive. ## Final review note

Check copyright and licensing before using any music. Keep proof of the licence or source with the project files. A well-balanced mix is still unsuitable if the music cannot legally be used on the intended channel, client campaign or monetized video.

Final takeaway

Build the mix around intelligible speech. Clean the voice, choose supportive music, automate levels and judge the result on the devices your audience actually uses.

Sources and references