Skip to content

ElevenLabs settings for natural AI voiceover

Use these ElevenLabs settings for natural AI voiceover, then format and generate each line so your finished video sounds less robotic.

RunbookJuly 29, 20267 min read
ElevenLabs settings for natural AI voiceover
FIG. 01 — FEATURED

This sheet contains partner links. A purchase through one earns Runbook a commission at no additional cost to you. How we make money.

You can turn a finished script into a natural-sounding voiceover in about 30 minutes with ElevenLabs Text to Speech, the screen that converts written words into audio. You'll set a repeatable Multilingual v2 baseline, shape the delivery in the script, and export separate clips that are easy to repair without regenerating the whole video.

What you'll build

You'll create a small voiceover production system, not one giant audio file. Each sentence or performance beat becomes its own clip. That gives you control over timing, emphasis, and corrections when the video changes.

The settings in this guide are for an energetic cloned voice using Eleven Multilingual v2. In our testing, the strongest starting point was:

ControlStarting valueWhat it changes
Stability0.45How consistent the delivery stays between takes
Similarity0.85How closely the result follows the source voice
Style exaggeration0.35How strongly the source performance style comes through
Speed1.0The basic speaking pace

These values are a starting line. ElevenLabs says voice choice matters more than the model settings, and its Text to Speech product guide explains that lower stability creates more emotional range while very high stability can sound monotonous. The same settings can produce slightly different takes because the system includes controlled randomness.

Your move

Set Multilingual v2 to 0.45 stability, 0.85 similarity, and 0.35 style, then fix emphasis in the written script before moving another slider.

Stack

You need an ElevenLabs account, an approved voice, a finished script, headphones, and the video editor you already use. A cloned voice is a computer-made version of a real speaker. Only upload a voice you have permission to use.

For a quick clone, open Voices, select the plus icon, and choose Instant Voice Clone. ElevenLabs asks you to confirm that you have the right and consent to clone that voice. Its current cloning instructions recommend roughly one to two minutes of clean audio without background noise or room echo.

If you're still deciding whether the tool fits your production budget, check the ElevenLabs pricing breakdown. For a wider view of the category, the AI voiceover tool comparison explains where browser-based alternatives differ.

Step 1: choose the voice before touching the sliders

Open Text to Speech from the left sidebar. Paste one paragraph from your real script, then click the voice selector at the lower left.

Test two or three voices against the same paragraph. Keep the words and model unchanged during this test. You are listening for the right accent, age, pace, and natural energy. A calm documentary voice will not become a lively promotion just because you lower stability.

If you use a clone, check the recording first. Fan noise, music, another speaker, heavy echo, or changing microphone distance becomes part of the voice the system tries to reproduce. Fixing the source recording usually does more than another hour of slider changes.

Step 2: select Multilingual v2 for consistency

Open the Model menu and choose Eleven Multilingual v2. ElevenLabs describes this model as its stable option for long-form generation, with support for 29 languages. It is the safer place to begin when you want the same cloned voice across a series of videos.

Do not apply this slider recipe to Eleven v3. The v3 model is built for more expressive performances and uses written audio tags such as [whispers] or [laughs]. Similarity is not available there, and its stability control uses Creative, Natural, and Robust choices. The official v3 prompting guide also warns that Professional Voice Clones are not fully optimized for v3.

Use v3 when a scene needs acting. Use Multilingual v2 when repeatability matters more than surprise.

Step 3: enter the baseline settings

Open Voice settings and enter stability at 0.45, similarity at 0.85, style exaggeration at 0.35, and speed at 1.0.

Generate the same paragraph twice. Pick the take with clear words and believable emphasis, not simply the loudest take.

Then change only one control:

  • If the read is flat, lower stability by 0.05.

  • If the voice wanders or rushes, raise stability by 0.05.

  • If the clone loses its identity, raise similarity slightly.

  • If you hear odd breaths, unstable speed, or mispronunciations, lower style exaggeration. ElevenLabs recommends returning style to zero when it creates artifacts.

Keep speed near 1.0 until the performance is right. Speed is a finishing adjustment, with a documented range from 0.7 to 1.2. Extreme values can reduce quality.

Step 4: direct the voice with the script

The text carries more of the performance than the sliders do. Capitalize the one word that needs emphasis. Add commas where the speaker should breathe. Use an ellipsis for a deliberate pause. Spell unusual names the way they should sound if the normal spelling fails.

Avoid a page of short fragments. The model may add a breath or change its energy at every stop. Merge closely related fragments into one sentence so it can hear the full thought.

Write numbers as words when pronunciation matters. Replace $149 with “one hundred forty-nine dollars.” Replace a company abbreviation with the words a person would actually say. ElevenLabs also advises against unusual symbols and digits when they can be written plainly.

Copy this

Paste this short test before generating the real script. Replace the bracketed details, but keep the punctuation pattern for the first pass.

Your [service] should not leave customers waiting.

When a request arrives, your team sees it RIGHT away... and the follow-up starts before the lead goes cold.

Set it once, check the handoff, and keep the personal reply for the conversation that needs it.

Listen for three things: whether “right” receives useful emphasis, whether the pause sounds intentional, and whether the final line keeps one steady pace. If one part fails, rewrite that line before changing the global voice settings.

Step 5: generate one beat at a time

Split the approved script into sentences or natural performance beats. Generate each beat separately. A useful clip may contain two short sentences if they share one idea.

Save the selected takes with names that preserve their order:

01-hook.mp3
02-problem.mp3
03-process.mp3
04-proof.mp3
05-next-step.mp3

Open the History panel on the right side of Text to Speech if you need an earlier take. Use the download icon to save MP3 or WAV. On a narrow screen, the history icon appears above Generate speech.

Drop every clip onto its own position in your video editor. Trim empty space, then move the clip to the picture instead of forcing the entire edit around one long recording. Keep music quiet beneath the words. Your customer should never have to work to hear the message.

The part that breaks

Most weak voiceovers fail before generation. The script was written to be read on a page, not spoken aloud.

Read each line yourself. If you run out of breath, split it. If the key word is buried at the end of a long sentence, move it forward. If three consecutive lines have the same length, vary the rhythm.

The other common failure is changing stability, similarity, style, speed, voice, and punctuation together. That destroys your test. You cannot tell which change helped. Keep one reference line, change one input, and label the winning take.

Credit waste follows the same pattern. Generating a full page to repair one mispronounced name spends more than regenerating a single clip. Approve the text first, then generate in small units.

Upgrade path

Once the baseline works, create three saved script patterns: an energetic opening, a calm explanation, and a direct closing. Record the chosen settings beside each pattern so another person can reproduce the delivery.

Test Eleven v3 separately for scenes that genuinely need emotion. Add one audio tag at a time and compare it with the Multilingual v2 take. Do not move the whole production to v3 because one dramatic line sounds good.

For a closer look at that buying decision, use the ElevenLabs versus Murf comparison. Then browse the next setup in AI voice and video builds. Get the next runbook.

Frequently asked questions

What are the best ElevenLabs settings for a natural voice?

For an energetic cloned voice on Multilingual v2, start with stability at 0.45, similarity at 0.85, and style exaggeration at 0.35. Treat these as a tested baseline, then adjust one control at a time.

Why does my ElevenLabs voice sound robotic?

High stability, a flat source voice, weak punctuation, or generating a long script in one pass can flatten the delivery. Fix the text and split it into short performance beats before changing every slider.

Should I use Eleven v3 or Multilingual v2 for voiceover?

Use Multilingual v2 when a consistent cloned voice matters. Test Eleven v3 when you need expressive directions such as whispers or laughs, but expect more variation between generations.

How should I generate ElevenLabs voiceover for video?

Generate one sentence or performance beat at a time, save the chosen take with a numbered filename, and place each clip separately on the video timeline.

About Runbook

AI tools and automation builds for marketers. What to use, how to wire it, and the workflow to copy this week. How we work

GET THE NEXT DISPATCH

Run the next build before your competitors read about it.

One short email when an AI tool or automation actually changes the work, with the build to copy.

No send unless there is a build worth running.

// keep_reading

Related builds