How to Clone Your Voice in ElevenLabs and Narrate a Video
From a clean recording to a finished narration track — how to train a usable voice clone, get natural delivery out of it, and avoid the credit maths that catches people out.
We may earn a commission when you buy through links on this page. It never changes what we recommend — we only write tutorials for tools we actually use, and the price you pay is the same either way. How we make money
Voice cloning is one of the few AI features that works about as well as the marketing suggests. The catch is that almost everything about the result is decided in the first five minutes — by the sample you upload — and most people rush that part, get a mediocre clone, and conclude the technology is not there yet.
This tutorial covers the whole loop: recording a sample worth cloning, training the voice, getting delivery that does not sound like a train announcement, and budgeting so you do not run out of credits halfway through a project.
Key takeaways
- 01Sample quality decides clone quality — a clean two minutes beats a noisy ten
- 02Record in the register you will actually narrate in, not your reading-aloud voice
- 03Stability and similarity are a trade-off, not a quality slider — tune them per project
- 04Credits are consumed per character generated, including every take you throw away
What you need first
- A quiet room. Not a studio — just somewhere without traffic, fans or hard echo.
- Any decent microphone. A USB condenser or a modern phone held close both work.
- Two to three minutes of you speaking naturally.
- An ElevenLabs account. Voice cloning sits on the paid tiers; the free tier is enough to audition the stock voices and the interface, not to clone yourself.
Only clone a voice you have the right to use
Cloning your own voice is fine. Cloning someone else’s without written permission is a legal problem in most jurisdictions and a violation of the terms of service everywhere. If it is a client’s voice, get the consent in writing before you upload anything.
Step 1 — Record the sample properly
This is the step that matters. Everything downstream is downstream of it.
Record two to three minutes of continuous natural speech. Not a script read in your announcer voice — talk. Explain something you know well to an imaginary person. The clone learns your rhythm, your hesitations and your default pitch, so if you record in a voice you would never actually use, you get a clone you will never actually use.
Practical rules:
- One consistent distance from the mic. Do not lean in and out.
- No background music, no other voices, no room echo. The model cannot separate them out; it learns them as part of you.
- Do not edit out every breath. Breaths make the clone sound human.
- Match the target energy. If the narration will be upbeat, record upbeat.
- Export uncompressed — WAV if you can, high-bitrate MP3 if you cannot.
More audio is not better past a few minutes. Three clean minutes outperforms ten minutes with a fan humming in the background, every time.
Step 2 — Train the voice
Upload the sample in the voice library, name it something you will recognise in
six months (Minh — narration, warm), and add a short description of the
accent and tone. The description is not cosmetic — it feeds the model.
Training itself is fast, usually a minute or two.
Ready to try ElevenLabs?
Start on the free plan — paid plans from $5/mo.
Step 3 — Audition before you commit
Generate the same three test sentences through your new voice, every time:
- A neutral factual sentence.
- A question.
- A sentence with a proper noun or acronym specific to your subject.
Sentence three is the one that exposes problems. Clones handle ordinary prose well and mangle domain-specific words — product names, technical terms, place names. Better to find that now than in the final render.
If the clone sounds close but slightly off, the fix is nearly always a better sample, not more tweaking. Re-record before you spend an hour on sliders.
Step 4 — Tune stability and similarity for the job
These two settings do most of the work, and they pull against each other.
Low stability
More expressive
Varies delivery between generations. Good for character work and conversational video. Can wander.
High stability
More consistent
Flat, predictable delivery. Good for long documentation narration where nothing should surprise the listener.
High similarity
Closer to you
Reproduces your sample faithfully — including any noise or quirks in it. Push it up only if the sample was clean.
A workable starting point for narration is stability around the middle and similarity high, then adjust in one direction only, one setting at a time. If you change both at once you will not know which one helped.
Step 5 — Generate in chunks, not in one block
Do not paste 4,000 words and press generate. Split at natural section breaks and generate a few paragraphs at a time. Three reasons:
- A bad take costs you one paragraph of credits, not the whole script.
- You can re-roll a single awkward sentence without touching the rest.
- Long single generations drift more in tone from start to finish.
Punctuation is your direction. A full stop is a real pause; a comma is a short one; an ellipsis produces hesitation. If a line reads too fast, break the sentence rather than reaching for a setting.
Step 6 — Do the credit maths before the project, not during it
This is where people get caught. Credits are consumed per character generated, and every discarded take counts. A realistic estimate:
- A 1,500-word script is roughly 9,000 characters.
- Expect to generate 1.5–2x your final length once re-rolls are included.
- So budget around 15,000–18,000 characters for a 1,500-word narration.
Check your plan's monthly character allowance against that number before you start a series, not after episode three. The pricing on the entry tiers looks generous until you multiply it by re-takes.
Keep your best takes
Download every take you are happy with immediately and keep them in a project folder. Regenerating a line you already had costs credits and may not come back identical.
Step 7 — Assemble
Bring the chunks into whatever editor you already use, lay them against the video, and adjust the gaps between sections rather than regenerating for timing. Silence is free; generation is not.
Two finishing touches that make a clone sound noticeably more professional:
- Leave the breaths in. Editors instinctively cut them. Do not.
- Add a touch of room tone under the whole track. Perfectly silent gaps are the single biggest giveaway that narration is synthetic.
Is it good enough to publish?
For narration over video, explainers, documentation and audiobooks — yes, on a good sample. For anything where a listener is looking for emotional performance, it is still recognisably synthetic to an attentive ear.
For the full picture on where it holds up and where the pricing bites, see the ElevenLabs review.
We may earn a commission when you buy through links on this page. It never changes what we recommend — we only write tutorials for tools we actually use, and the price you pay is the same either way. How we make money
ElevenLabs Review: The AI Voice That Finally Passes
How to Write a Blog Post With ChatGPT Without It Sounding Like ChatGPT
Get the AI tool roundup.
One email a week: what launched, what is actually worth paying for, and what to cancel. No hype, no spam.
Unsubscribe any time. We never sell your email.