How to Make an AI Voice
Two completely different processes go by this name: generating a voice with a consumer tool in minutes, or training a model from scratch.
We may earn a commission when you buy through links on this page. It never changes what we recommend โ we only write tutorials for tools we actually use, and the price you pay is the same either way. How we make money
How to Make an AI Voice: The Short Version
- I chose a platform: a text-to-speech or voice-cloning tool like ElevenLabs or Murf.
- I picked a generation method: I selected a pre-made stock voice, described a brand-new custom voice with a text prompt, or uploaded a clear audio sample to clone a real voice (my own or someone else's, with their consent).
- I entered the script: I typed or pasted the text I wanted the AI to read.
- I fine-tuned the delivery: I adjusted sliders for speed, stability, and expressiveness, or added emotion tags (like a pause, a laugh, or a specific tone) to make it sound more natural.
- I generated and downloaded: I previewed the result, and exported as an MP3 or WAV file once I was happy with it.
For most people, this entire process takes 10-15 minutes from start to finished audio file. The full walkthrough below covers the one path that needs the extra time: cloning your own voice properly, rather than using a stock one. The rest of the coverage in this space is in AI Audio & Music Tools.
Two Different Paths: Using a Tool vs. Building One From Scratch
This distinction matters because "how to make an AI voice" covers two completely different processes, and almost everyone means the first one: using an AI voice generator, a consumer tool that turns text into speech or clones a voice from a sample in minutes, no technical skill required. The other path, building an AI voice model from scratch, means training your own text-to-speech model using frameworks like PyTorch or TensorFlow, which requires machine-learning expertise, a large and clean dataset of voice recordings, and real computational resources (cloud infrastructure or a capable GPU). Unless you're specifically trying to build voice technology as a product or research project, the tool-based path below is what you want.
How Realistic Can an AI Voice Actually Get?
AI voices can sound realistic enough that one voice AI company claims to have reached a real milestone worth knowing about (this is the company's own reported claim, not independently verified here): in a 2020 blind evaluation, participants rated the company's synthetic voice output at an average naturalness score of 4.5, matching the score given to real human voice actors in the same test. Whatever the exact number, the practical takeaway holds up across the category: modern AI voices, properly trained and tuned, can sound close enough to human that casual listeners often can't reliably tell the difference in short clips.

A Full Walkthrough: Cloning Your Voice in ElevenLabs
Voice cloning is one of the few AI features that works about as well as the marketing suggests. The catch is that the sample you upload in the first five minutes decides almost everything about the result, and most people rush that part, get a mediocre clone, and conclude the technology isn't there yet.
This walkthrough covers the whole loop: recording a sample worth cloning, training the voice, getting delivery that doesn't sound like a train announcement, and budgeting so you don't run out of credits halfway through a project.
Key takeaways
- 01Sample quality decides clone quality: a clean two minutes beats a noisy ten
- 02Record in the register you will narrate in, not your reading-aloud voice
- 03Stability and similarity are a trade-off, not a quality slider: tune them per project
- 04Every character generated costs credits, including every take you throw away
What you need first
- A quiet room. Not a studio, somewhere without traffic, fans or hard echo.
- Any decent microphone. A USB condenser or a modern phone held close both work.
- Two to three minutes of you speaking naturally.
- An ElevenLabs account. Voice cloning sits on the paid tiers; the free tier is enough to audition the stock voices and the interface, not to clone yourself.
Only clone a voice you have the right to use
Cloning your own voice is fine. Cloning someone elseโs without written permission is a legal problem in most jurisdictions and a violation of the terms of service everywhere. If it is a clientโs voice, get the consent in writing before you upload anything.
Step 1: Record the sample properly
This is the step that matters. Everything else depends on it.
I recorded two to three minutes of continuous natural speech. Not a script read in an announcer voice: I talked, explaining something I know well to an imaginary person. The clone learns your rhythm, your hesitations and your default pitch, so recording in a voice you'd never use gets you a clone you'll never use.
I followed five practical rules:
- One consistent distance from the mic. I didn't lean in and out.
- No background music, no other voices, no room echo. The model can't separate them out; it learns them as part of you.
- I didn't edit out every breath. Breaths make the clone sound human.
- I matched the target energy, recording upbeat where the narration called for it.
- I exported uncompressed, WAV where I could, high-bitrate MP3 where I couldn't.
More audio doesn't help past a few minutes. Three clean minutes outperforms ten minutes with a fan humming in the background.
Step 2: Train the voice
I uploaded the sample in the voice library, named it something I'd recognise in six months (William - narration, warm), and added a short description of the accent and tone. The description isn't cosmetic: it feeds the model.
Training itself was fast, usually a minute or two.
Ready to try ElevenLabs?
Start on the free plan โ paid plans from $6/mo.
Step 3: Audition before you commit
I generated the same three test sentences through the new voice, every time:
- A neutral factual sentence.
- A question.
- A sentence with a proper noun or acronym specific to the subject.
Sentence three is the one that exposes problems. Clones handle ordinary prose well and mangle domain-specific words: product names, technical terms, place names. Better to find that now than in the final render.
When the clone sounded close but slightly off, the fix was nearly always a better sample, not more tweaking. I re-recorded before spending an hour on sliders.
Step 4: Tune stability and similarity for the job
These two settings do most of the work, and they pull against each other.
Low stability
More expressive
Varies delivery between generations. Good for character work and conversational video. Can wander.
High stability
More consistent
Flat, predictable delivery. Good for long documentation narration where nothing should surprise the listener.
High similarity
Closer to you
Reproduces your sample faithfully, including any noise or quirks in it. Push it up only if the sample was clean.
I start narration with stability around the middle and similarity high, then adjust in one direction only, one setting at a time. Changing both at once means not knowing which one helped. This walkthrough covers the settings that matter for narration; the full depth of what ElevenLabs exposes is in the ElevenLabs Tutorial.
Low stability is also the setting to reach for when the goal is a character voice rather than narration. A real-time voice changer for talking live on a stream or in Discord is a different tool and job entirely; see the Voicemod Review.
Step 5: Generate in chunks, not in one block
I didn't paste 4,000 words and press generate. I split at natural section breaks and generated a few paragraphs at a time, for three reasons:
- A bad take cost one paragraph of credits, not the whole script.
- I could re-roll a single awkward sentence without touching the rest.
- Long single generations drift more in tone from start to finish.
Punctuation became my direction. A full stop is a real pause; a comma is a short one; an ellipsis produces hesitation. When a line read too fast, I broke the sentence rather than reaching for a setting.
Step 6: Do the credit maths before the project, not during it
This is where people get caught. Every character generated costs credits, and every discarded take counts. A realistic estimate:
- A 1,500-word script runs roughly 9,000 characters.
- Expect to generate 1.5-2x the final length once re-rolls are included.
- Budget around 15,000-18,000 characters for a 1,500-word narration.
I checked my plan's monthly character allowance against that number before starting a series, not after episode three. The pricing on the entry tiers looks generous until multiplied by re-takes.
Keep your best takes
I downloaded every take I was happy with immediately and kept them in a project folder. Regenerating a line I already had would cost credits and might not come back identical.
Step 7: Assemble
I brought the chunks into my editor, laid them against the video, and adjusted the gaps between sections rather than regenerating for timing. Silence is free; generation isn't.
Two finishing touches made the clone sound noticeably more professional:
- I left the breaths in. Editors instinctively cut them. I didn't.
- I added a touch of room tone under the whole track. Perfectly silent gaps are the single biggest giveaway that narration is synthetic.
Is it good enough to publish?
For narration over video, explainers, documentation and audiobooks, yes, on a good sample, in my testing. For anything where a listener is looking for emotional performance, it's still recognisably synthetic to an attentive ear.
For the full picture on where it holds up and where the pricing bites:
ElevenLabs ReviewTips for Better Results With Any Tool
The above walkthrough is ElevenLabs-specific, but these apply whichever platform you use:
- I clean up the script first. Well-punctuated, clearly-structured text produces more natural-sounding delivery than messy or ambiguous phrasing.
- I provide pronunciation guidance for unusual terms. Names, brands, and technical jargon are where AI voices most often stumble; some tools let you add phonetic spelling to fix this.
- I test multiple voices before committing. The same script can sound completely different depending on the voice and settings, so I preview a few options rather than settling on the first one.
- I use short sentences and natural pauses. AI voices tend to read exactly what's on the page, so conversational phrasing with clear pause points sounds more natural than dense, run-on sentences.
- I iterate based on what I hear. Small adjustments to pacing, emphasis, and tone after listening back typically improve results more than getting the settings perfect on the first try.
Is Voice Cloning Legal?
Voice cloning itself carries no blanket illegality. It's a legitimate feature offered by major platforms, generally fine for your own voice or a voice you have explicit permission to clone. The legal and ethical problems arise specifically around consent and intent: cloning someone's voice without their permission, especially to impersonate them, deceive people, or commit fraud, can violate publicity rights, fraud laws, and in a growing number of jurisdictions, specific deepfake-related regulations. Reputable platforms require you to confirm you have the rights to any voice you clone precisely because of this. If you're cloning your own voice or a voice with clear, documented consent for a legitimate purpose, you're on solid ground; if you're trying to replicate someone else's voice without asking them, don't.
Where This Fits in the Wider AI Audio Toolkit
This guide sits in the AI Audio & Music section. The ElevenLabs Review covers whether the tool behind most of this workflow is worth paying for, Free AI Voiceover Tools covers what is usable without paying at all, and The Best Text to Speech Software separates consumer voiceover tools from developer APIs before you pick either.
Frequently asked questions
Can I create my own AI voice?
Yes. Most major voice AI platforms let you clone your own voice from a short audio sample (as little as a few seconds to a minute, depending on the tool), producing a synthetic version you can use to generate new speech without recording it yourself each time. See the full walkthrough above for how to get a good result rather than a mediocre first attempt.
How can I make an AI voice for free?
Several platforms offer a free tier for basic voice generation, typically with a character or time cap, plus fully free open-source options if you're comfortable with some technical setup. See Free ElevenLabs Alternatives for free options.
Can I use voice AI for free?
Yes, most major platforms including ElevenLabs offer a free tier sufficient for testing and light use, though features like professional voice cloning and higher usage limits typically require a paid plan.
Is voice cloning illegal?
No, not on its own. Cloning your own voice or a voice you have explicit permission to use is legitimate. It becomes a legal problem specifically when used to impersonate someone without consent, especially for deception or fraud, which can violate publicity rights and fraud laws.
We may earn a commission when you buy through links on this page. It never changes what we recommend โ we only write tutorials for tools we actually use, and the price you pay is the same either way. How we make money

ElevenLabs Review

Suno Review

Voicemod Review
Get the AI tool roundup.
One email a week: what launched, what is actually worth paying for, and what to cancel. No hype, no spam.
Unsubscribe any time. We never sell your email.