ElevenLabs Tutorial
I wasted credits early on by defaulting to the newest, most expressive model for jobs that needed something else — here's the routine that fixed it.
We may earn a commission when you buy through links on this page. It never changes what we recommend — we only write tutorials for tools we actually use, and the price you pay is the same either way. How we make money
Getting Started: Your First Voiceover in 5 Steps
- I created an account at elevenlabs.io. The free tier gives 10,000 monthly credits, enough to experiment with roughly 10 minutes of audio.
- I opened Text to Speech in the dashboard.
- I chose a voice: I browsed the Voice Library (filterable by accent, gender, and use case) or selected a voice I'd already cloned. I previewed it on the actual 100-150 words I planned to use, since a popular library voice isn't automatically right for a specific project.
- I selected a model for the job (see below, this is where most beginners waste credits).
- I pasted my script, generated, and listened before scaling up to a longer project.
For the rest of my AI Audio & Music Tools coverage, see the full hub.
Which Model Should You Actually Use?
The single most useful thing to get right before spending real credits: I don't default to the newest or most expressive model for every job.
| Use Case | Model | Why |
|---|---|---|
| YouTube narration, ads, storytelling | Eleven v3 | Best emotional range, audio tags, dramatic delivery |
| Courses, audiobooks, long-form lessons | Multilingual v2 | More stable and consistent across a long script |
| Fast drafts, agents, live/interactive use | Flash v2.5 | Lower latency and cheaper per-character cost |
Testing a voice clone, I don't judge it from a single generation. I run a short script through 2-3 models and compare neutral, excited, and slow-narration takes before committing to one.
Settings: Why v3 and v2-Class Models Work Completely Differently
This tripped me up early: ElevenLabs' settings don't work the same way across every model, and assuming they do wastes credits fast. For v2-class models (Multilingual v2, etc.), I start around 50% stability and 75% similarity, then adjust one control at a time. Higher stability means more consistency, lower stability allows more natural variation, and pushing similarity too high can reproduce artifacts from the source sample. Eleven v3 doesn't expose the same sliders at all. Instead, it uses a stability mode combined with the script's punctuation, line breaks, and audio tags to control delivery. One mistake I made early on: v3 does not support SSML break tags. If you're used to inserting them for pauses, they won't work on v3. I use audio tags, ellipses, and line breaks instead.
Adding Emotion and Pauses with Audio Tags
On Eleven v3, I can direct the performance directly inside my script using bracketed tags:
[whisper]before a sentence for an intimate, quiet delivery[excited]for an energetic, announcement-style tone[pause 0.5s]to insert deliberate dramatic timing
This lets me direct the AI the way I'd direct a voice actor, rather than accepting whatever flat, default delivery the model produces on a plain script.
Instant vs. Professional Voice Cloning
ElevenLabs offers two distinct cloning paths, and I found picking the wrong one for my goal wastes both time and credits:
- Instant Voice Cloning: ready in about a minute, needs only a 60-second audio sample, available from the Starter plan up. I found it good enough for quick social content or internal drafts, but not the highest fidelity.
- Professional Voice Cloning: needs 30+ minutes of clean, consistent audio and a longer review turnaround, so I planned ahead rather than expecting same-day results. Available from the Creator plan up. It gave meaningfully higher fidelity when the training audio was clean, and I'd treat it as a real business asset if I used it weekly.
Before I clone either way, I only clone my own voice or one I have explicit, documented permission to use. I record in a quiet room with a decent microphone and include some emotional range (neutral, energetic, slow) in the sample rather than one flat reading, since a clone that's only heard one mood struggles when I ask for range later. I start with a 150-word test script, not a 3,000-word one, to confirm the voice and settings before scaling up.
Common Mistakes That Waste Credits
- Using the free plan for commercial work. The free tier has no commercial license; monetized YouTube videos, client work, or ads need a paid plan.
- Generating a full script before testing settings. I always test a short sample first, then scale once the voice, model, and pacing are right.
- Assuming Eleven v3 is always the best choice. It isn't for long-form stability or cost-sensitive work; see the model table above.
- Using SSML break tags with v3. They're not supported; I use audio tags and punctuation instead.
- Training a voice clone on noisy audio. Background noise, music, and inconsistent mic distance all show up in the final clone.
- Relying on a shared library voice for long-term brand work without documenting it. Voice-library availability can change, so when a voice matters to my brand, I record its exact voice ID and settings.
Using the API
For developers integrating ElevenLabs into an app or workflow: I got my API key from my account's profile settings, then sent a POST request with my text and voice settings, and the response returned a ready-to-use audio file. The basics are straightforward; the API also supports multiple languages and typical response times around 400ms, which makes it usable for interactive or near-real-time applications, not only batch-generated audio files.
Where This Fits in the Wider AI Audio Toolkit
This walkthrough sits in the AI Audio & Music section. The ElevenLabs Review covers whether the subscription earns its price before you spend time on the workflow, and Free ElevenLabs Alternatives covers the options if it does not. For the same job done across tools rather than inside this one, see How to Make an AI Voice.
Frequently asked questions
How do I use ElevenLabs?
I sign up for a free account, open Text to Speech, choose a voice and a model matched to my use case, paste the script, adjust settings if needed, and generate. The whole first-time process takes a few minutes.
What are the best settings for ElevenLabs voice?
It depends on the model. For v2-class models, I start around 50% stability and 75% similarity and adjust one at a time. Eleven v3 doesn't use those same sliders; I direct it instead through its stability mode, punctuation, and audio tags like [whisper] or [excited].
Can I use ElevenLabs for free?
Yes. The free tier includes 10,000 monthly credits, enough to test the platform, though it doesn't include a commercial license. See my full ElevenLabs review for current paid-tier pricing, which sources don't fully agree on.
Can you provide a tutorial for using ElevenLabs text-to-speech?
Yes. The core workflow: I create an account, open Text to Speech, pick a voice and model suited to the specific use case, paste the script, generate a short test before scaling up, and adjust settings incrementally based on what I hear.
How do I add emphasis or emotion to an ElevenLabs voiceover?
On Eleven v3, bracketed audio tags in the script itself — [whisper], [excited], [pause 0.5s] — control delivery directly. See "Adding Emotion and Pauses with Audio Tags" above for how I use them.
How do I add more voices in ElevenLabs?
Beyond the default library, I clone a new voice from a short recording (instant cloning) or a longer, cleaner sample for a professional-grade clone — covered in the "Instant vs. Professional Voice Cloning" section above — or browse and add one of the 10,000+ community voices in the Voice Library.
How do I buy more credits on ElevenLabs if I run out?
Beyond upgrading your plan tier, ElevenLabs' own "Pay As You Go" top-up (in account settings) lets you buy additional credits once your monthly allowance runs out, on every self-serve tier including Free.
We may earn a commission when you buy through links on this page. It never changes what we recommend — we only write tutorials for tools we actually use, and the price you pay is the same either way. How we make money

ElevenLabs Review

Suno Review

Voicemod Review
Get the AI tool roundup.
One email a week: what launched, what is actually worth paying for, and what to cancel. No hype, no spam.
Unsubscribe any time. We never sell your email.