Text-to-speech is the category where our catalog’s price ladder maps most cleanly onto a real quality ladder. Three providers, three prices, three tiers: $0.036, $0.06, and $0.12 per thousand characters routed. That is a 3.3x spread from floor to ceiling, and unlike some categories where the premium is mostly branding, in TTS you can hear what you are paying for. Whether you should pay for it depends on who is listening.
The numbers, from our catalog. Routed prices are provider list plus 20%; list prices are $0.03, $0.05, and $0.10 respectively.
| Provider | Routed price per 1k chars | Quality score | Per 1M chars routed |
|---|---|---|---|
| Deepgram Aura-2 | $0.036 | 78 | $36 |
| LMNT | $0.06 | 81 | $60 |
| ElevenLabs | $0.12 | 93 | $120 |
A quick unit anchor, because characters are an awkward unit to feel: spoken English runs very roughly a thousand characters per minute of audio, give or take pace and punctuation. So the table above reads approximately as $0.036, $0.06, and $0.12 per spoken minute. (More on unit conversions in understanding per-unit pricing.)
1. Deepgram Aura-2: the utility tier
Aura-2 is the price floor at $0.036 per thousand characters routed, and Deepgram’s pitch is speed and cost for high-volume, real-time use. The voices are clear and professional: nobody will mistake them for a voice actor, and nobody will struggle to understand them either. We score it 78.
This is the tier for workloads where speech is a feature, not the product. Reading notifications aloud, voice responses in an internal tool, accessibility narration for UI text, agent status updates. In those contexts the listener wants information, not performance, and paying 3.3x more for beauty nobody asked for is just margin donation. At $36 per million characters, Aura-2 is also the only tier where truly high-volume TTS (think generating audio for every article on a large site) stays comfortable.
2. LMNT: the middle path
LMNT sits at $0.06 per thousand characters routed with a score of 81, and the middle tier is exactly what it sounds like: noticeably more natural than the utility tier, meaningfully cheaper than the premium one. LMNT has also focused on low-latency synthesis, which matters when the audio is part of a live exchange rather than a file you render once.
I think of LMNT as the conversational-agent tier. A voice agent that speaks with users all day cares about two things at once: latency (dead air kills conversations) and cost per minute (conversations are long). At roughly half the ElevenLabs price and a step up from Aura-2 in naturalness, the middle of the ladder is a genuinely rational place to land, not a compromise. Score-per-dollar, it is arguably the best value in the category.
3. ElevenLabs: the premium tier
ElevenLabs is the quality ceiling of our TTS catalog at 93, the highest score in the category by a wide margin, and the price ceiling at $0.12 per thousand characters routed. The voices are the most natural of the three: expressive, controllable, and consistent across long passages. One router-specific caveat: through our normalized API you get mp3 output only, so if your pipeline needs other formats from ElevenLabs specifically, that is a constraint to know upfront.
The premium tier makes sense when the audio is the product. Audiobooks, podcast production, video narration, character voices, anything a listener chose to spend time with. In those contexts, voice quality is conversion rate, and $120 per million characters against $36 is not the right comparison anyway; the comparison is against human narration, which costs orders of magnitude more. For a 60,000-word audiobook (roughly 350,000 characters), ElevenLabs costs about $42 routed. That is not a number that should scare anyone producing audio for sale.
Picking a tier without lying to yourself
The honest framework: identify who hears the audio and for how long.
- Machines-to-humans, briefly (alerts, confirmations, UI speech): Aura-2. Nobody reviews a notification’s prosody.
- Conversation (voice agents, interactive anything): LMNT, for the latency-cost-quality balance.
- Content (audiobooks, narration, media): ElevenLabs. Quality is the product; pay for the product.
Voice quality is partly subjective, and our scores (methodology at /docs/quality) are hand-curated judgments, not physics. Prices, on the other hand, are not subjective at all, and a 3.3x spread deserves a deliberate decision rather than a default. The pattern I would push back on hardest is reflexively buying the premium tier for utility speech: it is the single most common overspend I see in this category, precisely because the premium product is genuinely good and choosing it feels safe.
One more routing note: because all three providers take the same input (text, a voice selection) and return audio, TTS is one of the easier categories to keep provider-flexible. Saved routing preferences mean switching tiers later is a config change, and failover across providers covers you when one has a bad day.
Side-by-side numbers for the whole category, generated from the same open catalog our router reads, live on the text-to-speech comparison page.