Text-to-speech is the category where I most often watch people pay the premium price for a workload that would never notice the difference, and occasionally the reverse, which is worse. The three providers carrying our TTS traffic span a clean 3.3x price range, and unlike some categories, the spread here maps onto something real. The catch is that the something is voice quality, which is the most subjective metric in the entire catalog. Price, on the other hand, is not subjective at all, so let us start there.
The spread, in dollars
Routed prices are provider list plus 20%, per the pricing docs. Quality scores are our hand-curated 0-100 ratings.
| Provider | Model | Routed per 1k characters | Quality score | Router note |
|---|---|---|---|---|
| Deepgram | Aura-2 | $0.036 | 78 | |
| LMNT | $0.06 | 81 | ||
| ElevenLabs | $0.12 | 93 | mp3 output only via router |
Characters are the billing unit, so translate to content. A thousand words of English runs roughly five to six thousand characters. Call it a long blog post or ten minutes of speech: about $0.20 through Aura-2, $0.33 through LMNT, $0.66 through ElevenLabs. An audiobook-length project at 500,000 characters costs $18, $30, or $60. For a notification system speaking short alerts all day, multiply tiny numbers by huge volumes and the 3.3x becomes the whole story.
Deepgram Aura-2: the volume tier
At $0.036 per thousand characters routed, Aura-2 is the floor of our TTS catalog, and Deepgram positions it (per their docs) squarely at real-time agent use cases: low-latency synthesis for voice bots, IVR, and interactive systems. The 78 we score it says the voices are clean and professional rather than remarkable. For utilitarian speech, meaning the listener wants the information and not the performance, that is not a criticism. Status alerts, phone-tree prompts, screen-reader-style narration of dynamic content: nobody replays those for the timbre.
My rule: if the audio is disposable, synthesize it at the floor.
LMNT: the middle that earns its slot
LMNT at $0.06 routed scores 81, and its pitch (again per the docs) leans hard on speed and stability for real-time conversation. Qualitatively, I would describe its voices as a step warmer than the utility tier, and the API as pleasantly focused: it does synthesis, quickly, without a sprawling feature surface. The honest question for any middle option is whether workloads actually land there, or whether everything polarizes to the cheap and premium ends. In my experience LMNT wins precisely when both neighbors half-fit: the floor sounds too robotic for your product’s voice, and the ceiling costs double for polish your users will not notice in a two-second reply.
ElevenLabs: the ceiling, priced like it
ElevenLabs is the highest quality score in our speech catalog at 93, and the only TTS provider whose output people recognize by reputation. The voice quality is the closest thing this subjective category has to a consensus: expressive, natural pacing, emotional range that the cheaper tiers visibly lack on long-form content. At $0.12 per thousand characters routed you are paying 3.3x the floor for it.
One router-specific caveat you should know before building on it through us: via route.tools, ElevenLabs output is mp3 only. If your pipeline needs raw PCM or another container, that constraint matters more than any quality score, and you should factor it in before the price discussion even starts.
Deciding without a benchmark
Voice quality resists benchmarking; our scores here are explicitly editorial judgment (methodology at /docs/quality), and yours should be too. The practical procedure: take one paragraph of your actual content, synthesize it through all three, and listen on the device your users will use. Phone speaker and studio headphones disagree about these voices. Ten minutes of listening beats any comparison table, including mine.
Then let the workload sort itself by exposure. Speech that represents your brand for minutes at a time (narration, content, characters) justifies the ceiling. Speech inside a fast conversational loop justifies the middle. Speech that exists to convey a fact and be forgotten belongs at the floor. The pattern rhymes with what I wrote in the TTS listicle: this category is tiers, not rankings.
The routing angle is the same one I always end on, because it keeps being true: these do not have to be one decision. Alerts can route cheapest-first while your narration pipeline pins quality-sorted, from the same key, via saved routing preferences. Failover matters here too, since a voice feature that goes silent during a provider outage is very noticeable, and a 78-quality fallback beats no audio at all.
Current pricing and scores for every TTS provider we route are side by side on the text-to-speech comparison page.