Skip to content
Journal / speech

Building voice agents in 2026: picking TTS and STT APIs that add up

How the TTS and STT loop works in a voice agent, what latency budgets to set, and what a conversational minute costs with real routed API prices.

Voice agents are two API categories glued together with latency anxiety. You transcribe what the user said (STT), think, then synthesize a reply (TTS), and the whole loop has to finish fast enough that the human on the other end does not start saying “hello?”. I route traffic in both categories, so I want to walk through the loop, the latency budget, and the part most posts skip: what a minute of conversation actually costs.

The loop

Every turn of a voice conversation is the same pipeline:

  1. Capture audio until the user stops talking (endpointing).
  2. Transcribe it (STT).
  3. Run your model or agent logic on the transcript.
  4. Synthesize the reply (TTS).
  5. Play it back, ideally streaming so playback starts before synthesis finishes.

Steps 2 and 4 are the tool calls. Everything else is your problem, but those two are the ones you pay per unit for, and the ones where provider choice moves both cost and feel.

The latency budget

In human conversation, a response gap much beyond a second starts to feel broken. That gives you a hard budget to divide across endpointing, STT, your model, and time-to-first-audio on TTS. I will not invent millisecond numbers for each provider here (the honest answer is that it varies by load, audio length, and region), but structurally the guidance is:

  • Stream everything you can. TTS providers that stream audio let you start playback while synthesis continues, which is the single biggest perceived-latency win.
  • Your LLM is usually the slowest step. Before optimizing STT vendors, measure how long your model takes to produce the first sentence.
  • Budget for failure. A provider timeout that triggers a retry blows the entire turn. This is a category where failover needs to be fast or the conversation dies; a router that tries a second provider automatically beats client-side retry logic you have to tune yourself.

What the categories cost

Routed prices from our catalog (provider list plus 20%; full tables at the TTS comparison and transcription comparison):

TTS, per 1k characters synthesized:

ProviderRouted priceQuality score
Deepgram Aura-2$0.03678
LMNT$0.0681
ElevenLabs$0.1293

STT, per audio minute:

ProviderRouted priceQuality score
Groq Whisper large-v3-turbo$0.0008480
Deepgram Nova-3$0.0051688
ElevenLabs Scribe$0.0080491

Note the shape of the two markets. TTS has a 3.3x spread and the expensive option is genuinely better in a way you can hear (I compared them in ElevenLabs vs LMNT vs Deepgram Aura). STT has nearly a 10x spread and the cheap option is a well-known open model, Whisper, served fast.

The cost of a conversational minute

Here is the worksheet. You need two assumptions, and I will state mine so you can swap in your own.

Assumption one: speech runs somewhere around 750 characters per minute (roughly 150 words). Assumption two: in a typical support-style conversation the agent speaks about half the time.

So one minute of conversation is about one minute of audio to transcribe (you generally transcribe the whole inbound channel) and about 375 characters to synthesize.

Budget stack (Groq Whisper plus Deepgram Aura-2):

  • STT: $0.00084
  • TTS: 0.375 x $0.036 = $0.0135
  • Total: roughly $0.014 per conversational minute, call it $0.86 per hour.

Premium stack (ElevenLabs Scribe plus ElevenLabs TTS):

  • STT: $0.00804
  • TTS: 0.375 x $0.12 = $0.045
  • Total: roughly $0.053 per conversational minute, or about $3.18 per hour.

Two things jump out. First, TTS dominates the bill in every configuration; synthesis is where the money goes, so that is the decision to sweat. Second, even the premium stack costs a few dollars per hour, which means for most products the LLM tokens in step 3, not the audio APIs, will be the biggest line item. Do this math for your own talk ratio before assuming the speech APIs are the problem.

Mixing tiers is usually right

Nothing forces you to buy STT and TTS from the same shelf. A combination I like structurally: cheap STT (Groq Whisper is scored 80 in our catalog, which is genuinely solid) paired with mid or premium TTS, because users hear your voice quality directly but never see your transcripts. Flip it for analytics-heavy products where transcript accuracy feeds downstream systems and the voice just needs to be pleasant.

One router-specific note: ElevenLabs TTS through our router is mp3 only, which is fine for playback but worth knowing if your pipeline expects raw PCM.

Start with the loop, not the vendor

Build the pipeline with any provider, instrument each step’s latency, then shop. The categories are normalized enough that swapping vendors should be a config change, not a rewrite; that is the whole argument for putting a routing layer under a voice agent, since a mid-conversation provider failure is the worst possible place to discover you have no plan B.

Full pricing for both categories, generated from the same catalog our router reads, is on the TTS and transcription comparison pages.