Back to Blog
Voice AI2026 Guide

Best AI Voice Agent Platforms 2026

Voice agents went from demo to production line item in 18 months. But almost every platform advertises a per-minute price that isn't what you actually pay. We compare Vapi, Retell AI, LiveKit, ElevenLabs, and Bland on true all-in cost, real-world latency, self-hosting, and compliance — and tell you which to build on.

14 min read
5 platforms analyzed
Published August 2026

Quick Verdict

Vapi

The best developer experience and the widest model/voice choice. $0.05/min platform fee, and $0 in model costs if you bring your own API keys.

Best DXBYO keys
Retell AI

No platform fee, HIPAA included, warm transfer and compliance built in. The most predictable pick for regulated, operations-heavy call volume.

ComplianceOutbound
LiveKit

The only genuinely open-source, self-hostable option here. You own the media pipeline — and the ops burden that comes with it.

Self-hostOpen source
ElevenLabs

Unmatched voice quality and expressiveness, now with a full Conversational AI product. Pick it when the voice is the brand.

Voice qualityCloning
Bland

One flat, all-inclusive rate — roughly $0.09/min covering platform, voice, and LLM. The easiest number to put in a spreadsheet.

Flat rateSDR calls

TL;DR

Building a product? Start with Vapi. Regulated or ops-heavy? Retell. Need control, data residency, or video too? LiveKit. Voice quality above all? ElevenLabs. Want one flat number? Bland.

What you're actually buying (and why prices look fake)

A voice agent is a pipeline, not a product. Every call runs through four layers: telephony (getting the call in and out over the PSTN), speech-to-text (transcribing the caller), an LLM (deciding what to say), and text-to-speech (saying it). On top sits the orchestrator — the thing that manages turn-taking, interruptions, tool calls, and transfers.

This is why the sticker prices mislead. When Vapi says $0.05/min, that is the orchestrator only — STT, TTS, LLM, and telephony are billed separately or passed through at cost. When Bland says $0.09/min, that number is all-in. Comparing them directly is a category error, and it is the single most common way teams blow their voice budget.

The number that matters: total cost per minute, with your model and voice choices. A tuned Vapi stack on cheap models lands near $0.09–0.12/min; the same stack with a premium ElevenLabs voice and a frontier LLM can hit $0.30+/min. Price your configuration, not the platform.

The second number that matters: end-to-end latency — the gap between the caller finishing a sentence and hearing a reply. Under ~800ms feels conversational; past ~1.2s callers start talking over the agent. Note that a "sub-150ms" TTS claim is one stage of the pipeline, not the round trip.

Platform Overview

Five platforms, three very different philosophies: developer infrastructure, productized agent, and open-source media stack.

Best Developer Experience

Vapi

Developer-first voice orchestration

$0.05 / min

platform fee; models at cost

  • Swap any STT, LLM, and TTS provider
  • Bring your own API keys — $0 model markup
  • Cleanest SDKs and Twilio integration
  • Tool calling and function handoffs
  • SMS/chat billed at $0.005 per message
  • Telephony billed by your provider, not Vapi

Best for: Product teams building voice into their own app, who want component-level control

Best for Regulated

Retell AI

Productized agent with compliance built in

$0.07–0.31 / min

component-billed, no platform fee

  • No platform fee, no feature gating
  • Voice infrastructure at $0.055/min
  • HIPAA support included, PII removal add-on
  • Warm transfer and human handoff
  • Visual workflow builder for ops teams
  • $10 free credits, no contract to start

Best for: Healthcare, insurance, and outbound campaign teams that need compliance and handoff

Most Control

LiveKit

Open-source realtime media + Agents framework

Free self-hosted

Cloud from $0, Ship $50/mo

  • Agents framework and media server are open source
  • Free tier: 1,000 agent minutes, 5 concurrent
  • $0.01/min agent sessions past included minutes
  • Audio, video, and text in one pipeline
  • v1.5 adds adaptive interruption + dynamic endpointing
  • Self-host for full data residency control

Best for: Teams needing self-hosting, HIPAA/data residency, video, or multi-agent orchestration

ElevenLabs

Best-in-class voice, now a full agent product

$0.08–0.10 / min

Conversational AI 2.0

  • The most expressive TTS available in 2026
  • 11,000+ voices across 70+ languages
  • Voice cloning for branded agents
  • Sub-150ms model latency on TTS stage
  • Free tier through $990/mo plans
  • Also usable as the TTS layer inside Vapi or LiveKit

Best for: Consumer-facing and branded agents where voice quality is the differentiator

Bland

Flat all-inclusive per-minute pricing

~$0.09 / min

platform + voice + LLM

  • One number covers the whole pipeline
  • No component math, no surprise line items
  • Voice cloning and programmable call flows
  • Built for high-volume outbound
  • Averages ~800ms latency in 2026 tests
  • No drag-and-drop visual builder

Best for: SDR automation and follow-up calling where cost predictability beats flexibility

Also worth knowing

Pipecat (Daily) is the other open-source orchestration framework — free self-hosted, with Pipecat Cloud at roughly $0.01–0.03/min. Deepgram (Nova-3 STT) and Cartesia (Sonic-3, sub-100ms TTS) are components you drop into any of these stacks rather than platforms you pick between.

Synthflow and Telnyx target the no-code and carrier-grade ends respectively — useful if you want a visual builder or bundled PSTN, less so if you're writing code.

True Cost Breakdown (2026)

Advertised rate versus what a real production call costs. The "real-world" column assumes a typical stack: a mid-tier LLM, a standard voice, and US telephony.

Vapi

Advertised

$0.05/min platform

Real-world

~$0.12–0.33/min all-in

Models at cost, or $0 with your own API keys. Telephony billed by your provider. The spread depends almost entirely on which voice and model you pick.

Retell AI

Advertised

$0.07–0.31/min

Real-world

~$0.10–0.15/min typical

Voice infra $0.055 + TTS $0.015 (ElevenLabs $0.040) + LLM $0.003–0.16 + US telephony ~$0.015. Add-ons: knowledge base +$0.005, denoising +$0.005, PII removal +$0.01.

LiveKit

Advertised

Free / $0.01 per agent min

Real-world

Infra cost only if self-hosted

Cloud: Build free (1,000 agent min, 5 concurrent), Ship $50/mo (5,000 min, 20 concurrent), Scale $500/mo (50,000 min, up to 600). You still pay STT/LLM/TTS separately.

ElevenLabs

Advertised

$0.08–0.10/min

Real-world

$0.08–0.10/min + telephony

Conversational AI 2.0 bundles the voice layer. Plans run from a free tier up to $990/mo, with usage accumulating against your credit allowance.

Bland

Advertised

$0.09/min flat

Real-world

$0.09–0.14/min

The only genuinely all-inclusive rate here — platform, voice, and LLM in one number. Lowest cognitive overhead for forecasting.

Prices verified against provider pricing pages as of August 2026. Voice AI pricing changes frequently and varies by region, model, and volume commitment — always confirm current rates before committing.

Do the volume math early. At 10,000 minutes/month, the gap between $0.09/min and $0.25/min is $1,600/month — more than most teams' entire infrastructure bill. And a single upgrade from a standard voice to a premium cloned one can move you across that range.

Feature Comparison

Control, latency, and compliance across the five platforms.

Feature
Vapi
Retell
LiveKit
ElevenLabs
Bland

Control

Open source
Self-host
Swap LLM freely
limited
Swap TTS voice
own only
Bring your own keys
partial

Experience

Typical end-to-end latency
~800ms–1.2s
~580–800ms
depends on stack
~700ms
~800ms
Visual workflow builder
Video support
Voice cloning
via provider
via provider
via provider
Best-in-class

Operations

HIPAA support
enterprise
self-host
enterprise
enterprise
Warm transfer to human
Built in
build it
Flat all-in pricing
partial
Free tier to start
$10 credits
1,000 min

Deep Dive: Vapi

Vapi is the platform most developers should try first. It treats the voice pipeline as swappable components: pick your STT, your LLM, your TTS, wire in tool calls, and Vapi handles turn-taking, interruption, and telephony glue. The pricing model reflects that philosophy — you pay Vapi $0.05/min for orchestration and hosting, and model costs pass through at cost, or at $0 if you bring your own API keys.

The catch is that flexibility becomes your problem. Default pipelines commonly land in the 800ms–1.2s latency range, and getting to a genuinely snappy ~500ms means tuning endpointing, picking a fast model, and choosing a low-latency voice. Real all-in costs land anywhere between $0.12 and $0.33/min depending on those choices. Vapi also asks you to pre-purchase concurrency, which is a real planning cost if your volume is spiky.

Best for: Product teams embedding voice into their own app who want component-level control. Skip if: you want one number on an invoice or a no-code builder for a non-technical ops team.

Deep Dive: Retell AI

Retell is the pick when the agent is a business process rather than a product feature. It charges no platform fee and bills components transparently: $0.055/min voice infrastructure, $0.015/min standard TTS (or $0.040 for ElevenLabs voices), $0.003–$0.16/min for the LLM, and roughly $0.015/min for US telephony. Add-ons are explicit line items — knowledge base +$0.005, advanced denoising +$0.005, PII removal +$0.01.

What sets it apart operationally is the stuff you'd otherwise build yourself: HIPAA support included, warm transfer to a human, and a visual workflow builder that a non-engineer can maintain. Independent 2026 tests put it at a steadier 580–800ms than Vapi's default pipeline, which matters more for perceived quality than any benchmark number.

Best for: Healthcare, insurance, and outbound campaign teams needing compliance and human handoff. Skip if: you want to self-host or you need deep custom control over the media layer.

Deep Dive: LiveKit

LiveKit is the only platform here you can genuinely run yourself. Both the media server and the Agents framework are open source, so self-hosting costs you infrastructure and nothing else — which is why it dominates in HIPAA, data-residency, and EU-sovereignty conversations. It's also the only one that handles audio, video, and text in one pipeline, and the natural choice for multi-agent orchestration.

LiveKit Cloud removes the ops work: Build is free with 1,000 agent minutes and 5 concurrent sessions, Ship starts at $50/mo (5,000 minutes, 20 concurrent), and Scale at $500/mo (50,000 minutes, up to 600 concurrent). Past included minutes, agent sessions run $0.01/min and US local telephony $0.01/min. Version 1.5 shipped adaptive interruption handling and dynamic endpointing — meaningful, because turn-taking quality is the hardest thing to build yourself.

Best for: Teams that need self-hosting, data residency, video, or multi-agent orchestration. Skip if: you don't have deployment experience — you're assembling a pipeline, not buying one.

Deep Dive: ElevenLabs

ElevenLabs won the voice-quality race and then built a platform on top of it. Conversational AI 2.0 runs roughly $0.08–0.10/min and bundles the voice layer with turn-taking and agent logic. With 11,000+ voices across 70+ languages and best-in-class cloning, it's the obvious pick when the voice is a brand asset rather than a utility.

Two caveats. First, the agent product is newer than the TTS engine — orchestration features lag Vapi and Retell. Second, its advertised sub-150ms latency is the model stage; the full round trip lands nearer 700ms in practice. Many teams get the best of both worlds by using ElevenLabs purely as the TTS layer inside a Vapi or LiveKit pipeline.

Best for: Consumer-facing and branded agents where voice quality wins deals. Skip if: you need heavy custom orchestration or the cheapest possible per-minute rate.

Deep Dive: Bland

Bland's pitch is refreshingly boring: one flat rate, around $0.09/min, covering platform, voice, and LLM. No component math, no pass-through surprises. For a sales team modelling 50,000 outbound minutes a month, that predictability is worth more than a theoretical 20% saving on a stack someone has to tune.

The trade-off is control. You get voice cloning and programmable call flows, but not the free choice of model and voice that Vapi or LiveKit give you, and there's no drag-and-drop builder for ops teams. Latency averages around 800ms — the slowest of the group in 2026 testing, though still inside conversational range.

Best for: High-volume outbound SDR and follow-up calling with predictable unit economics. Skip if: you need component-level control or the lowest possible latency.

Decision Guide

You're building voice into your product

Use Vapi. The best SDKs, free choice of every pipeline component, and $0 model markup with your own keys. Budget a week to tune latency before you judge it.

Vapi

You're in a regulated industry

Use Retell AI. HIPAA included, PII removal as a toggle, warm transfer built in, and no platform fee. If you must own the data end to end, self-hosted LiveKit instead.

Retell AI

You need self-hosting, video, or multi-agent

Use LiveKit. Open source top to bottom, one pipeline for audio and video, and a free Cloud tier to prototype before you commit to running it yourself.

LiveKit

Voice quality is the product

Use ElevenLabs — either as the full agent platform, or as the TTS layer inside a Vapi or LiveKit pipeline if you need stronger orchestration.

ElevenLabs

You need one predictable number

Use Bland. A flat ~$0.09/min all-in makes unit economics trivial to model at high outbound volume — worth trading some control for.

Bland

Building a voice AI product? Get the whole stack

Your voice platform is one piece. Use our AI-powered generator to build a complete stack — database, hosting, auth, LLM provider, and more — tailored to your budget and team size.

Frequently Asked Questions

What does an AI voice agent actually cost per minute?

Realistically $0.09 to $0.30 per minute all-in, depending on your model and voice choices. Bland is flat at roughly $0.09; Retell typically lands at $0.10–0.15; a Vapi stack ranges from $0.12 to $0.33. Advertised platform fees (like Vapi's $0.05/min) exclude STT, LLM, TTS, and telephony.

Vapi vs Retell: which should I choose?

Choose Vapi if you're a developer embedding voice into a product and want to swap every pipeline component. Choose Retell if you're running a business process — it has no platform fee, includes HIPAA, ships warm transfer and a visual builder, and holds steadier latency out of the box.

Can I self-host an AI voice agent?

Yes — LiveKit and Pipecat are both open source and free to self-host, leaving you paying only infrastructure plus your STT/LLM/TTS providers. Vapi, Retell, ElevenLabs, and Bland are all vendor-hosted. Self-hosting is the standard answer for strict data-residency or sovereignty requirements.

What latency do I need for a natural conversation?

Aim for under 800ms end to end. Past roughly 1.2 seconds, callers assume the agent didn't hear them and start talking over it. Be careful reading vendor claims: a "sub-150ms" number usually refers to the TTS model alone, not the full transcribe → think → speak round trip.

Which platform is best for HIPAA compliance?

Retell AI includes HIPAA support on its standard pricing and offers a PII removal add-on, making it the fastest path to a compliant deployment. For full control over where data lives, self-hosted LiveKit is the stronger answer. Vapi, ElevenLabs, and Bland handle HIPAA through enterprise agreements.

Do I need a separate TTS provider like ElevenLabs or Cartesia?

Not necessarily — every platform here ships usable default voices. You add a dedicated provider when voice quality is a differentiator (ElevenLabs) or when you're chasing the lowest possible latency (Cartesia's Sonic-3 claims sub-100ms at the model stage). Expect premium voices to add roughly $0.025/min over standard ones.

Related Articles