Vapi vs Retell AI: The Short Answer
Pick Vapi if you want the cheapest all-in cost and maximum control over which TTS, STT and LLM providers you use — it charges a $0.05/min orchestration fee and passes provider costs through at cost. Pick Retell AI if you want predictable, single-line billing and a curated stack that just works — it charges $0.07–$0.31/min depending on the components you turn on. Both are strong; the decision is really unbundled-and-flexible (Vapi) vs bundled-and-simple (Retell).
The catch almost every comparison misses: these two platforms don't price the same way, so the headline per-minute numbers aren't comparable. Below is a component-by-component breakdown built from each company's current pricing page (checked September 2026), plus modeled all-in costs at 1,000, 10,000 and 50,000 minutes a month so you can see where each one actually wins.
Winner by Category
| Category | Winner | Why |
|---|---|---|
| Lowest all-in cost | Vapi | At-cost provider pass-through; ~$0.08–$0.13/min typical |
| Simplest billing | Retell AI | One per-minute rate, components you toggle on |
| Provider flexibility | Vapi | Bring your own TTS/STT/LLM keys |
| Fastest to ship | Retell AI | Curated defaults, less configuration |
| Lowest latency TTS | Tie | Both can run Cartesia (~40ms TTFA) |
| Predictable at scale | Retell AI | No compounding line items to model |
The Real Difference: Two Pricing Architectures
Vapi and Retell solve the same problem — orchestrating a real-time voice pipeline (speech-to-text → LLM → text-to-speech → telephony) — but they bill for it in opposite ways.
Vapi is unbundled. You pay a flat $0.05/min orchestration fee, and every provider you plug in bills at its own rate with no markup. Vapi's pricing page lists the pass-through ranges directly: Deepgram transcription at $0.0095–$0.0099/min, an OpenAI LLM at $0.0077–$0.0452/min, and ElevenLabs voice at $0.0146–$0.0238/min (verified September 2026). Internal transport is free; external telephony (Twilio, Vonage, Telnyx) is extra.
Retell is bundled. Its headline rate is $0.07–$0.31/min for AI voice agents, and that range is the sum of components you turn on: $0.055/min for the voice infrastructure, plus text-to-speech ($0.015/min for its standard voices — Minimax, Fish, Cartesia, OpenAI, Inworld — or $0.040/min for ElevenLabs), plus your LLM ($0.0016–$0.064/min depending on model), plus US telephony at $0.015/min (verified September 2026). Add-ons like a knowledge base (+$0.005/min) or PII removal (+$0.01/min) stack on top.
Why this matters more than the sticker price
Vapi's $0.05/min looks 30% cheaper than Retell's $0.07 floor, but that's not a like-for-like number — Vapi's $0.05 is orchestration only, while Retell's $0.07 already includes a voice engine. Once you add the same TTS, STT and LLM to both, they land within a few cents of each other. The decision isn't "which is cheaper per minute" — it's whether you'd rather manage provider keys and get at-cost rates (Vapi) or accept a small bundling premium for one predictable invoice (Retell).
Modeled All-In Cost at 1K / 10K / 50K Minutes
These figures model a typical inbound agent: standard-quality TTS, a mid-tier LLM, and US telephony. They're built from each platform's published component rates (checked September 2026), not from a billed account — your real numbers move with LLM choice, call length and volume discounts.
| Volume / month | Vapi | Retell AI |
|---|---|---|
| 1,000 min | ~$82–$129 (Vapi's own example) | ~$110–$150 |
| 10,000 min | ~$800–$1,300 | ~$1,100–$1,500 |
| 50,000 min | ~$4,000–$6,500 (or Pro package, $999/mo min) | Negotiate enterprise pricing |
Vapi's per-minute stays flat because the orchestration fee doesn't scale down on the usage plan — but its $999/mo Pro package drops the hosting fee to 10%, which pays for itself above roughly 20,000 min/mo. Retell publishes a $0.07–$0.31/min range and moves large accounts to custom pricing. Model your exact stack with our TTS cost calculator.
TTS Provider Flexibility
This is Vapi's clearest advantage. Vapi is designed around bring-your-own-keys, so you can wire in essentially any TTS provider and pay that provider's rate directly. That matters because the voice layer is where cost and latency diverge most: running Cartesia (~40ms time-to-first-audio) versus ElevenLabs (~75ms, higher quality) is a real UX-and-price trade-off you control.
Retell also supports multiple voices, but as a curated menu with fixed line-item pricing: most of its TTS options (Minimax, Fish, Cartesia, OpenAI, Inworld and its own platform voices) are billed at a flat $0.015/min, while ElevenLabs voices cost $0.040/min. That's simpler to reason about, but you're choosing from Retell's list at Retell's rates rather than plugging in your own account. If you want the cheapest possible voice, note that Retell's $0.015/min covers Cartesia — the same low-latency engine you'd otherwise pay Vapi to orchestrate. For a deeper look at which engine fits a live agent, see our best TTS for voice agents guide.
Latency: It's About the Provider, Not the Platform
For a phone agent, total round-trip latency under about 300ms is what keeps a conversation feeling natural, and the biggest single lever is your TTS provider's time-to-first-audio. Because both Vapi and Retell can run Cartesia, the floor is roughly the same on either platform — neither has a structural latency moat over the other. Where they differ is the orchestration overhead and how aggressively each streams tokens into the TTS engine, which is hard to compare from published specs and worth prototyping on your own call patterns before committing.
Developer Experience & Scale
Retell exposes clear operational controls in its pricing: concurrency beyond 20 simultaneous calls costs $8 per concurrency slot per month, extra knowledge bases are $8/month each, and verified phone numbers are $10/month. Those are the knobs you'll tune as an inbound support or scheduling agent scales, and having them itemized makes capacity planning straightforward.
Vapi trades that simplicity for control. You're assembling the stack, which means more integration work up front and more moving parts to monitor, but also the freedom to swap any component when a cheaper or better provider appears. Its Success Packages — free to start, $29/month Core, $999/month Pro, and custom Premier — gate support level and the hosting-fee discount rather than the core feature set.
Which to Choose, by Use Case
- Cost-sensitive, engineering-heavy team: Vapi. At-cost pass-through plus your own provider negotiations gets you the lowest all-in rate, and you already have the team to manage the keys.
- Ship-fast product team: Retell. One invoice, curated defaults and itemized scaling knobs beat the config overhead of assembling a stack when speed matters more than shaving cents.
- Latency-critical phone agent: Either — run Cartesia on whichever platform you prefer for the rest of the experience.
- High-volume outbound only: Look at Bland AI too (below) — its flat bundled rate is built for that exact case.
The Third Option: Bland AI
If your workload is high-volume outbound, Bland AI is worth a look. Its Start plan is $0.14/min with no monthly platform fee (and includes 2 credits plus an inbound number worth $15/month); the Build plan drops the rate to $0.12/min for a $299/month platform fee (verified September 2026). Everything — LLM, speech-to-text and text-to-speech — is bundled into the per-minute rate with no separate token charges, which makes it the most predictable of the three. The trade-off is less provider flexibility than Vapi and a higher per-minute floor than a lean Retell or Vapi build. We break all of this down further in our voice agent pricing guide.
Related Guides
By TextToLab Research Team · Researched comparison (not a hands-on billing test). Pricing verified against the live Vapi (vapi.ai/pricing), Retell AI (retellai.com/pricing) and Bland AI (bland.ai/pricing) pricing pages on September 25, 2026. All-in figures are modeled from each platform's published component rates and will vary with your LLM choice, call length and volume. Provider-level TTS and latency detail from our individual pricing deep-dives (linked above).