All posts

12 Best AI Voice Agents to Try in 2026

Retell AI was the fastest platform in one independent voice-agent benchmark and the slowest in another. The difference comes down to where each test places the stopwatch. This comparison separates measured latency from architecture, pricing and the work each product leaves to your team.

21 min read
On this page

12 Best AI Voice Agents to Try in 2026

Retell AI was the fastest platform in one independent voice-agent benchmark and the slowest in another. The disagreement is useful. It shows why a vendor latency claim cannot tell you how long a caller will wait.

Tested Media ran 500 production calls per platform in March 2026. It measured from the end of caller speech to the start of agent speech and put Retell first, with a 680 ms median and 920 ms p95.

Openbenchmarks used a robot caller on a real phone line, reading a fixed script. Its researchers analysed independent dual-channel recordings with Silero VAD and pooled 2,078 usable turns. On that test, Retell finished last among the five platforms, with a 1,740 ms median and 2,259 ms p95.

The stopwatch explains much of the disagreement. Openbenchmarks' methodology distinguishes a platform reading, which may stop when synthesised audio leaves the server, from time to first audible byte in an independent recording. The latter includes endpointing delay, network travel on both legs and telephony overhead. Its comparison found vendor readings roughly 490 ms lower than the corresponding call audio. A platform recording could also begin roughly 550 ms before the same audio appeared in an independent recording.

As the researchers put it, “A platform measures from where it stands, and the caller is not standing there.”

That gap is the basis for this field guide. Treat a vendor number as a component measurement. For planning, add roughly 500 ms and ask for end-to-end p95 on the carrier route you will use.

How I evaluated these AI voice agents

The products below do not all solve the same problem. Retell runs a managed voice stack. Deepgram exposes an agent API. Pipecat is a Python framework. PolyAI sells an enterprise contact-centre programme. Putting them into a single league table would imply a common job and score that do not exist.

I grouped them by the layer a buyer is choosing. Within each group, the useful questions are different.

For managed platforms, latency and the full production bill matter together. A low hosting rate means little if telephony, model use, compliance or call capacity changes the total. For model APIs and frameworks, ownership matters more: who handles transport, provider selection, state, retries and monitoring? That is the same orchestration question any agent stack raises, with a phone line attached. Enterprise products need to be judged on integration and governance because public rate cards are generally unavailable.

The latency evidence also has boundaries. Tested Media covered Retell, Vapi, Bland and Synthflow. Openbenchmarks covered Telnyx, ElevenLabs, Bland, Vapi and Retell. A missing independent result remains missing; a vendor claim is not a substitute.

Independent latency results

These figures are separated from the product field guide because only five of the twelve products have a result in either independent benchmark. Compare down a benchmark column. Do not compare a Tested Media result directly with an Openbenchmarks result as if the clocks began and ended at the same points.

PlatformTested Media median / p95Openbenchmarks median / p95What the result exposes
Retell AI680 / 920 ms1,740 / 2,259 msThe benchmark method can reverse the apparent winner
Vapi720 / 1,050 ms1,558 / 2,008 msConfiguration and measurement boundary both matter
Bland AI850 / 1,180 ms1,520 / 2,248 msIts 1.48× Openbenchmarks tail was the widest
Synthflow920 / 1,250 msNo independent resultTested Media's slowest median in its set
ElevenLabs AgentsNo independent result1,424 / 1,768 msIts 1.24× tail was the steadiest

Openbenchmarks also measured Telnyx at a 1,296 ms median and 1,856 ms p95. Telnyx is an honourable mention rather than one of the twelve, but its 1.43× tail is a useful reminder that the fastest median does not guarantee the tightest distribution.

Twelve options by architectural layer

This table answers a different question from the benchmark table: what are you buying, and what remains your responsibility?

LayerProductBest fitWhat your team still operatesPricing evidence
Managed platformRetell AIProduction phone agentsWorkflow tuning and call reviewRetell pricing: $0.07–$0.31/min
Managed platformVapiConfigurable developer buildsProvider choices and failure handlingVapi pricing: $0.05/min hosting, providers extra
Managed platformBland AIOutbound campaignsCampaign behaviour and tail testingTested Media: $0.09–$0.24/min
Managed platformSynthflowNo-code call flowsWorkflow limits and escalation designTested Media: $0.13–$0.20/min
Model/APIElevenLabs AgentsVoice quality with consistent measured turnsTools and escalationNot verified in the research
Model/APIOpenAI Realtime APINative speech-to-speechThe surrounding agent productNot verified in the research
Model/APIDeepgram Voice Agent APIConsolidated agent connectionTelephony behaviour and monitoringNot verified in the research
Model/APIAssemblyAI Voice Agent APIFlat API meter with framework pluginsTelephony and orchestrationAssemblyAI roundup: $4.50/hour
FrameworkLiveKit AgentsTransport with pluggable modelsProvider stack and deploymentProviders charged separately
FrameworkPipecatVendor-neutral Python buildsDeployment and state handlingProviders charged separately
EnterprisePolyAIVoice-focused contact centresProcurement and programme governanceCustom quote
EnterpriseParloaMultilingual regulated operationsProcurement and programme governanceCustom quote

Managed voice-agent platforms

Managed products are the direct shortlist for teams that want phone automation without running media infrastructure. The platform operates the voice stack while the buyer designs the agent and connects business systems. The difficult comparison is the all-in production cost, not the smallest number on a pricing page.

Retell AI: the benchmark contradiction in one platform

Retell AI homepage

Retell deserves close attention because it demonstrates how benchmark boundaries change a result. Tested Media put it first at a 680 ms median. Openbenchmarks put it last at 1,740 ms. Those measurements do not cancel each other out. One shows performance under Tested Media's production-call method; the other captures audio over a real phone route in an independent recording.

For a buyer, the right response is to reproduce the call path that matters. Ask Retell for p95 rather than an average, then measure the same prompt over your carrier. Keep endpointing settings fixed. A polished demo can hide slow turns in the tail, and an internal platform recording can begin before the caller hears audio.

Retell is still a strong managed-platform shortlist choice. It lets a team deploy phone agents without selecting and wiring every audio component. It also makes its component pricing unusually visible. Retell's pricing page lists voice infrastructure at $0.055 per minute, platform voices at $0.015 and US Twilio telephony at $0.015. LLM use ranges from $0.003 to $0.16 per minute.

Production controls add to that bill. Knowledge-base access, advanced denoising and safety guardrails each cost $0.005 per minute. PII removal adds $0.01. AI QA costs $0.10 per minute after the first 100 minutes. The first 20 concurrent calls are included; further capacity costs $8 per concurrent line each month.

The arithmetic matters more than the advertised floor. A deployment using voice infrastructure, a platform voice, US Twilio telephony, a modest model and the production controls can reach roughly $0.20 per minute or more. Call review and workflow maintenance also stay with the buyer. Retell operates the stack, but it does not decide whether a retry is safe or whether an escalation passed enough context to a human.

Retell is best treated as a production managed platform with transparent component costs. It has no uncontested latency crown. Its value is the combination of managed infrastructure, tuning room and a bill that can be modelled before traffic grows.

Pros

  • Fastest Tested Media median, with a 920 ms p95.
  • Transparent component pricing makes the production bill easier to model.
  • Managed infrastructure avoids assembling every audio component.
  • Initial concurrency is included before paid capacity begins.

Cons

  • Slowest Openbenchmarks median among the platforms it compared.
  • Production controls can push the floor price to roughly $0.20 per minute or more.
  • Workflow tuning and call review still belong to the buyer.

Vapi: managed infrastructure with provider choice

Vapi homepage

Vapi gives developers more control over the provider stack than the other managed products here. Teams can choose speech recognition, model and voice providers, or bring their own API keys, while Vapi runs the core voice infrastructure. This makes Vapi useful when a team wants to tune each component without taking responsibility for the entire media layer.

That control affects latency. Tested Media measured a 720 ms median and 1,050 ms p95, while noting that Vapi can produce lower results with the right configuration and that defaults are slower. Openbenchmarks measured a 1,558 ms median and 2,008 ms p95 over its phone-line route. A buyer should therefore evaluate a named configuration, not “Vapi” in the abstract. Changing endpointing or a provider changes the system being measured.

Vapi pricing starts with a $0.05 per-minute hosting charge. Speech recognition, model and speech generation usage pass through at cost. Bringing your own keys can remove those provider charges from Vapi's invoice, but the providers still charge for their services.

The Build tier includes more than 60 minutes and 10 concurrent lines. Each additional line costs $10 per month. Compliance can dominate the platform bill before call volume becomes large: the HIPAA add-on costs $2,000 per month, and Zero Data Retention costs $1,000 per month.

Provider choice also creates operational work. Someone has to decide which combination meets the latency, voice and data-handling requirements. Observability must cross vendor boundaries, and an incident may require separating a carrier delay from endpointing behaviour or a slow model response. Managed infrastructure narrows that problem; it does not erase it.

Vapi is the better fit when provider choice and bring-your-own keys are requirements rather than nice extras. Teams that want a fixed stack with fewer decisions may prefer Retell or a no-code product. Teams that choose Vapi should budget for the full provider chain and test the default configuration before tuning it.

Pros

  • Bring-your-own keys and provider choice support a configurable stack.
  • The hosting charge is separated from provider usage.
  • The Build tier includes 10 concurrent lines.
  • Tested Media documents the effect of configuration on latency.

Cons

  • Default settings can be slower than a tuned configuration.
  • HIPAA and Zero Data Retention add substantial fixed monthly costs.
  • Observability has to cross the providers selected by the team.

Bland AI: outbound volume with a wide latency tail

Bland AI homepage

Bland is aimed at outbound campaigns. The sanctioned research could not verify its often-quoted $0.09 all-in rate from a first-party pricing page; Tested Media gives a broader $0.09–$0.24 per-minute range.

Its latency tail deserves more attention than its median. Openbenchmarks measured a 1,520 ms median and 2,248 ms p95. The resulting 1.48× tail ratio was the widest in that benchmark. For an outbound campaign, test p95 on the actual routes and listen to slow turns before choosing on price.

Pros

  • Designed for outbound campaigns at volume.
  • Measured by both independent benchmarks used in this article.

Cons

  • Openbenchmarks found the widest latency tail in its comparison.
  • The research supports a price range, not a verified flat all-in rate.

Synthflow: a shorter no-code route

Synthflow homepage

Synthflow suits bounded appointment-setting and receptionist flows where a visual builder is more useful than a configurable provider stack. Tested Media measured a 920 ms median and 1,250 ms p95, the slowest result in its four-platform set. The same source lists a $0.13–$0.20 per-minute range, which also gives Synthflow the highest floor in that comparison.

The trade-off appears when the workflow needs unusual state logic or deep debugging. A no-code route can reduce initial engineering, but teams should test escalation and corrections before committing to it.

Pros

  • No-code tooling suits bounded receptionist and appointment-setting flows.
  • Tested Media publishes both median and p95 results.

Cons

  • Slowest median in Tested Media's managed-platform set.
  • No Openbenchmarks phone-line result is available.

Model and API layer

These services expose the real-time model or agent connection. They give product teams architectural control while leaving more of the phone agent around that connection to the buyer.

ElevenLabs Agents: the strongest measured consistency

ElevenLabs homepage

ElevenLabs is usually shortlisted for voice quality. The more defensible reason to include it here is consistency in an independent phone-line benchmark. Openbenchmarks recorded a 1,424 ms median and 1,768 ms p95 across 429 usable turns from 432 attempts. Its 1.24× tail ratio was the tightest in that comparison, and its p95 was lower than every other platform in the table.

That does not make it the fastest by every measure. Telnyx had the lower Openbenchmarks median, and Tested Media did not include ElevenLabs. What the published evidence supports is narrower: among the products Openbenchmarks compared, ElevenLabs kept slow turns closer to its typical turn.

Consistency matters because callers experience individual turns, not an aggregate average. A system with a respectable median and a stretched p95 will periodically leave a long silence. Those delays can be especially damaging around confirmations, tool calls and transfers, where the caller may assume the line has failed and start speaking again.

Voice quality also cannot compensate for a weak workflow. The product still needs safe tool calls, state handling and escalation rules. A warm handoff should carry the reason for the call and the actions already attempted. If a tool times out after completing an action, the retry path needs an idempotency check so it does not book, cancel or charge twice.

The research pack did not provide comparable, verified agent pricing for ElevenLabs, so this article does not manufacture a per-minute figure. That evidence gap belongs in a proof of concept. Price the complete route and compare it using the same prompt, carrier and tools as the managed alternatives.

ElevenLabs is the most interesting model/API-layer choice here when natural voice and stable turn timing matter. Its independent result is stronger than a vendor latency claim, while its operational boundary remains different from Retell or Vapi. The buyer still has to build and run more of the agent around it.

Pros

  • Steadiest Openbenchmarks tail ratio at 1.24×.
  • Lowest p95 in that independent comparison.
  • Strong fit when voice quality is central to the caller experience.
  • The benchmark publishes usable-turn counts as well as latency.

Cons

  • No comparable agent price was verified in the research pack.
  • Tested Media did not include it.
  • Tool safety, state and escalation remain product responsibilities.

OpenAI Realtime API: native speech-to-speech

OpenAI Realtime API documentation

OpenAI Realtime API processes speech without exposing a separate STT-to-LLM-to-TTS chain. That can preserve vocal cues and simplify the visible audio path. It also removes some of the clean stage boundaries available in a chained system.

Neither benchmark in the research pack measured it, and the pack does not provide a verified price. Choose it for a native speech-to-speech architecture, then validate end-to-end phone behaviour independently.

Pros

  • Native speech-to-speech removes the visible STT and TTS split.
  • Direct audio processing can preserve vocal cues.

Cons

  • Neither independent benchmark published a latency result.
  • The research pack contains no verified price.

Deepgram Voice Agent API: a consolidated agent connection

Deepgram homepage

Deepgram puts the agent path behind a single socket, reducing the real-time connections a developer must wire together. A consolidated API does not give the agent its tools or operate the phone agent for you. Your team still has to handle carrier behaviour and monitor tool calls.

The research pack contains no independent latency result or verified price for this API. Its case rests on architecture and developer fit, not an unsupported performance comparison.

Pros

  • A single socket reduces the real-time connections developers must wire together.
  • The surrounding product logic stays under the team's control.

Cons

  • No independent latency result was published in either benchmark.
  • Telephony behaviour and monitoring remain with the buyer.

AssemblyAI Voice Agent API: a flat API meter with framework plugins

AssemblyAI homepage

AssemblyAI's roundup describes its Voice Agent API as one WebSocket at a flat $4.50 per hour, with no concurrency commitment, plus plugins for LiveKit and Pipecat. That meter covers the API layer rather than the all-in cost of a phone agent. Telephony and the work around orchestration remain outside it. Neither independent benchmark measured the product.

Pros

  • Flat hourly API pricing is simpler to model than a pile of component rates.
  • The stated rate has no concurrency commitment.
  • Plugins connect it to LiveKit and Pipecat.

Cons

  • Telephony and orchestration sit outside the hourly figure.
  • Neither independent benchmark published a latency result.

Orchestration frameworks

A framework suits a team that wants to select providers and own deployment behaviour. It is a different purchase from a managed calling product, which is why these options are not ranked beneath one.

LiveKit Agents: transport and orchestration

LiveKit Agents homepage

LiveKit Agents combines transport with orchestration while leaving speech recognition, speech generation and the model pluggable. It fits teams making voice part of a broader product. Every selected provider brings its own bill and operational boundary, and neither benchmark supplies an end-to-end latency result for the framework.

AssemblyAI's roundup and the sanctioned production overview mention an agent-session rate around $0.01 per minute for LiveKit or Pipecat, plus providers, but do not provide product-specific first-party evidence. That figure is therefore omitted from the comparison table rather than presented as verified pricing.

Pros

  • Transport and orchestration live in the same framework.
  • Model providers remain pluggable.

Cons

  • Each provider adds a bill and an operational boundary.
  • Neither benchmark supplies an independent end-to-end result.

Pipecat: vendor-neutral Python orchestration

Pipecat homepage

Pipecat is a vendor-neutral Python framework built by the Daily team. It supports more than 60 integrations and leaves the pipeline in the developer's hands. The Deepgram roundup is the sanctioned source for that architecture context, but the research pack provides no independent end-to-end latency result.

Pipecat leaves deployment, state and interruption handling to your team. That is useful when provider portability matters and the engineering team wants direct control. It is excessive when the job is a narrow phone workflow that a managed platform already operates.

Pros

  • Vendor-neutral Python keeps orchestration separate from one model provider.
  • More than 60 integrations support a configurable pipeline.
  • The team controls deployment behaviour directly.

Cons

  • The team owns state, interruption handling and deployment.
  • No independent end-to-end benchmark result is available.

Enterprise contact-centre platforms

These products are bought as programmes rather than swipe-a-card APIs. Integration depth and governance carry more weight than a public per-minute rate. Both use custom pricing in the sanctioned research.

PolyAI: voice-specialised enterprise CX

PolyAI homepage

PolyAI focuses on voice automation for large contact centres, where containment and integration with service infrastructure shape the purchase. It has no public rate card or independent latency result in the research pack.

Expect a sales cycle and implementation programme. A small team that needs a focused receptionist flow should shortlist a self-serve managed product instead.

Pros

  • Voice-specialised positioning suits large contact-centre programmes.
  • Enterprise integration is central to the product's fit.

Cons

  • No public rate card or independent latency result is available.

Parloa: multilingual automation for regulated enterprises

Parloa homepage

Parloa targets Fortune 500 contact centres and supports more than 130 languages, according to the Parloa product site. That makes it relevant to regulated organisations that need broad multilingual coverage. Its pricing is custom, and the research pack contains no independent latency result.

A sales-led implementation can support governance and deep integration. It also makes cost and time-to-value difficult to compare with a self-serve API.

Pros

  • Supports more than 130 languages.
  • Targets regulated, large-enterprise contact-centre work.

Cons

  • Pricing requires a custom quote.
  • Neither benchmark published an independent latency result.

Run a fair bake-off

Use the same carrier route, phone number, prompt and tool schema for every candidate. Keep endpointing settings fixed. Record caller speech-end and first audible agent byte from outside the platform, then calculate median and p95 from those recordings. Log failed turns separately rather than dropping them without explanation. The same discipline that keeps any agent reliable applies once the agent is on a phone line.

The functional tests matter too. Interrupt the agent, correct a detail mid-sentence and introduce background noise. Test a warm transfer and inspect the context received by the human. For a write action, force a webhook timeout after the action succeeds; the retry must check an idempotency key before attempting the action again. Compare the complete bill for that configuration, including telephony, providers, compliance and concurrency.

How to pick

Start with ownership, exactly as you would for any agentic workflow. A managed platform is appropriate when the team wants a working phone stack and accepts platform constraints. An API or framework fits a product team that wants provider control and can operate the resulting boundaries. Enterprise contact-centre products belong in a procurement process centred on governance and integration.

Then measure latency on the route that will carry production traffic. Ask for p95 rather than an average. Tail ratio is p95 divided by median; Openbenchmarks found Bland at 1.48× and Telnyx at 1.43×, compared with ElevenLabs at 1.24×. A high p95 produces noticeably slow turns even when the median looks acceptable.

Cost follows the architecture. Retell pricing makes individual production controls visible, while Vapi pricing separates hosting from provider use and compliance. The AssemblyAI roundup advises budgeting two to three times an advertised figure for a real-world all-in cost. Use that only as a planning check; build the actual bill from the components in your chosen route.

Honourable mentions

  • NICE Cognigy: an enterprise conversational-AI option for a sales-led shortlist.
  • Regal: a sales-led enterprise platform whose pricing page routes buyers to a custom quote.
  • Telnyx Voice AI: the fastest Openbenchmarks median at 1,296 ms, with a 1.43× tail.
  • Cartesia: a model-layer option to compare when speech generation is the priority.
  • Thoughtly: another managed option for business call flows.

Frequently Asked Questions

What is the best AI voice agent in 2026?

There is no universal winner because the products occupy different layers, in the same way AI agents differ from fixed automation. Retell and Vapi are the strongest managed-platform shortlist from the available evidence. ElevenLabs had the steadiest measured tail in Openbenchmarks. Large contact centres should compare PolyAI and Parloa on integration and governance rather than API pricing.

Which AI voice agent has the lowest latency?

It depends on the measurement boundary. Tested Media put Retell first at a 680 ms median. Openbenchmarks put Telnyx first at 1,296 ms over a real phone line. Request end-to-end p95 on your carrier route.

How much does an AI voice agent cost per minute?

Vapi charges $0.05 per minute for hosting before provider costs. Retell lists a $0.07–$0.31 per-minute range, with paid production controls. Tested Media lists Synthflow at $0.13–$0.20. The all-in total also depends on telephony and the model stack.

Is Vapi better than Retell AI?

Vapi is the better fit when provider choice and bring-your-own keys are requirements. Retell exposes component costs clearly and requires fewer provider decisions. Their benchmark order also changes with the method, so compare both using the same call flow and carrier.

Can an AI voice agent transfer a call to a human?

Yes. A warm handoff should pass the reason for the call, verification status and attempted actions so the caller does not repeat the story. Repeated misunderstanding or a direct request for a person should trigger escalation.

What should I test before launching an AI voice agent?

Start with a repetitive, low-risk intent. Appointment confirmation is a safer first job than open-ended support, and the same scoping logic applies to any agent you deploy. Test interruptions, corrections and unclear requests. Verify that a retry cannot duplicate a completed tool action. Roll out to a limited traffic segment first and review transfers alongside business outcomes.

Sources

Last updated: Sep 1, 2026

Build your agent team in 30 seconds.

Build agent teams that work along with your team. Free to start, no card required.