Compare Retell AI and ElevenAgents, ElevenLabs’ conversational-agent product, on telephony, workflow control, testing, pricing, concurrency, client delivery, and production operations, with Dasha as a managed-runtime alternative.
Retell AI and ElevenLabs now compete as full voice agent platforms. The decision affects much more than voice quality. Telephony, conversation control, testing, concurrency, client SDKs, and the cost of each connected minute all matter once an agent reaches production. Here is the practical comparison for technical teams choosing a platform for phone, web, or mobile conversations.
Retell AI vs ElevenLabs at a glance
Retell AI is centered on phone-call operations. ElevenAgents combines conversational agents with ElevenLabs' own speech stack and broader client-channel support. Both can place and receive calls, transfer callers, run tools, use knowledge bases, test agents, and analyze completed conversations.
For a serious conversational AI product, we recommend evaluating Dasha's managed voice AI backend alongside both. We combine a managed runtime with REST APIs, flexible text-to-speech and language-model providers, telephony, web voice and chat, bulk execution, testing, logs, and monitoring. This is especially relevant when you are building a multitenant product or expect high concurrency.
Dasha, our pick
- Primary fit: Production conversational AI products that need API control, provider choice, and scale
- Provider stack: Supported text-to-speech providers and language models
- Agent control: Prompts, tools, webhooks, MCP connections, knowledge, and API configuration
- Phone operations: Inbound and outbound SIP, call transfers, IVR handling, and bulk scheduled calls
- Web delivery: Web voice and chat over WebRTC and WebSocket
- Production operations: Call Inspector recordings, transcripts, and timeline evidence; separate Activity Logs; deployment; and concurrency monitoring
Retell AI
- Primary fit: Phone operations with detailed telephony controls
- Speech stack: Platform voices plus third-party providers, including ElevenLabs
- Agent control: Single Prompt for simple conversations and Conversation Flow for new structured agents; Multi-Prompt remains a legacy option
- Phone operations: Managed numbers, custom SIP, batch calling, SMS on eligible numbers, branded calling for approved US numbers, transfers, and IVR controls
- Web delivery: Web voice plus a chat and callback widget
- Production operations: Simulation and batch tests, versions, A/B tests, live call takeover, analytics, and post-call analysis
ElevenAgents
- Primary fit: Voice-led agents across web, mobile, phone, chat, and WhatsApp
- Speech stack: ElevenLabs speech recognition and synthesis, with voices from ElevenCreative's library and cloning tools
- Agent control: Prompts, visual workflows, task-specific procedures, tools, MCP, knowledge, and alpha guardrails
- Phone operations: Twilio and SIP connections, batch calling, transfers, DTMF, voicemail handling, and SMS
- Web and app delivery: Web and mobile SDKs, chat, WebSocket, and WhatsApp
- Production operations: Automated tests, simulation, versions, traffic experiments, enterprise live monitoring, analytics, and OpenTelemetry
The practical split is clear. Retell has more phone-operation detail in its self-serve platform. ElevenAgents keeps the agent close to ElevenLabs' voice models and supports more client application surfaces. Dasha gives technical teams a managed production runtime without tying the product to one text-to-speech provider or charging per concurrent line.
What each platform actually is
Dasha

Dasha is a managed production platform for real-time conversational AI. Technical teams configure agents through a web application or REST API, connect phone and web channels, call backend tools, and attach knowledge. The Call Inspector shows completed-call recordings, transcripts, timeline events, model interactions, tool executions, and latency evidence. Activity Logs separately record organization-wide call lifecycle, webhook, tool, configuration, and API events.
Our fit is strongest when the agent is part of a product rather than a standalone campaign. A SaaS team can configure separate SIP trunks, prompts, knowledge bases, and API keys per customer, with isolated logging and analytics for each tenant. Dasha supports multiple text-to-speech providers and multiple language-model providers, so a provider change does not require replacing the surrounding agent platform.
Dasha is a weaker fit when your main deliverable is generated media, a voice marketplace, or a voice-cloning workflow. ElevenCreative is built around those audio capabilities. Retell also has phone-specific extras, such as a branded-calling add-on for US numbers. It requires a verified business profile and an application, and Retell recommends adding its separate verified-number service as well. The displayed name and character limit vary by carrier.
Retell AI

Retell AI is a hosted voice and chat agent platform with a strong telephony center. A Single Prompt agent covers simple conversations. For new structured, multi-step agents, Retell recommends a graph built with Conversation Flow nodes for logic, code, tools, transfers, SMS, and agent handoffs. Multi-Prompt remains available as a legacy option for existing agents. Retell also supports custom functions, remote MCP servers, knowledge bases, web calls, chat agents, and website widgets.
Its phone tooling goes beyond basic inbound and outbound calling. Retell can sell US and Canadian numbers, import numbers through SIP, connect common carriers and contact-center platforms, schedule batch calls, navigate interactive voice response (IVR) menus with keypad tones, and perform cold or warm transfers. SMS from an agent's own number requires an eligible US, non-toll-free Twilio number or custom number that has passed A2P approval. An SMS-approved Retell number can send Retell's preset template during a call without A2P approval, but it cannot send custom text or support two-way messaging. Live monitoring supports transcripts, silent listening, call takeover, and termination.
Retell's speech layer is modular. Platform voices sit beside third-party voice options, including ElevenLabs. The custom voice workflow supports searching ElevenLabs community voices, importing a voice clone, and training a clone. That makes Retell AI versus ElevenLabs a platform decision, a speech-provider decision, or both.
ElevenAgents

ElevenAgents is the conversational agent product. ElevenCreative is the media and voice-creation product. Treating the company as only a text-to-speech API misses most of the current comparison.
ElevenAgents has a visual workflow builder, task-specific free-form and structured procedures, custom tools, MCP connections, knowledge retrieval, personalization, alpha guardrails, and a choice of supported or custom language models. ElevenCreative supplies the voice library and voice-cloning ecosystem. ElevenAgents' voice controls apply those voices to live agents with pronunciation dictionaries, speed control, expressive mode, multi-voice conversations, and language-specific configuration.
The platform also handles telephony. It supports native Twilio integration and SIP trunking, additional carrier connections, integrations with contact-center systems, batch calls, transfer to a number or another agent, keypad tones, voicemail detection, and text messaging through Twilio. Web and mobile SDKs, chat, and WhatsApp support give it a wider client-channel footprint than a phone-centered platform.
Feature-by-feature comparison
Voice design and provider flexibility
ElevenLabs keeps voice creation and agent delivery in the same vendor ecosystem. A team can design or clone a voice in ElevenCreative, then select it and tune pronunciation, speed, language behavior, or expressive delivery in ElevenAgents. This reduces integration work for products where a recognizable voice, emotional delivery, or multilingual speech is a central part of the experience.
Retell can use ElevenLabs voices while retaining Retell's orchestration and telephony. It also supports other voice providers and curated platform voices. This gives teams more room to change the speech vendor for cost, language, reliability, or pronunciation reasons.
Dasha takes a similar modular approach. Supported text-to-speech options include ElevenLabs, Cartesia, Inworld, and LMNT, while the surrounding conversation runtime, telephony, tools, and logs stay in Dasha. Provider choice matters when a product serves several customers or regions, because one TTS model rarely leads on every voice, language, and price point.
Fit: ElevenAgents suits products that want close access to ElevenLabs' voice ecosystem. Retell suits phone agents that want access to ElevenLabs voices plus other providers. Dasha suits products that treat language models and text-to-speech providers as replaceable parts of a managed runtime.
Agent logic and integrations
Retell offers two current paths for new agents. A Single Prompt covers straightforward conversations. Conversation Flow is the recommended path for structured agents, adding a visual graph with deterministic rules, reusable subflows, code, tools, and LLM-directed transitions. Multi-Prompt is a legacy option for teams maintaining agents already built around prompted states.
ElevenAgents combines prompts with visual workflows and agent procedures. Each procedure contains instructions for one task and loads when its trigger matches. Structured procedures run a fixed sequence of typed steps, while free-form procedures adapt task instructions to the conversation. Procedures belong to one agent and cannot be shared as workspace-level resources across agents. Agent tools can run in the client, call webhooks, use system functions, or connect to MCP servers. Environment variables support the same agent across development, staging, and production.
Both platforms connect to calendars, customer systems, and custom APIs. Retell includes Salesforce and HubSpot contact synchronization. ElevenAgents lists a broader set of packaged business integrations, including Salesforce, HubSpot, Zendesk, ServiceNow, Intercom, Slack, and Google Drive. The deciding question is whether you need a prebuilt connector or a stable API and webhook contract for your own integration layer.
Telephony and channels
Claims that ElevenLabs cannot make phone calls are obsolete. Both products support inbound and outbound phone agents, SIP, batch calls, transfers, IVR keypad tones, voicemail handling, and messaging. The difference is operational emphasis.
Retell exposes more call-operation products in one place. Managed and imported numbers, verified numbers, batch dialing, contacts, CRM mappings, live takeover, custom SIP headers, and call efficiency controls form a coherent phone operations surface. Its branded caller ID is limited to US numbers, requires an approved business profile and application, and is separate from the verified-number service that Retell recommends using with it. The displayed business name and supported character limit vary by carrier.
ElevenAgents covers phone delivery while investing more broadly in endpoints. Its client SDKs span web, iOS, Android, React Native, Python, and JavaScript. It also supports text chat and WhatsApp. This makes it relevant for a voice experience that follows a user across an app, website, messaging channel, and phone line.
Dasha covers inbound and outbound SIP, browser voice, web chat, bulk and scheduled calling, transfers, and per-customer telephony configuration through API. That last point matters for embedded voice products because SIP settings can remain part of each tenant's configuration.
Testing, releases, and live operations
Retell's testing methods include a manual playground, simulated users, pass criteria, batch regression runs, browser audio tests, and real phone tests. Agent versions and environment tags separate drafts from production, and A/B testing can split live traffic. After launch, teams can monitor and take over calls, build analytics dashboards, and extract structured post-call fields.
ElevenAgents supports automated agent tests and conversation simulation. Agent versioning covers branches, controlled traffic deployment, and production experiments. Its operations layer includes success evaluation, structured data extraction, sentiment, semantic conversation search, a real-time analytics dashboard, enterprise monitoring controls, and OpenTelemetry trace export.
Dasha's dashboard testing runs browser voice, chat, and real-phone tests from one mode selector. The Call Inspector contains recordings, transcripts, a timeline of model and tool events, and latency evidence for completed calls. Separate Activity Logs capture searchable organization-wide lifecycle, webhook, tool, configuration, and API events, while concurrency monitoring covers live capacity. Teams that need a managed runtime and deep per-call evidence get that production foundation without assembling an observability layer around a bare speech API.
Retell currently has a clear call-floor advantage through silent listening and human takeover. ElevenAgents has a clear software-release advantage through branches, traffic deployment, experiments, and OpenTelemetry. Dasha is the choice we prefer for product teams that need text-to-speech and language-model flexibility, multitenant configuration, and operational traceability in one managed backend.
Latency and conversational timing
Vendor latency numbers often measure different spans. A text-to-speech time-to-first-byte figure excludes speech recognition, turn detection, language-model inference, tool calls, audio transport, and the phone network. An end-to-end figure can include some or all of them. Comparing those values as one leaderboard produces a false result.
The useful delay measure starts at the end of the user's turn and ends with the first audible response at the device or phone endpoint. The 50th and 95th percentiles belong beside interruption stop time, false interruptions, and recovery after overlapping speech. A cross-language turn-taking study found that people consistently minimize silence and overlapping talk, even though conversational timing varies across languages.
A fair comparison holds the carrier, region, language model, voice, prompt, and utterances constant. A fast median can hide a poor tail. The 95th percentile and the frequency of awkward failures usually say more about production quality.
Pricing and concurrency
The headline per-minute price is only one line of the cost model. Total cost also includes the agent runtime, speech, language-model usage, carrier minutes, phone numbers, concurrency, failed or repeated tool calls, transfers, recording, and campaign add-ons. Our voice agent pricing guide explains how those layers change total cost.
Dasha pricing
Our current pricing bills Growth usage from $0.08 per minute and does not add a fee per concurrent line.
- Free entry: 1,000 minutes and one concurrent call
- Growth price: Starts at $0.08 per connected minute
- Language model: Separate
- Telephony: Separate
- Concurrency: One call on Developer. Growth supports up to 1,000 calls per agent or more, with no per-line fee
- Other cost drivers: Carrier and language-model usage
Retell AI pricing
Retell advertises a usage range of $0.07 to $0.31 per minute for AI voice agents. That range combines voice infrastructure, a selected text-to-speech voice, and a selected language model, while telephony and add-ons remain separate. Detailed speech-to-speech model configurations on the same pricing page can price above the $0.31 headline range, so $0.31 is not a hard ceiling. Retell's concurrency documentation explains the per-workspace default, reserved inbound capacity, and optional burst behavior.
- Free entry: $10 usage credit
- Usage price: Advertised at $0.07 to $0.31 per minute, based on voice and model choices
- Language model: Included in the advertised configuration range
- Telephony: Separate by destination or carrier
- Concurrency: Pay-As-You-Go workspaces default to 20 active calls. Additional standard slots cost $8 each per month
- Reserved inbound capacity: Reserves part of the standard pool for inbound traffic and reduces the portion available to outbound and web calls. It does not add slots
- Burst: If an admin enables burst, the workspace can reach the lower of three times its standard limit or its standard limit plus 300. A call that starts above the standard limit adds $0.10 per minute for its entire duration
- Other cost drivers: Numbers, knowledge bases, batch dials, branded calls, SMS, and other add-ons
ElevenAgents pricing
The ElevenAgents pricing page separates hosted agent minutes from language-model usage and telephony. Its burst documentation explains the higher rate above a plan's concurrency limit.
- Free entry: 15 call minutes and four concurrent calls
- Plan price: Monthly plans run from $6 to $990 with included minutes. Additional minutes cost $0.08
- Language model: Billed separately by model usage
- Telephony: Billed separately at carrier cost
- Concurrency: Plans include four to 40 simultaneous calls. Burst must be enabled per agent. For non-enterprise customers, it accepts excess calls up to the lower of three times the plan limit or 300 concurrent calls and charges those calls at twice the standard rate, $0.16 per minute on the listed $0.08 rate. Enabled burst calls are deprioritized and may have higher speech-processing latency
- Other cost drivers: Plan commitment, extra minutes, burst calls, language-model usage, telephony, and text messages
A 10,000-minute example shows why the formulas matter:
- Dasha: the starting runtime cost is about $800, with VoIP and language-model usage added. Growth does not add a per-line fee as concurrency rises.
- Retell AI: the advertised range produces $700 to $3,100 before telephony and optional add-ons. That headline-range estimate already includes Retell's advertised voice infrastructure, text-to-speech, and language-model components, but detailed speech-to-speech configurations can cost more than $0.31 per minute. The exact result depends heavily on the selected voice and model. A Pay-As-You-Go workspace defaults to 20 concurrent calls, and calls accepted through enabled burst mode add the documented surcharge.
- ElevenAgents: the $299 Scale plan includes 3,738 minutes and 30 concurrent calls. Another 6,262 minutes at $0.08 adds $500.96, for about $800 before language-model and telephony charges. Calls above the plan's concurrency limit use burst pricing only when burst is enabled and within its cap; otherwise they are rejected. The $990 Business plan includes 12,375 minutes and 40 concurrent calls.
The scenario assumes 10,000 connected agent minutes, a US-dollar monthly plan, no taxes, and no volume agreement. It excludes carrier charges for all three and excludes language-model charges for Dasha and ElevenAgents. Retell's advertised range already includes its listed model component. If traffic arrives in a short daily window, concurrency can change the result more than total minutes.
How to choose with a production pilot
A polished demo call cannot expose queueing, provider failures, off-script behavior, or a broken transfer. A fair decision uses the same small production pilot for every platform.
- One real workflow. The build connects the actual carrier, knowledge source, calendar or CRM, transfer destination, and post-call webhook. A prompt-only demo leaves the operational work unmeasured.
- A fixed scenario set. Cases cover noisy audio, accents, interruptions, corrections, silence, off-topic requests, slow tools, tool errors, voicemail, IVR menus, and a human handoff.
- Target concurrency. The run includes the expected steady load and a short burst, with connection failures, queueing, tail latency, and cost recorded.
- Business completion. The scorecard covers task completion, transfer success, correct structured data, tool error recovery, hang-up rate, and cost per completed task.
- One controlled release. A prompt, tool schema, and voice change exposes version history, test coverage, staged rollout, rollback, and audit evidence.
- A portability exercise. A speech or language-model change shows which prompts, voices, tests, tools, phone settings, and historical traces would have to move with a future platform change.
For outbound calls, platform compliance features do not make a campaign lawful by themselves. Consent, identification, calling windows, suppression lists, recording rules, and industry restrictions still belong in your application and operating process. The FCC has ruled that AI-generated voices fall under the Telephone Consumer Protection Act's restrictions on artificial or prerecorded voice calls (FCC declaratory ruling).
FAQ
Can Retell AI use ElevenLabs voices?
Yes. Retell's custom voice workflow lets teams search and add ElevenLabs community voices, import a voice clone, or train a clone. ElevenLabs usage is reflected in the selected Retell voice cost.
Does ElevenAgents support phone calls?
Yes. ElevenAgents supports inbound and outbound calls through Twilio, SIP trunks, and several carrier integrations. It also supports batch calls, transfers, voicemail detection, keypad tones, and SMS through Twilio. ElevenAgents still charges telephony separately from the hosted agent minute.
Is ElevenAgents the same as ElevenCreative?
No. The ElevenAgents platform handles live conversational agents. The ElevenCreative suite covers media creation for speech, dubbing, music, voiceovers, and related creative workflows. The products share ElevenLabs' speech technology but have different builders, pricing, and use cases.
Which platform is cheaper, Retell AI or ElevenLabs?
At the low end, Retell starts at $0.07 per agent minute and ElevenAgents charges $0.08 for additional hosted agent minutes. That does not settle total cost. Retell's voice and model combination changes its per-minute rate. ElevenAgents adds a monthly plan, language-model use, telephony, and possible burst pricing. A valid total-cost comparison uses the actual concurrency and provider mix.
If your pilot calls for a managed runtime, flexible text-to-speech and language-model providers, API-controlled multitenancy, and high concurrency without per-line fees, start building with Dasha.
