AI voice agent pricing in the self-serve examples reviewed here starts with advertised base rates from $0.05 to $0.14 per minute, before exclusions and required plan fees. That is only a starting point. Some rates cover the runtime alone, others include speech or model usage, and others add a monthly or annual platform fee. Configured, premium, and enterprise offers can cost substantially more.
The useful number is your total cost per successful outcome. Start with every billable component and fixed operating cost, then divide the total by resolved calls, booked appointments, qualified leads, or whichever result the agent must deliver.
The rates in this guide were verified on official pricing pages on July 29, 2026. They are public USD list prices, not quotes. Recheck them against current pricing and your contract before making a purchasing decision.
What AI voice agents cost in 2026
Current pricing pages demonstrate why a bare per-minute comparison does not work. Two providers can both advertise $0.05 per minute while including different parts of the voice stack.
| Platform | Current public price | What the published rate covers | Important costs or conditions outside it |
|---|---|---|---|
| Dasha | Growth starts at $0.08/min | Managed platform, described as all-inclusive except the listed exclusions | VoIP and large language model (LLM) tokens; volume and committed-usage discounts are available |
| Vapi | $0.05/min Build hosting | Vapi hosting and orchestration | Speech-to-text (STT), LLM, text-to-speech (TTS), and network transport are passed through or billed by the provider |
| Retell AI | $0.07-$0.31/min published pay-as-you-go headline range | Retell voice infrastructure plus the selected TTS and model configuration | Telephony, optional features, capacity above 20 active calls, and prompt or short-call exceptions; detailed model and voice selections can exceed the headline range |
| Telnyx | $0.05/min base | Orchestration, Telnyx-hosted STT, and listed hosted TTS | LLM usage; telephony and applicable Voice API fees; premium third-party voices or models, recording, messaging, and knowledge-base add-ons where selected |
| ElevenAgents | Plans from $0; extra minutes $0.08/min | Subscription minute allowance and the hosted speech/agent layer | LLM and telephony at cost; plan fee, concurrency limits, and higher-priced burst minutes |
| Bland AI | $0.14/min Start; $0.12 or $0.11/min on paid tiers | LLM, STT, and TTS in the AI minute | Telephony through your own carrier or Bland at pass-through cost; plan-specific transfer charges; $0.015 Bland-telephony outbound-attempt minimum |
| Deepgram Voice Agent API | Standard pay as you go: $0.075/min ($4.50/hour); Advanced: $0.163/min | STT, managed LLM processing and orchestration, and TTS | Lower bring-your-own-model and Growth rates are available; deployment choices change the commercial structure |
| Twilio ConversationRelay | $0.07/min | Hosted STT/TTS speech interface between a live call and the customer's AI application | Programmable Voice charges, phone numbers, and the customer's reasoning/model layer |
| Synthflow | For new deployments, Enterprise starts at $30,000/year | A scoped enterprise deployment and support package | Final pricing depends on call volume, concurrency, telephony, integrations, security, and launch support |
These are different commercial units. Vapi's $0.05 is a hosting layer. Telnyx's $0.05 includes STT and several voices but not the model or phone call. Bland bundles the AI components but not carrier cost. Synthflow prices an enterprise program rather than a self-serve minute. The lowest visible number is not necessarily the lowest total.
Why the advertised minute is not the total cost
An AI voice agent creates costs from the first dial or inbound connection through the business outcome. The exact boundary depends on the architecture and contract.

Call attempts and telephony
Telephony depends on direction, destination, number type, and the number of live call legs. An outbound attempt may add answering-machine detection, branded calling, or dialer fees before a useful conversation occurs. A transfer can create another phone leg or conference participant charge.
Billing increments matter too. Twilio says partial minutes for Programmable Voice are rounded up to the next whole minute. When a two-way flow uses two Call SIDs, each call is billed and rounded separately. That is not a universal carrier rule, but it shows why a workload with many 20-second calls can have a different effective rate from one long call with the same total raw duration.
Speech, model, and runtime
A cascaded voice agent usually has three AI meters:
- STT converts caller audio to text and is commonly priced by processed audio time.
- The LLM reads instructions, conversation history, retrieved knowledge, and tool results, then produces response tokens.
- TTS turns the response into audio and may be priced by characters, tokens, or generated audio time.
The runtime coordinates turn-taking, interruptions, tools, routing, deployment, and call state. Some platforms bundle speech into the runtime rate. Others pass every provider invoice through separately. A native speech-to-speech model uses audio-token or audio-minute pricing, so do not add separate STT and TTS unless the architecture actually invokes them.
Operations and fixed requirements
Production cost continues after the call. Include phone numbers, concurrency, recordings, storage, analytics, automated evaluation, test calls, support, security requirements, and data retention. Then include the engineers and operators who maintain integrations, review failures, update knowledge, test releases, monitor latency and completion, and handle incidents.
Calculate your monthly AI voice agent cost
Use this formula as the starting point:
monthly total cost = connected minutes x loaded connected-minute rate + attempted-call, transfer, and event fees + fixed platform, number, capacity, support, and compliance fees + implementation and monthly engineering labor
cost per successful outcome = monthly total cost / successful outcomes
The loaded connected-minute rate is the sum of every component that scales with a connected minute: runtime, telephony, STT, TTS, LLM usage, recording, and any minute-priced add-ons. Build it in six steps.
1. Measure attempts, connects, duration, and outcomes
For outbound calls, keep attempted, connected, transferred, and successful calls as separate fields. For inbound work, start with connected calls and their duration distribution. Do not use average handle time alone if you have many short calls, because minimums and rounding apply at the call or leg level.
2. Convert each component to its own billable unit
Use provider records where possible:
telephony cost = sum(billable minutes for every live call leg x leg rate) STT cost = processed audio minutes x STT rate TTS cost = billable characters or audio units x TTS rate LLM cost = input tokens x input rate + output tokens x output rate runtime cost = billable agent minutes x runtime rate
Before launch, estimate token and character usage from representative transcripts. After launch, replace those estimates with the usage values returned by each API.
3. Add attempts, transfers, and optional features
Per-event charges can be material at high volume. Retell, for example, lists Batch Call at $0.005 per dial, while Bland documents a $0.015 minimum for outbound attempts through its telephony. Knowledge retrieval, denoising, guardrails, personal-data removal, quality assurance, recording, and messaging may also have separate rates.
4. Size peak capacity separately
Monthly minutes do not tell you how many calls arrive at the busiest moment. Track peak active calls and call-start rate in short intervals. Then ask what happens at the limit: queue, reject, throttle, burst at a higher rate, or require more purchased capacity.
Public plans illustrate the differences. Vapi includes 10 concurrent calls and lists extra lines at $10 per month. Retell includes 20 and lists extra standard capacity at $8 per active-call slot per month. ElevenAgents uses plan limits and, when enabled, prices burst minutes at twice its standard extra-minute rate. Dasha's Growth plan advertises no per-line fee.
5. Add fixed and operating costs
List monthly platform fees, phone numbers, minimum commitments, support, compliance, storage, monitoring, and internal labor separately. Do not hide these costs inside a very large minute forecast. A $1,000 fixed requirement adds $1 per minute at 1,000 minutes and only $0.01 per minute at 100,000 minutes.
6. Reconcile to outcomes
Once traffic is live, join the billing export to call outcomes. Track cost per attempt, connected call, transfer, and successful result. Cost per successful outcome is the strongest comparison because it includes quality. A low-priced agent that retries more calls, transfers more often, or completes fewer tasks can be more expensive in practice.
Three monthly cost scenarios
The scenarios below show the method with one unbundled US outbound stack. They use public list prices for Vapi hosting, Twilio telephony and recording, Deepgram speech, OpenAI GPT-5 mini, and Datadog logs. The labor assumptions combine May 2024 BLS median wages for software developers and customer service representatives with March 2026 employer-compensation data. These are national planning rates, not contractor or vendor quotes. The scenarios are not Dasha pricing, vendor quotes, or forecasts.
| Example rate | List-price input |
|---|---|
| US local outbound telephony | $0.014 per billable minute for each outbound call/Call SID + $0.0075 answering-machine detection per completed call with AMD enabled + $0.0018 per US conference participant-minute + $1.15 per US local number/month |
| Speech | Deepgram Flux English streaming PAYG promotional rate at $0.0065 per STT minute + Aura-2 PAYG at $0.03 per 1,000 TTS characters |
| LLM | $0.25 per 1M input tokens + $2.00 per 1M output tokens |
| Platform hosting and concurrency | $0.05 per Vapi call minute; 10 concurrent calls included, then $10 per additional line/month |
| Recording and storage | $0.0025 per recorded minute + $0.0005 per stored recording minute/month above the first 10,000 average stored minutes per project |
| Logs | $0.10 per GB ingested + $1.70 per 1M indexed log events (15-day retention, billed annually) |
| Labor | $92 per engineering hour + $29.50 per operations hour |
For modeling, the scenarios use AI-active duration as a proxy for Vapi billable call minutes and apply Twilio answering-machine detection to every outbound attempt as a deliberately conservative ceiling, even though Twilio bills it per completed call. Actual answering-machine detection spend depends on the number of completed calls with the feature enabled. The examples also apply marginal rates from the first unit and ignore promotional or included minutes so the comparison remains conservative and every term stays visible.
| Planning input | Pilot | Growing production | High-volume program |
|---|---|---|---|
| Outbound attempts | 2,000 | 20,000 | 120,000 |
| Connection rate | 35% | 45% | 50% |
| Connected calls | 700 | 9,000 | 60,000 |
| Average AI-active time | 3 min | 4 min | 5 min |
| Connected AI minutes | 2,100 | 36,000 | 300,000 |
| Transfer share | 5% | 12% | 20% |
| Engineering hours/month | 12 | 60 | 240 |
| Human-only minutes per transfer | 4 | 5 | 6 |
| Warm overlap per transfer | 0.5 min | 0.75 min | 1 min |
| AI speaking share | 45% | 50% | 55% |
| LLM input tokens/AI minute | 1,500 | 3,000 | 6,000 |
| LLM output tokens/AI minute | 150 | 200 | 250 |
| Planned concurrency | 1 | 11 | 73 |
| Phone numbers | 1 | 5 | 20 |
| Recording retention | 1 month | 3 months | 6 months |
| Calls manually reviewed | 2% | 3% | 4% |
The model assumes 900 TTS characters per spoken minute, 20 minutes of manual review per flagged call, 0.25 MB of logs and traces per AI-active minute, and 20 indexed events per AI-active minute. Replace every assumption with your own data.
| Monthly output | Pilot | Growing production | High-volume program |
|---|---|---|---|
| Cash infrastructure | $203 | $3,600 | $32,564 |
| Engineering and exception review | $1,242 | $8,175 | $45,680 |
| Total before live handoff labor | $1,445 | $11,775 | $78,244 |
| Optional live handoff labor | $77 | $3,053 | $41,300 |
| Full workflow total | $1,522 | $14,828 | $119,544 |
| Cost before handoff labor per connected call | $2.06 | $1.31 | $1.30 |
| Illustrative success rate among connected calls | 40% | 50% | 60% |
| Cost before handoff labor per successful outcome | $5.16 | $2.62 | $2.17 |
The success rates are planning assumptions, not performance claims. The cash infrastructure works out to roughly $0.097, $0.100, and $0.109 per AI-active minute in these examples. For readability, the outputs use aggregate call-leg duration and do not apply Twilio's per-call or per-recording whole-minute rounding; actual charges can be higher, especially when calls or transfer legs are short. Internal labor dominates the pilot and remains material at scale. Live handoffs can exceed the AI infrastructure cost in the high-volume scenario, so keep them visible rather than treating escalation as free.
Sensitivity is easy to calculate. Every $0.01 change in the loaded minute rate changes monthly spend by $21 at 2,100 minutes, $360 at 36,000 minutes, and $3,000 at 300,000 minutes. One extra AI-active minute on every connected call affects telephony, speech, runtime, model use, recording, and possibly capacity.
Seven costs that change the invoice
1. Failed calls, voicemail, and short conversations
Ask when each layer starts billing: dial, ring, answer, agent audio, or successful connection. Dasha says it charges its agent layer only for successful connections. Retell says failed-to-connect calls are not billed at its AI layer, but connected voicemail and silence remain billable while the agent is active. Bland can apply an attempt minimum on its telephony. No single rule applies across the whole stack.
2. Minimums, rounding, silence, and hold time
Retell documents a 10-second minimum for certain dynamic AI-first openings and proportional billing adjustments for long prompts. ElevenLabs measures the full connection and applies a 95% discount only to qualifying silence periods longer than 10 seconds. The carrier or runtime can keep billing even when a model layer discounts silence.
3. Transfers and overlapping call legs
Determine whether a transfer is a SIP refer, bridge, conference, or new call. Ask whether the AI remains connected and which legs continue billing. Retell says its AI-agent fee stops after transfer while telephony continues. Another architecture may keep the agent attached during a warm handoff and bill three live participants for part of the call.
4. Model, prompt, and voice choices
Long system prompts, replayed history, retrieved knowledge, and tool output increase input tokens. Premium models or voices may improve completion enough to justify their rate, but test that with the same workflow. Cache static greetings or responses only when the provider terms and product behavior allow it.
5. Concurrency, throughput, and burst behavior
Price peak capacity with arrival data, not monthly averages. Confirm active-call limits, calls per second, queue limits, failover headroom, and the price or behavior above the limit. Simulation and test traffic may consume capacity too.
6. Recording, storage, analytics, and evaluation
Retention compounds as new recordings are added every month. Automated quality scoring can also be a large minute-priced line item; Retell lists AI Quality Assurance at $0.10 per analyzed minute after the first 100 analyzed minutes per workspace. Choose retention and evaluation coverage from risk and operational needs rather than enabling every feature by default.
7. Support, compliance, integration, and maintenance
Enterprise requirements can outweigh a few cents of usage-rate difference. Vapi publicly lists HIPAA at $2,000 per month and Zero Data Retention at $1,000 per month. Dedicated support, service-level agreements, single sign-on, data residency, custom retention, and launch help may require a contract. A custom stack must budget equivalent integration, monitoring, release, and on-call responsibilities.
Compare bundled, modular, and do-it-yourself pricing
| Model | Where it helps | What to include in total cost |
|---|---|---|
| Bundled managed platform | Faster forecasting and fewer vendor seams | Confirm which telephony, models, support, capacity, and enterprise controls remain outside the bundle |
| Modular voice platform | More provider choice and component-level cost control | Add every speech, model, carrier, runtime, storage, testing, and support meter; budget integration ownership |
| Do-it-yourself or open-source stack | Maximum architectural control and potential infrastructure savings at scale | Count engineering, deployment, observability, provider failover, evaluation, upgrades, security controls, and on-call support |
Dasha's managed voice AI backend is designed for teams that want control without operating every seam themselves. If you are comparing a developer platform, our guide to voice AI backends versus API platforms explains the ownership tradeoff. If you are assembling the full stack, compare the same responsibilities in our managed versus do-it-yourself breakdown.
None of the three models is automatically cheapest. The right choice depends on scale, provider requirements, available engineering capacity, reliability targets, and how much operational ownership creates product advantage.
How Dasha pricing fits the model
Our current pricing uses a free Developer plan and a usage-based Growth plan:
- Developer includes 1,000 free minutes, one concurrent call, API access, and email support.
- Growth starts at $0.08 per minute and excludes VoIP and LLM tokens from that rate.
- We bill connected usage to the second with no minimum call duration.
- We do not charge our agent layer for attempts that do not connect.
- Growth advertises no per-line fees, with volume and committed-usage discounts available.
To forecast a Dasha deployment, add the carrier, model, optional service, and internal operating costs required by your workload. Then calculate cost per outcome using the same method above. Dasha's getting-started path includes 1,000 free minutes with no card required. Use those test calls to estimate your own duration distribution before discussing paid production terms.
Use this quote-comparison checklist
Ask every shortlisted provider the same questions:
- What does the published rate include: runtime, STT, TTS, LLM, telephony, numbers, recording, and support?
- When does billing start and stop for an attempt, connected call, voicemail, silence, hold, and transfer?
- What is the minimum or rounding increment for each component and each call leg?
- How do destination country, direction, local versus toll-free numbers, and carrier surcharges change the rate?
- How do model, prompt length, conversation history, voice, language, and optional features change spend?
- What concurrency and call-start rate are included, and what happens at the limit?
- Which recording, retention, analytics, evaluation, privacy, and compliance requirements add fixed or usage fees?
- Which support, uptime, dedicated-capacity, or implementation terms require an enterprise agreement?
- What commitments, renewal terms, overages, ramp periods, and unused-volume rules apply?
- Can usage exports be joined to business outcomes so you can measure cost per successful result?
Build the first estimate from public prices, then replace assumptions with pilot traffic and a written quote. That produces a budget you can defend, and it prevents a low headline rate from hiding a high cost per completed call.
