AI voice agent pricing: a practical guide to total cost

Technical product leader reviewing the connected cost layers of a live AI voice agent
Technical product leader reviewing the connected cost layers of a live AI voice agent

AI voice agent pricing in the self-serve examples reviewed here starts with advertised base rates from $0.05 to $0.14 per minute, before exclusions and required plan fees. That is only a starting point. Some rates cover the runtime alone, others include speech or model usage, and others add a monthly or annual platform fee. Configured, premium, and enterprise offers can cost substantially more.

The useful number is your total cost per successful outcome. Start with every billable component and fixed operating cost, then divide the total by resolved calls, booked appointments, qualified leads, or whichever result the agent must deliver.

The rates in this guide were verified on official pricing pages on July 29, 2026. They are public USD list prices, not quotes. Recheck them against current pricing and your contract before making a purchasing decision.

What AI voice agents cost in 2026

Current pricing pages demonstrate why a bare per-minute comparison does not work. Two providers can both advertise $0.05 per minute while including different parts of the voice stack.

PlatformCurrent public priceWhat the published rate coversImportant costs or conditions outside it
DashaGrowth starts at $0.08/minManaged platform, described as all-inclusive except the listed exclusionsVoIP and large language model (LLM) tokens; volume and committed-usage discounts are available
Vapi$0.05/min Build hostingVapi hosting and orchestrationSpeech-to-text (STT), LLM, text-to-speech (TTS), and network transport are passed through or billed by the provider
Retell AI$0.07-$0.31/min published pay-as-you-go headline rangeRetell voice infrastructure plus the selected TTS and model configurationTelephony, optional features, capacity above 20 active calls, and prompt or short-call exceptions; detailed model and voice selections can exceed the headline range
Telnyx$0.05/min baseOrchestration, Telnyx-hosted STT, and listed hosted TTSLLM usage; telephony and applicable Voice API fees; premium third-party voices or models, recording, messaging, and knowledge-base add-ons where selected
ElevenAgentsPlans from $0; extra minutes $0.08/minSubscription minute allowance and the hosted speech/agent layerLLM and telephony at cost; plan fee, concurrency limits, and higher-priced burst minutes
Bland AI$0.14/min Start; $0.12 or $0.11/min on paid tiersLLM, STT, and TTS in the AI minuteTelephony through your own carrier or Bland at pass-through cost; plan-specific transfer charges; $0.015 Bland-telephony outbound-attempt minimum
Deepgram Voice Agent APIStandard pay as you go: $0.075/min ($4.50/hour); Advanced: $0.163/minSTT, managed LLM processing and orchestration, and TTSLower bring-your-own-model and Growth rates are available; deployment choices change the commercial structure
Twilio ConversationRelay$0.07/minHosted STT/TTS speech interface between a live call and the customer's AI applicationProgrammable Voice charges, phone numbers, and the customer's reasoning/model layer
SynthflowFor new deployments, Enterprise starts at $30,000/yearA scoped enterprise deployment and support packageFinal pricing depends on call volume, concurrency, telephony, integrations, security, and launch support

These are different commercial units. Vapi's $0.05 is a hosting layer. Telnyx's $0.05 includes STT and several voices but not the model or phone call. Bland bundles the AI components but not carrier cost. Synthflow prices an enterprise program rather than a self-serve minute. The lowest visible number is not necessarily the lowest total.

Why the advertised minute is not the total cost

An AI voice agent creates costs from the first dial or inbound connection through the business outcome. The exact boundary depends on the architecture and contract.

AI voice agent cost model from attempts through connected minutes, voice stack, operations, total cost, and successful outcomes

Call attempts and telephony

Telephony depends on direction, destination, number type, and the number of live call legs. An outbound attempt may add answering-machine detection, branded calling, or dialer fees before a useful conversation occurs. A transfer can create another phone leg or conference participant charge.

Billing increments matter too. Twilio says partial minutes for Programmable Voice are rounded up to the next whole minute. When a two-way flow uses two Call SIDs, each call is billed and rounded separately. That is not a universal carrier rule, but it shows why a workload with many 20-second calls can have a different effective rate from one long call with the same total raw duration.

Speech, model, and runtime

A cascaded voice agent usually has three AI meters:

  • STT converts caller audio to text and is commonly priced by processed audio time.
  • The LLM reads instructions, conversation history, retrieved knowledge, and tool results, then produces response tokens.
  • TTS turns the response into audio and may be priced by characters, tokens, or generated audio time.

The runtime coordinates turn-taking, interruptions, tools, routing, deployment, and call state. Some platforms bundle speech into the runtime rate. Others pass every provider invoice through separately. A native speech-to-speech model uses audio-token or audio-minute pricing, so do not add separate STT and TTS unless the architecture actually invokes them.

Operations and fixed requirements

Production cost continues after the call. Include phone numbers, concurrency, recordings, storage, analytics, automated evaluation, test calls, support, security requirements, and data retention. Then include the engineers and operators who maintain integrations, review failures, update knowledge, test releases, monitor latency and completion, and handle incidents.

Calculate your monthly AI voice agent cost

Use this formula as the starting point:

monthly total cost = connected minutes x loaded connected-minute rate + attempted-call, transfer, and event fees + fixed platform, number, capacity, support, and compliance fees + implementation and monthly engineering labor
cost per successful outcome = monthly total cost / successful outcomes

The loaded connected-minute rate is the sum of every component that scales with a connected minute: runtime, telephony, STT, TTS, LLM usage, recording, and any minute-priced add-ons. Build it in six steps.

1. Measure attempts, connects, duration, and outcomes

For outbound calls, keep attempted, connected, transferred, and successful calls as separate fields. For inbound work, start with connected calls and their duration distribution. Do not use average handle time alone if you have many short calls, because minimums and rounding apply at the call or leg level.

2. Convert each component to its own billable unit

Use provider records where possible:

telephony cost = sum(billable minutes for every live call leg x leg rate) STT cost = processed audio minutes x STT rate TTS cost = billable characters or audio units x TTS rate LLM cost = input tokens x input rate + output tokens x output rate runtime cost = billable agent minutes x runtime rate

Before launch, estimate token and character usage from representative transcripts. After launch, replace those estimates with the usage values returned by each API.

3. Add attempts, transfers, and optional features

Per-event charges can be material at high volume. Retell, for example, lists Batch Call at $0.005 per dial, while Bland documents a $0.015 minimum for outbound attempts through its telephony. Knowledge retrieval, denoising, guardrails, personal-data removal, quality assurance, recording, and messaging may also have separate rates.

4. Size peak capacity separately

Monthly minutes do not tell you how many calls arrive at the busiest moment. Track peak active calls and call-start rate in short intervals. Then ask what happens at the limit: queue, reject, throttle, burst at a higher rate, or require more purchased capacity.

Public plans illustrate the differences. Vapi includes 10 concurrent calls and lists extra lines at $10 per month. Retell includes 20 and lists extra standard capacity at $8 per active-call slot per month. ElevenAgents uses plan limits and, when enabled, prices burst minutes at twice its standard extra-minute rate. Dasha's Growth plan advertises no per-line fee.

5. Add fixed and operating costs

List monthly platform fees, phone numbers, minimum commitments, support, compliance, storage, monitoring, and internal labor separately. Do not hide these costs inside a very large minute forecast. A $1,000 fixed requirement adds $1 per minute at 1,000 minutes and only $0.01 per minute at 100,000 minutes.

6. Reconcile to outcomes

Once traffic is live, join the billing export to call outcomes. Track cost per attempt, connected call, transfer, and successful result. Cost per successful outcome is the strongest comparison because it includes quality. A low-priced agent that retries more calls, transfers more often, or completes fewer tasks can be more expensive in practice.

Three monthly cost scenarios

The scenarios below show the method with one unbundled US outbound stack. They use public list prices for Vapi hosting, Twilio telephony and recording, Deepgram speech, OpenAI GPT-5 mini, and Datadog logs. The labor assumptions combine May 2024 BLS median wages for software developers and customer service representatives with March 2026 employer-compensation data. These are national planning rates, not contractor or vendor quotes. The scenarios are not Dasha pricing, vendor quotes, or forecasts.

Example rateList-price input
US local outbound telephony$0.014 per billable minute for each outbound call/Call SID + $0.0075 answering-machine detection per completed call with AMD enabled + $0.0018 per US conference participant-minute + $1.15 per US local number/month
SpeechDeepgram Flux English streaming PAYG promotional rate at $0.0065 per STT minute + Aura-2 PAYG at $0.03 per 1,000 TTS characters
LLM$0.25 per 1M input tokens + $2.00 per 1M output tokens
Platform hosting and concurrency$0.05 per Vapi call minute; 10 concurrent calls included, then $10 per additional line/month
Recording and storage$0.0025 per recorded minute + $0.0005 per stored recording minute/month above the first 10,000 average stored minutes per project
Logs$0.10 per GB ingested + $1.70 per 1M indexed log events (15-day retention, billed annually)
Labor$92 per engineering hour + $29.50 per operations hour

For modeling, the scenarios use AI-active duration as a proxy for Vapi billable call minutes and apply Twilio answering-machine detection to every outbound attempt as a deliberately conservative ceiling, even though Twilio bills it per completed call. Actual answering-machine detection spend depends on the number of completed calls with the feature enabled. The examples also apply marginal rates from the first unit and ignore promotional or included minutes so the comparison remains conservative and every term stays visible.

Planning inputPilotGrowing productionHigh-volume program
Outbound attempts2,00020,000120,000
Connection rate35%45%50%
Connected calls7009,00060,000
Average AI-active time3 min4 min5 min
Connected AI minutes2,10036,000300,000
Transfer share5%12%20%
Engineering hours/month1260240
Human-only minutes per transfer456
Warm overlap per transfer0.5 min0.75 min1 min
AI speaking share45%50%55%
LLM input tokens/AI minute1,5003,0006,000
LLM output tokens/AI minute150200250
Planned concurrency11173
Phone numbers1520
Recording retention1 month3 months6 months
Calls manually reviewed2%3%4%

The model assumes 900 TTS characters per spoken minute, 20 minutes of manual review per flagged call, 0.25 MB of logs and traces per AI-active minute, and 20 indexed events per AI-active minute. Replace every assumption with your own data.

Monthly outputPilotGrowing productionHigh-volume program
Cash infrastructure$203$3,600$32,564
Engineering and exception review$1,242$8,175$45,680
Total before live handoff labor$1,445$11,775$78,244
Optional live handoff labor$77$3,053$41,300
Full workflow total$1,522$14,828$119,544
Cost before handoff labor per connected call$2.06$1.31$1.30
Illustrative success rate among connected calls40%50%60%
Cost before handoff labor per successful outcome$5.16$2.62$2.17

The success rates are planning assumptions, not performance claims. The cash infrastructure works out to roughly $0.097, $0.100, and $0.109 per AI-active minute in these examples. For readability, the outputs use aggregate call-leg duration and do not apply Twilio's per-call or per-recording whole-minute rounding; actual charges can be higher, especially when calls or transfer legs are short. Internal labor dominates the pilot and remains material at scale. Live handoffs can exceed the AI infrastructure cost in the high-volume scenario, so keep them visible rather than treating escalation as free.

Sensitivity is easy to calculate. Every $0.01 change in the loaded minute rate changes monthly spend by $21 at 2,100 minutes, $360 at 36,000 minutes, and $3,000 at 300,000 minutes. One extra AI-active minute on every connected call affects telephony, speech, runtime, model use, recording, and possibly capacity.

Seven costs that change the invoice

1. Failed calls, voicemail, and short conversations

Ask when each layer starts billing: dial, ring, answer, agent audio, or successful connection. Dasha says it charges its agent layer only for successful connections. Retell says failed-to-connect calls are not billed at its AI layer, but connected voicemail and silence remain billable while the agent is active. Bland can apply an attempt minimum on its telephony. No single rule applies across the whole stack.

2. Minimums, rounding, silence, and hold time

Retell documents a 10-second minimum for certain dynamic AI-first openings and proportional billing adjustments for long prompts. ElevenLabs measures the full connection and applies a 95% discount only to qualifying silence periods longer than 10 seconds. The carrier or runtime can keep billing even when a model layer discounts silence.

3. Transfers and overlapping call legs

Determine whether a transfer is a SIP refer, bridge, conference, or new call. Ask whether the AI remains connected and which legs continue billing. Retell says its AI-agent fee stops after transfer while telephony continues. Another architecture may keep the agent attached during a warm handoff and bill three live participants for part of the call.

4. Model, prompt, and voice choices

Long system prompts, replayed history, retrieved knowledge, and tool output increase input tokens. Premium models or voices may improve completion enough to justify their rate, but test that with the same workflow. Cache static greetings or responses only when the provider terms and product behavior allow it.

5. Concurrency, throughput, and burst behavior

Price peak capacity with arrival data, not monthly averages. Confirm active-call limits, calls per second, queue limits, failover headroom, and the price or behavior above the limit. Simulation and test traffic may consume capacity too.

6. Recording, storage, analytics, and evaluation

Retention compounds as new recordings are added every month. Automated quality scoring can also be a large minute-priced line item; Retell lists AI Quality Assurance at $0.10 per analyzed minute after the first 100 analyzed minutes per workspace. Choose retention and evaluation coverage from risk and operational needs rather than enabling every feature by default.

7. Support, compliance, integration, and maintenance

Enterprise requirements can outweigh a few cents of usage-rate difference. Vapi publicly lists HIPAA at $2,000 per month and Zero Data Retention at $1,000 per month. Dedicated support, service-level agreements, single sign-on, data residency, custom retention, and launch help may require a contract. A custom stack must budget equivalent integration, monitoring, release, and on-call responsibilities.

Compare bundled, modular, and do-it-yourself pricing

ModelWhere it helpsWhat to include in total cost
Bundled managed platformFaster forecasting and fewer vendor seamsConfirm which telephony, models, support, capacity, and enterprise controls remain outside the bundle
Modular voice platformMore provider choice and component-level cost controlAdd every speech, model, carrier, runtime, storage, testing, and support meter; budget integration ownership
Do-it-yourself or open-source stackMaximum architectural control and potential infrastructure savings at scaleCount engineering, deployment, observability, provider failover, evaluation, upgrades, security controls, and on-call support

Dasha's managed voice AI backend is designed for teams that want control without operating every seam themselves. If you are comparing a developer platform, our guide to voice AI backends versus API platforms explains the ownership tradeoff. If you are assembling the full stack, compare the same responsibilities in our managed versus do-it-yourself breakdown.

None of the three models is automatically cheapest. The right choice depends on scale, provider requirements, available engineering capacity, reliability targets, and how much operational ownership creates product advantage.

How Dasha pricing fits the model

Our current pricing uses a free Developer plan and a usage-based Growth plan:

  • Developer includes 1,000 free minutes, one concurrent call, API access, and email support.
  • Growth starts at $0.08 per minute and excludes VoIP and LLM tokens from that rate.
  • We bill connected usage to the second with no minimum call duration.
  • We do not charge our agent layer for attempts that do not connect.
  • Growth advertises no per-line fees, with volume and committed-usage discounts available.

To forecast a Dasha deployment, add the carrier, model, optional service, and internal operating costs required by your workload. Then calculate cost per outcome using the same method above. Dasha's getting-started path includes 1,000 free minutes with no card required. Use those test calls to estimate your own duration distribution before discussing paid production terms.

Use this quote-comparison checklist

Ask every shortlisted provider the same questions:

  1. What does the published rate include: runtime, STT, TTS, LLM, telephony, numbers, recording, and support?
  2. When does billing start and stop for an attempt, connected call, voicemail, silence, hold, and transfer?
  3. What is the minimum or rounding increment for each component and each call leg?
  4. How do destination country, direction, local versus toll-free numbers, and carrier surcharges change the rate?
  5. How do model, prompt length, conversation history, voice, language, and optional features change spend?
  6. What concurrency and call-start rate are included, and what happens at the limit?
  7. Which recording, retention, analytics, evaluation, privacy, and compliance requirements add fixed or usage fees?
  8. Which support, uptime, dedicated-capacity, or implementation terms require an enterprise agreement?
  9. What commitments, renewal terms, overages, ramp periods, and unused-volume rules apply?
  10. Can usage exports be joined to business outcomes so you can measure cost per successful result?

Build the first estimate from public prices, then replace assumptions with pilot traffic and a written quote. That produces a budget you can defend, and it prevents a low headline rate from hiding a high cost per completed call.

Related Posts

We use cookies for functional and analytical purposes. Please refer to our Privacy Policy for details.