Voice AI conversion rates are easy to inflate by changing the denominator. A rate based on qualified conversations cannot be compared with one based on every eligible lead. Teams need a defined funnel, CRM-confirmed outcomes, and a controlled comparison before they can claim lift. Here is a practical framework for measuring voice AI across outbound and inbound sales calls.
There is no universal voice AI conversion rate
The useful definition is:
Voice AI conversion rate = verified target outcomes / all eligible opportunities assigned to the voice AI
The target outcome may be a qualified lead, completed transfer, booked appointment, attended appointment, order, or closed sale. Name it. Then name the population and time window in the denominator.
A single industry average cannot do that work. Published rates often mix inbound and outbound calls, warm and cold leads, calls and unique people, or bookings and sales. Some start counting only after a person answers. Others exclude technical failures and unanswered attempts. Those choices can produce very different percentages from the same campaign.
Use external figures as planning context only. Your decision benchmark should be a recent baseline from the same lead source, segment, offer, geography, contact policy, and conversion window. A concurrent control is stronger because it reduces the effect of seasonality and changes in lead mix.
Dasha belongs in this workflow as the production voice runtime and call-evidence layer. We help technical teams execute and inspect the conversations. Your CRM or other business system supplies the conversion truth. That boundary shapes every metric below.
Define the funnel before calculating conversion
Voice AI affects several gates between a lead and revenue. Keep those gates separate so a gain at one stage cannot conceal a loss at another.
For outbound sales, a practical funnel is:
- Eligible leads: unique people or accounts that meet the campaign rules after consent and suppression filters.
- Attempted leads: eligible leads that received at least one attempted call.
- Reached leads: leads where the intended person answered.
- Substantive conversations: reached leads that progressed beyond the greeting and identity check into the defined sales workflow.
- Qualified leads: prospects that met the same explicit qualification rules used by the control process.
- Converted leads: prospects that completed the primary outcome within the attribution window.
- Realized outcomes: attended meetings, activated accounts, retained orders, or recognized revenue.
An inbound funnel usually starts with offered calls, then moves through answered calls, resolved or qualified conversations, bookings or orders, and realized outcomes.
The conversion metrics that belong on the scorecard
| Metric | Calculation | What it tells you |
|---|---|---|
| Attempt coverage | Unique leads attempted / eligible leads | Whether the system worked the intended population |
| Reach rate | Intended people reached / eligible leads | The combined effect of list quality, timing, number reputation, and dialing |
| Conversation rate | Substantive conversations / people reached | Whether the opening and early interaction kept people engaged |
| Qualification rate | Qualified leads / substantive conversations | Whether targeting and qualification work once a conversation starts |
| Primary conversion rate | Verified outcomes / eligible leads assigned | The clean business-level rate for comparing variants |
| Downstream success rate | Realized outcomes / initial conversions | Whether early conversions were valuable and correctly qualified |
| Cost per verified outcome | Total campaign and follow-up cost / verified outcomes | Whether the economics improved |
| Opt-out rate | Explicit opt-outs / people reached | Whether growth is creating unacceptable customer or compliance risk |
Measure unique leads as well as calls. An AI agent can make more attempts than a human team, so conversions per call may fall while conversions per eligible lead rise. Both facts matter, but the lead-level rate answers the business question.
Use the CRM, scheduling system, billing system, or other system of record to confirm the outcome. An agent saying “your appointment is booked” is conversation evidence. The appointment record is outcome evidence.
How voice AI can change conversion rates
Voice AI does not add a fixed number of percentage points. It changes specific funnel inputs. The size and direction of the result depend on which constraint your current process has.
| Conversion lever | Where it appears | Measure it with | Common failure mode |
|---|---|---|---|
| Faster first contact | Reach and qualification | Time from assignment or inbound lead event to first attempt; reach rate by delay band | Fast contact to a low-quality or ineligible list |
| Broader coverage | Attempt coverage and reach | Coverage by hour, day, source, and segment | More attempts without better lead selection |
| Consistent follow-up | Reach and conversion | Attempts per lead and cumulative conversion by attempt | Excessive cadence, opt-outs, or number reputation damage |
| Consistent qualification | Qualification and downstream success | Required-question completion and CRM-confirmed lead quality | Rigid handling of ambiguous or unusual cases |
| Responsive conversation | Conversation and qualification | Turn latency, interruptions, repair loops, hang-up stage | Slow turns, talk-over, misheard entities, or long answers |
| Timely human handoff | Conversion and downstream success | Transfer completion, context delivered, and post-transfer outcome | A transfer that rings out or loses collected context |
Speed matters most when intent decays quickly. A large lead-response study reported in Harvard Business Review found that companies contacting an online lead within an hour were nearly seven times as likely to qualify it as companies that waited even one hour longer. That finding concerns lead qualification, comes from 2011 data, and does not establish a current AI close-rate benchmark. It does show why response time deserves its own measurement rather than a vague claim about “higher conversion.” (lead-response study)
Conversation quality also needs a separate scorecard. The 2026 EVA-Bench research evaluates voice-agent accuracy and experience as distinct dimensions across 213 enterprise scenarios. That separation is useful in production: a completed task can still contain long silence, poor interruption handling, or a confusing repair loop. (EVA-Bench research)
Our voice agent evaluation framework covers task outcome, entity accuracy, tool behavior, latency, conversation repair, and reliability in more detail. These are diagnostic and guardrail metrics. Revenue or another verified business event remains the conversion outcome.
Choose the comparison your decision requires
Two voice AI pilots can use the same software and answer different questions.
Test the complete operating model
Compare the current human process with the proposed AI process under the capacity and cadence each would actually use. This captures coverage, response time, consistency, operating cost, and conversation quality together. It answers whether replacing or augmenting the existing workflow creates business value.
Isolate conversation performance
Match the AI and human arms on lead source, dialing window, attempt policy, offer, caller-ID setup, and qualification rules. This removes much of the operational advantage and focuses the comparison on the conversation itself.
State which comparison you are running. A system-level result should not be presented as proof that the AI held a better individual conversation. A conversation-level experiment does not show the throughput advantage of around-the-clock execution.
Run a controlled voice AI conversion experiment
1. Define eligibility and one primary outcome
Write the inclusion, exclusion, and suppression rules before launch. Choose a primary outcome that the business system can confirm, such as an attended appointment within the predefined attribution window. Use secondary outcomes for diagnosis.
Keep compliance constraints fixed across variants. In the United States, the FCC treats an AI-generated voice as an artificial voice under the Telephone Consumer Protection Act. The FTC Telemarketing Sales Rule adds requirements around do-not-call suppression, caller identification, opt-outs, and campaign practices. These rules affect the eligible population and campaign design, so they belong upstream of the conversion calculation. (FCC AI voice ruling, FTC telemarketing guidance)
2. Pick the experimental unit
Randomize unique leads or accounts, rather than individual call attempts. Keep every attempt to the same person in one arm. If several contacts at one company can influence the same opportunity, randomize at the account level to avoid contamination.
3. Balance factors that strongly affect response
Stratify or verify balance by lead source, intent, geography, time zone, customer segment, and lead age. Use the same attribution window in both arms. A before-and-after comparison is vulnerable to changes in season, offer, list quality, staffing, and campaign timing.
4. Instrument the whole path
Join call records to business outcomes with a stable lead or account ID. Record:
- experiment arm, campaign, agent version, lead source, and segment;
- assignment, attempt, answer, conversation, transfer, and call-end timestamps;
- machine, wrong-party, opt-out, and technical-failure outcomes;
- qualification fields and where they came from;
- booking, attendance, opportunity, sale, and cancellation events; and
- telephony, platform, human follow-up, and operating costs.
Use explicit event names and timestamps. Free-form call summaries are useful for review, but they are a weak source for funnel calculations.
5. Set the sample and stopping rule before launch
Estimate the required sample from the control rate, the smallest lift worth detecting, the desired confidence level, and statistical power. Small baseline rates require more observations to distinguish signal from noise. Run through the full outcome window, then report the numerator, denominator, absolute lift in percentage points, relative lift, and confidence interval.
Do not stop after an unusually good day. Repeatedly checking and ending a test when the result looks favorable increases false positives. If sales mature slowly, use a defined leading indicator for operations and wait for the complete revenue window before making the revenue claim.
6. Analyze all assigned opportunities
The primary analysis should retain every eligible lead in its assigned arm, including unanswered calls and technical failures. This intention-to-treat view measures the system a buyer would actually deploy. A second analysis of reached leads or substantive conversations can explain where the difference arose, but it should not replace the primary result.
7. Protect the downside with guardrails
A higher booking rate can coexist with lower attendance, poorer lead quality, more opt-outs, or more failed transfers. Set limits for:
- opt-outs, complaints, and wrong-party contacts;
- critical qualification-field errors;
- duplicate bookings or CRM writes;
- transfer failures;
- technical failures and dropped calls;
- p50 and tail latency; and
- downstream cancellation, attendance, and close rates.
Run representative call scenarios before live traffic and turn production failures into regression cases. Our voice agent testing guide gives a practical release workflow for voice, tools, transfers, failures, and rollback.
Read the result as a funnel, not a headline
Start with the primary conversion rate, then locate the stage that moved.
| Result pattern | Likely interpretation | Evidence to inspect |
|---|---|---|
| Attempt coverage rises, reach is flat | More of the list was worked, with no improvement in contactability | Attempt timing, contact data, number reputation, suppression rules |
| Reach rises, substantive conversations fall | More people answered, but the opening or interaction lost them | First-turn transcript, hang-up stage, latency, disclosure, list intent |
| Qualifications rise, bookings stay flat | The qualification definition is too loose or the close is weak | Qualification fields, booking offer, objection paths |
| Bookings rise, attendance falls | Early conversion quality declined | Reminder path, expectation setting, cancellations, lead fit |
| Meetings rise, won revenue stays flat | The AI improved the top of the funnel without improving sales quality | Opportunity value, sales acceptance, close rate, sales-cycle length |
| Conversion is flat, cost per outcome falls | The economics improved through lower handling cost or greater capacity | Fully loaded campaign cost and verified outcomes |
Segment every finding. A global lift can hide a decline for a language, region, channel, lead source, or high-value account tier. Preserve both the overall business view and the unweighted slice table.
Measure voice AI conversion rates with Dasha
Dasha helps technical teams build and run production voice AI agents through a managed runtime, REST APIs, and a web application, with telephony, integrations, testing, monitoring, and large-scale call execution.
For conversion measurement, keep the boundary clear. We provide call-level evidence such as transcripts, model activity, tool executions, timeline events, and latency breakdowns through Call Inspector. Your team defines eligibility, qualification, the experiment, and the authoritative CRM or revenue outcome.
A sound implementation uses a stable business ID to connect each call with its later outcome, attaches the campaign and agent version to that record, inspects failures by funnel stage, and adds reproducible failures to the regression set. This lets you improve the agent without losing attribution or mistaking a better call summary for a verified sale.
Build one measured production call path with Dasha, then compare it with your current process using the same eligible population and outcome window.
