Replacing Amazon Transcribe starts with a scope decision. A speech-to-text API, a self-hosted model, and a managed voice-agent runtime solve different problems even when all three produce text. This buyer's guide compares seven options by workload, deployment, published pricing, and migration cost, then gives you a practical way to test them on your own audio.
Choose the replacement layer before the vendor
Amazon Transcribe is a managed speech recognition service. A close substitute accepts audio and returns transcripts through batch or streaming APIs. A wider platform can also manage turn-taking, telephony, application logic, and production operations. A self-hosted model gives your team control over the inference environment while also giving it the operating burden.
Those categories should not share a universal rank. The right shortlist depends on the system boundary you want to own.
| Option | What it replaces | Main modes | Deployment | Published entry pricing | Strongest fit |
|---|---|---|---|---|---|
| Dasha | Live voice-agent runtime, including the speech layer | Real-time voice conversations | Managed service | 1,000 minutes on Developer; Growth from $0.08 per connected minute, excluding telephony and language-model usage | Teams shipping production voice agents rather than a transcription feature |
| Deepgram | Speech API | Streaming and prerecorded | Managed API, with deployment options on enterprise plans | Prerecorded Nova-3 from $0.0043/min; promotional streaming Nova-3 from $0.0048/min | Developers who want a speech-focused API and explicit streaming models |
| AssemblyAI | Speech APIs plus a separate Speech Understanding API | Realtime and pre-recorded | Managed API | Realtime: Universal-Streaming from $0.15/hr; Universal-3.6 Pro Realtime at $0.45/hr. Pre-recorded: Universal-3.5 Pro at $0.21/hr | Product teams that want transcription and separately priced post-call analysis |
| Google Cloud Speech-to-Text | Speech API | Synchronous, asynchronous, streaming, and dynamic batch | Google Cloud | V2 standard recognition starts at $0.016/min; dynamic batch at $0.003/min | Google Cloud workloads and large asynchronous queues |
| Azure Speech | Speech API | Real-time, fast, and batch transcription | Azure cloud and containers | Region and agreement dependent | Microsoft estates or teams that require container deployment options |
| OpenAI transcription API | Speech API | File transcription and real-time transcription | Managed API | Estimated $0.0045/min for gpt-transcribe; $0.017/min for gpt-live-transcribe | Applications already built around OpenAI APIs |
| Self-hosted Whisper | Speech model and inference | Primarily file-oriented in the reference implementation | Your infrastructure | No API fee; compute and operations are yours | Teams that need model-level control and can run speech inference |
These prices are not normalized. Vendors charge by different units, include different features, and treat channels, add-ons, idle streaming time, support, and infrastructure differently. Use them to form a shortlist, then model the bill for the same workload. Procurement should use current quotes because published rates can change.
Amazon Transcribe is still a sensible AWS-native baseline

Amazon Transcribe handles recorded media and live streams, supports more than 100 languages, and includes features such as custom vocabulary, automatic language identification, speaker diarization, confidence scores, and personally identifiable information redaction. Amazon Transcribe Call Analytics adds contact-center functions such as call categorization and conversation characteristics.
The published US East examples are $0.006 per minute for batch and $0.01 per minute for streaming. Billing is by the second with no minimum. Standard pricing includes features such as vocabulary filters, diarization, and language identification, while content redaction, custom language models, and Call Analytics add cost. Two-channel audio is included in the standard service price.
Staying with Amazon Transcribe is rational when audio already lands in Amazon Web Services (AWS), identity and access management controls are established, and the service meets your measured accuracy and latency gates. A vendor change adds less value when it only recreates a working pipeline with new credentials and schemas.
1. Dasha: replace the voice-agent runtime, not a transcription endpoint

Choose Dasha when Amazon Transcribe is one component inside a live voice agent and the actual goal is to replace the surrounding production stack. We manage the real-time voice runtime, turn-taking, Session Initiation Protocol (SIP) and web call paths, integrations, testing, call history, and monitoring. The application remains yours: your business rules, customer data, downstream systems, carrier choices, compliance decisions, and acceptance criteria stay under your control. Our voice AI infrastructure guide explains that boundary in more detail.
The Developer plan includes 1,000 minutes. Growth starts at $0.08 per connected minute, excluding telephony and language-model usage. This cannot be compared directly with Amazon's recognition-only minute because we cover a wider operating layer.
We are a poor fit for bulk subtitles, archived-call transcription, media indexing, or another file-only workload. For those jobs, a direct speech API or a self-hosted model preserves the smaller system boundary. For a production voice agent, consolidating the speech and runtime path can remove glue services, although it is a broader migration than swapping one transcription endpoint.
2. Deepgram: a speech-focused API with separate streaming models

Deepgram separates its current speech models by interaction pattern. Flux is a streaming model with model-native turn detection. Nova-3 supports live and prerecorded transcription without that turn-detection behavior. That distinction is useful when endpointing is part of the evaluation rather than an afterthought.
Current Deepgram pricing lists promotional streaming rates of $0.0065 per minute for English Flux and $0.0048 per minute for mono Nova-3. Prerecorded mono Nova-3 starts at $0.0043 per minute. Multichannel use and speech add-ons can change the bill.
Deepgram is a close substitute when the team wants a modular speech layer behind its own orchestration. Adopting Deepgram's wider Voice Agent API changes that boundary, so assess it as a platform migration rather than as the same speech-to-text swap.
3. AssemblyAI: transcription with optional understanding models

AssemblyAI separates its transcription products by API. The Realtime Speech-to-Text API offers Universal-Streaming and Universal-3.6 Pro Realtime. The Pre-recorded Speech-to-Text API offers Universal-3.5 Pro.
Its Speech Understanding API analyzes transcripts for outputs such as summaries, entity detection, sentiment, and speaker identification. These models are separately priced and fit pre-recorded or post-call workflows rather than live transcription.
AssemblyAI pricing lists Universal-Streaming at $0.15 per hour and Universal-3.6 Pro Realtime at $0.45 per hour under the Realtime Speech-to-Text API. Universal-3.5 Pro is the pre-recorded model at $0.21 per hour. Streaming charges use total session duration. Optional services are billed separately, including speaker diarization at $0.12 per hour; Speech Understanding models have their own rates.
This option fits teams that need a managed transcription API plus separately priced post-call enrichment for recordings or completed live sessions. Be precise about which outputs are requirements. Buying diarization or analysis by default can make an attractive base rate irrelevant.
4. Google Cloud Speech-to-Text: multiple processing modes in Google Cloud

Google Cloud Speech-to-Text provides synchronous, asynchronous, and streaming recognition, plus language, adaptation, diarization, and automatic punctuation features. Dynamic batch offers slower asynchronous processing at a lower price, which creates a useful cost lever for backlogs that do not need immediate results.
For V2 standard recognition, Google Cloud pricing starts at $0.016 per minute for the first 500,000 minutes each month. Dynamic batch is $0.003 per minute. Successful requests are rounded to one-second increments, and each audio channel is billed separately.
This is a practical shortlist choice when Google Cloud already owns storage, identity, monitoring, and procurement. It deserves extra cost scrutiny for stereo or multichannel archives because billed channel-minutes can diverge from the duration displayed in a media player.
5. Azure Speech: cloud transcription with container options

Azure Speech supports real-time, fast, and batch transcription. Custom Speech can adapt recognition for domain vocabulary and acoustic conditions. Microsoft also supports speech containers for teams whose deployment or data path cannot rely entirely on a public cloud endpoint.
Azure's current generally available speech-to-text REST API is version 2025-10-15. Azure pricing varies by region, usage tier, and agreement, so there is no honest single global per-minute figure for the comparison table. Container use also introduces host infrastructure and licensing conditions alongside the API economics.
Azure belongs on the shortlist when the application already uses Microsoft identity, networking, observability, or procurement. Customization and container deployment can justify the integration work, but they do not replace a representative audio test.
6. OpenAI: hosted file and live transcription models

OpenAI now directs file transcription to gpt-transcribe and ongoing microphone or call audio to gpt-live-transcribe. The speech-to-text guide accepts common audio formats and caps file uploads at 25 MB, so larger assets need segmentation or another ingestion path.
OpenAI's API pricing estimates gpt-transcribe at $0.0045 per minute and gpt-live-transcribe at $0.017 per minute. These estimates derive from audio-token usage, which means the final invoice is usage based rather than a fixed audio-minute tariff.
There is also a migration deadline hidden by the familiar Whisper name. OpenAI's deprecation schedule lists whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-transcribe-diarize for removal on February 26, 2027. New integrations should use the stated replacements instead of treating the older hosted Whisper endpoint as the long-term default.
7. Self-hosted Whisper: model control with an operating burden

OpenAI released Whisper code and model weights under the MIT license. The repository provides multilingual speech recognition, translation, language identification, several model sizes, and published memory guidance. Its reference implementation processes complete files in sliding 30-second windows.
There is no managed API charge when you run Whisper yourself. Graphics processing unit (GPU) time, capacity planning, autoscaling, observability, model packaging, security patching, and on-call ownership become part of the cost instead. A third-party streaming wrapper can lower perceived delay, but that wrapper is another production system your team must validate and operate.
Self-hosting is justified when deployment control, data locality, or model access outweighs that work. It is less attractive when the team mainly needs a supported endpoint with predictable scaling. Keep self-hosted Whisper separate from OpenAI's hosted transcription API in both architecture diagrams and procurement models.
Compare accuracy on the work your product performs
A single public word error rate cannot settle this purchase. Build a fixed evaluation set from the audio the system will actually receive, with appropriate consent and data handling. Include the codecs, sample rates, telephone routes, microphones, languages, accents, background noise, overlap, long silences, and domain terms that appear in production.
| Measure | What to record | Why it matters |
|---|---|---|
| Word error rate | Insertions, deletions, and substitutions against a human transcript | Provides a repeatable overall baseline |
| Critical-entity error | Names, dates, amounts, addresses, account terms, and product codes | A low average error rate can still break the business task |
| Partial stability | How often interim words change before the final transcript | Unstable partials make live UI and agent logic harder to use |
| End-of-turn behavior | False endpoints, missed endpoints, and time from final speech to final result | Determines whether a conversation feels interruptible and responsive |
| Speaker attribution | Diarization or channel assignment errors | Affects call review, summaries, compliance, and analytics |
| Slice-level quality | Results by route, language, noise, device, and call type | Prevents a high-volume easy segment from hiding a failing segment |
Score every candidate on the same files and live replays. For real-time systems, capture the full event stream rather than only the final text. Log first partial, stable partial, endpoint event, final transcript, reconnect behavior, and any dropped audio. Set acceptance gates for the critical slices before negotiating price.
Model the total cost of the same workload
Start with monthly audio minutes, then calculate the unit that each vendor actually bills. A useful cost model includes:
- recognition minutes or audio tokens;
- separately charged channels, session duration, and idle streaming time;
- diarization, redaction, custom models, analysis, and support;
- storage, network egress, queues, and retry traffic;
- human correction and failed-task handling;
- orchestration and monitoring services around a raw API;
- GPU capacity and ML operations for self-hosting;
- telephony, language-model usage, and the runtime for a voice agent.
The lowest headline transcription rate can lose once two-channel billing, enrichment, or operational labor is included. A wider managed runtime can cost more per connected minute while removing services that appear elsewhere in the architecture budget. Compare the completed workload, not the label on one line item.
Migrate without tying application logic to the next API
- Freeze a baseline. Store consented test audio, reference transcripts, current outputs, latency traces, and cost assumptions before changing providers.
- Put events behind an adapter. Convert vendor-specific partials, finals, timestamps, speaker labels, confidence values, errors, and reconnect signals into an internal schema.
- Rebuild dependent features deliberately. Custom vocabulary, personally identifiable information redaction, language detection, diarization, and Call Analytics do not have identical replacements across vendors.
- Shadow traffic where policy permits. Send a controlled share of audio through both paths and compare outputs without changing the customer experience. Apply the same privacy, retention, and access controls to the second path.
- Cut over one slice at a time. Start with a language, call type, or batch queue that passed its gate. Keep the adapter and a tested rollback path until error and cost trends are stable.
This approach preserves leverage. The next model or vendor change becomes an adapter and validation project rather than a rewrite of business logic.
Frequently asked questions
Is Whisper better than Amazon Transcribe?
They expose different tradeoffs. Amazon Transcribe is a managed AWS service with batch, streaming, security integration, and optional call analysis. Self-hosted Whisper provides model and deployment control while leaving inference operations to your team. The better choice is the one that clears your accuracy, latency, privacy, and total-cost gates on representative audio.
Is there a free Amazon Transcribe alternative?
Self-hosted Whisper has no API license fee under its MIT license, but it still consumes compute and engineering time. Managed services may offer free quotas or credits with limits. A free tier is useful for integration work, while a production decision needs the steady-state bill.
Is Dasha a drop-in replacement for Amazon Transcribe?
No. We replace a wider live voice-agent runtime. Dasha is relevant when transcription feeds turn-taking, application logic, telephony, integrations, and operations for a voice agent. A batch or file-only pipeline should use a direct speech API or a self-hosted model instead.
How should we compare transcription accuracy?
Use a consented, labeled sample from production conditions. Measure overall word error rate, critical entities, speaker attribution, and each important language, route, and noise slice. For streaming, also measure partial-result churn, endpoint errors, time to final text, and recovery after reconnects.
If Amazon Transcribe is one component inside a production voice agent, start with Dasha and run one real call path against the acceptance gates your team already uses.



