“Vertex AI Speech” is ambiguous. In Google’s public cloud, transcription, synthesis, and live multimodal audio are separate services. Google Distributed Cloud also uses the Vertex AI Speech-to-Text name. A useful comparison therefore starts with three buyer jobs: component speech APIs, native speech-to-speech APIs, and managed voice-agent runtimes. The right alternative depends on which layer you need to replace and which production responsibilities your team wants to keep.
First decide which layer you are replacing
The closest alternative depends on the job. A component speech application programming interface (API) changes audio into text or text into audio. A native speech-to-speech API accepts live audio, maintains a model conversation, and returns audio. A managed voice-agent runtime also owns more of the session lifecycle, channel integration, testing, inspection, and production operation.
These boundaries matter because the products do not expose equivalent units of work:
- Component speech APIs give you speech-to-text (STT), also called automatic speech recognition (ASR), or text-to-speech (TTS). Your team assembles the conversational pipeline and operates it.
- Native speech-to-speech APIs give you a stateful, direct-audio model session with turn handling and tools. Your team still owns the surrounding application and production controls.
- Managed voice-agent runtimes coordinate the live session, agent configuration, tools, channels, and operational evidence. Your team retains the business workflow, authoritative data, permissions, compliance decisions, and acceptance criteria.
Google spans these layers through separate products. Cloud Speech-to-Text handles transcription, Cloud Text-to-Speech handles synthesis, and Gemini Live API handles stateful live multimodal interaction. In Google Distributed Cloud air-gapped deployments, Google also documents a Vertex AI Speech-to-Text API. That deployment-specific name should not be used as an umbrella for every Google speech and voice service.
Vertex AI speech alternatives at a glance
| Option | Buyer job and layer | Interface or deployment boundary | Your team still operates | Official source |
|---|---|---|---|---|
| Dasha | Run production voice agents through a managed runtime and operational surface | Web application and REST APIs, with Session Initiation Protocol (SIP) and web voice or chat paths | Workflow rules, business data, authorization, carrier configuration, compliance, and acceptance criteria | Dasha runtime boundary |
| Google Cloud Speech-to-Text and Text-to-Speech | Transcribe or synthesize audio with component APIs | Separate recognition and synthesis APIs | Channels, model orchestration, tools, state, testing, observability, and recovery | Chirp 3 documentation and Cloud TTS basics |
| Gemini Live API | Build a native live audio or multimodal interaction | Stateful bidirectional session with audio input and output, interruption handling, and function calling | Secure client path, business tools, application state, telephony, and production operations | Gemini Live reference |
| OpenAI Realtime API | Build a native speech-to-speech agent | Direct-audio sessions over Web Real-Time Communication (WebRTC) in browsers or WebSocket on servers, with conversation state and tools | Application control, business systems, channel policy, reliability, and operational evidence | Realtime API guide |
| Deepgram | Replace ASR or TTS components, or use a managed conversational pipeline | Separate streaming STT and TTS APIs, or one Voice Agent API WebSocket that combines STT, a large language model (LLM), and TTS | Channel edge, business state, tool safety, deployment controls, and the operations outside the API | Streaming STT guide, TTS guide, and Voice Agent guide |
| AssemblyAI | Transcribe audio, analyze transcripts, or configure a hosted voice agent | Separate recorded and streaming STT, transcript analysis, and Voice Agent API surfaces | Product policy, business data, permissions, carrier account, and any operating controls outside the service | Model overview, Speech Understanding, and Voice Agent API |
| Microsoft Azure Speech | Use component STT or TTS, or a managed live voice interface | Separate Azure Speech APIs plus Voice Live API, which combines recognition, generative AI, and synthesis | Client and channel integration, business systems, tool policy, and surrounding production operations | Speech-to-text documentation, text-to-speech overview, and Voice Live overview |
This table is an ownership map, not a scorecard. A streaming transcription API cannot be declared faster than a complete voice-agent runtime without defining the channel, audio, region, endpointing policy, model path, tool work, playback point, and measurement window. The same problem applies to accuracy and voice quality claims. Compare equivalent boundaries on the same workload.
1. Dasha: a managed voice-agent runtime and operational surface

Dasha is the first option to evaluate when your real replacement target is the orchestration and operating layer around a production voice agent. It is not a drop-in replacement for Chirp or Cloud Text-to-Speech.
We provide a managed real-time runtime, REST APIs, and a web application. Current product surfaces cover SIP and web voice or chat connections, external tools, testing, call inspection, activity and call history, queue and concurrency monitoring, and SIP traces. This gives a technical team one place to run and examine live agents while keeping the product workflow and business systems under its control.
The ownership split is explicit:
- Dasha owns: managed session execution, agent configuration and deployment surfaces, supported channel connections, tool invocation plumbing, and runtime evidence exposed by the platform.
- Your team owns: the workflow and business rules, systems of record, tool authorization, carrier configuration, consent and compliance policy, data retention requirements, and acceptance criteria.
That boundary fits a SaaS or product team that wants to stop maintaining the live runtime and its operational seams. A component API is the better fit when you only need transcription or synthesis. A custom stack is the better fit when source-level control of media buffers, codecs, transport behavior, or every runtime component is a product requirement.
Migration scope follows the layer you adopt. Moving away from a managed runtime means replacing its session, channel, configuration, and inspection contracts. Keeping authoritative data and business actions behind your own tools reduces the amount of domain logic tied to that runtime.
2. OpenAI Realtime API: a native speech-to-speech model session

OpenAI Realtime API belongs in the native speech-to-speech layer. The model works directly with audio, retains conversation state, and can call tools. Browser clients can use WebRTC, while server applications can use WebSocket connections. The session handles turns, interruptions, and handoffs.
This boundary can simplify a voice experience because the model session owns more of the audio conversation than a separately chained STT, LLM, and TTS pipeline. It also changes what you can inspect and tune. A component chain exposes separate recognition output, language-model events, and synthesis events. A native session exposes the event and control surface defined by the Realtime API.
OpenAI Realtime is a fit when direct audio interaction is the requirement and your team is prepared to own the rest of the product: client authentication, channel and telephony policy, tool execution, business state, monitoring, retries, incident handling, and release controls. It does not by itself replace a managed production runtime.
Pilot it with interruptions during both speech output and a consequential tool call. Confirm that canceled work, spoken confirmations, and committed business state remain consistent.
3. Deepgram: separate speech components and a managed voice pipeline

Deepgram exposes products at two layers, so the integration choice has to be named.
Its streaming Speech-to-Text API is a component replacement for Google Cloud Speech-to-Text. Audio enters a streaming recognition connection and transcript events return to your application. Its Text-to-Speech API is a separate synthesis component that turns text into streamed audio. With either API, your application still supplies the agent model, turn policy, tools, state, channel handling, and operational controls.
Deepgram Voice Agent API has a different boundary. One WebSocket combines STT, an LLM, and TTS into a managed conversational stream. The application configures the agent and sends audio, then receives transcripts, agent events, and audio output. This is a cascaded voice pipeline behind one interface. It should be compared with other managed conversational APIs, rather than with a standalone ASR model.
The component route fits teams that want to choose and operate each stage. The Voice Agent route fits teams that want one live pipeline contract while keeping the surrounding channel, business systems, tool safety, deployment process, and incident response. Switching between those routes changes the interface and operating burden even when both carry the Deepgram name.
4. AssemblyAI: transcription, transcript intelligence, and a separate agent API

AssemblyAI also spans more than one layer. Its recorded and streaming STT models are component speech services. Speech Understanding runs on transcripts to extract structured information such as entities, topics, sentiment, action items, and summaries. That is useful for call analytics, meeting products, and post-call workflows. It is not a TTS component.
AssemblyAI now documents a separate Voice Agent API for real-time agents deployed to a browser, application, or phone. The product surface includes agent configuration, tools, turn-taking and interruptions, a Twilio SIP path, tests, recordings and transcripts, and webhooks. Teams considering it for live execution should evaluate that API boundary separately from AssemblyAI transcription.
Older comparisons often recommend LeMUR for transcript reasoning. That advice is obsolete. AssemblyAI ended LeMUR after March 31, 2026 and directed users to LLM Gateway for continued language-model access, according to its LeMUR deprecation notice.
Use AssemblyAI's component APIs when transcripts and transcript-derived outputs are the main product artifact. Evaluate its Voice Agent API as a managed voice-agent API when live interaction is the job. Those are different architecture decisions with different ownership and migration costs.
5. Microsoft Azure Speech: component APIs plus Voice Live

Microsoft offers the closest structural parallel to Google's multi-product speech stack in this set.
Azure Speech-to-Text supports real-time and batch transcription. Azure Text-to-Speech is a separate synthesis service available through the Speech SDK and REST APIs. These component services leave the conversational pipeline, channels, tools, state, deployment, and operations with your team.
Voice Live API moves up a layer. Microsoft describes it as a managed speech-to-speech interface that combines speech recognition, generative AI, and text-to-speech. The client sends audio and receives audio output and action triggers through one managed service. This removes manual component orchestration inside the session, while your application continues to own the channel integration, tool authorization, business state, and production policy around it.
Azure therefore needs two evaluations. Compare its STT and TTS services with Google's corresponding component APIs. Compare Voice Live with Gemini Live, OpenAI Realtime, Deepgram Voice Agent, AssemblyAI Voice Agent, or a managed runtime according to the amount of session and operational ownership each one assumes.
When staying with Google's stack is the right decision

An alternative is useful only when it improves a requirement that matters to your product. Google's current boundaries already cover several distinct jobs:
- Cloud Speech-to-Text supports streaming, synchronous, and batch recognition methods. Model, method, feature, locale, and region availability do not form one universal support list.
- Cloud Text-to-Speech converts text or Speech Synthesis Markup Language (SSML) to audio through a separate API.
- Gemini Live API provides live multimodal sessions, audio output, interruptions, and function calling.
- Google Distributed Cloud has its own Vertex AI Speech-to-Text documentation and deployment constraints.
Keeping Google is reasonable when those interfaces meet your acceptance criteria and your team can sustain the orchestration and operating work around them. Replacing one component can also be enough. A transcription problem does not automatically justify moving the model, synthesis, telephony, and runtime layers.
Run a pilot that compares equivalent boundaries
A vendor demo proves that one prepared path can work. A production pilot should expose the ownership boundary and failure behavior you will live with.
1. Freeze the test contract
Use the same audio corpus, languages, accents, noise conditions, region, codecs, channels, business tools, and expected outcomes. Record the exact API, model, configuration, prompt, and voice used for each run.
2. Measure the job you are buying
- For STT: measure critical-entity accuracy, final transcript accuracy, interim-result stability, endpointing behavior, diarization where required, and recovery from dropped audio.
- For TTS: measure pronunciation of domain terms, intelligibility, time to playable audio, streaming continuity, and how quickly canceled speech stops reaching the user.
- For native speech-to-speech: measure task completion, interruption behavior, state consistency, tool-call correctness, and end-to-end time from the end of relevant user speech to audible response.
- For managed runtimes: add channel setup, transfer behavior, version control, queueing, retries, call inspection, operational alerts, failure recovery, and tenant isolation.
3. Inject production failures
Interrupt the agent during generation, playback, and tool execution. Force a tool timeout, duplicate webhook, partial database write, provider error, and unavailable region. Confirm that the spoken result matches the committed business result and that one trace reconstructs the failure.
4. Map ownership and exit cost
List who owns phone numbers, trunks, prompts, agent configuration, recordings, transcripts, tool schemas, business data, traces, and evaluation fixtures. Then identify what must be rewritten if you change the speech component, model session, or runtime. A provider switch is manageable when the contract boundary and retained assets are clear.
5. Choose against acceptance criteria
Weight the results according to your product. A transcription service can win an ASR evaluation and still be the wrong answer for a team trying to remove runtime operations. A managed runtime can reduce operating work and still be unnecessary for a batch transcription product.
If runtime operation is the layer you want to stop owning, use the Dasha evaluation path to run one real call with one business tool. Interrupt the agent and force a tool timeout, then inspect the resulting call evidence before expanding the pilot.



