OpenAI Realtime API alternatives: 5 production options

Direct real-time voice APIs and managed runtimes compared
Direct real-time voice APIs and managed runtimes compared

Choosing an OpenAI Realtime API alternative starts with an architecture decision. A direct speech-to-speech API can replace the model session. A managed runtime takes on telephony, tools, testing, and production operations too. Technical teams need to compare those boundaries before comparing voices or demo latency.

The short answer

The strongest OpenAI Realtime API alternative depends on what you want to replace:

  • Choose Dasha when you want a managed production runtime for phone and browser voice, with REST APIs, tools, testing, call inspection, and operational surfaces.
  • Choose Gemini Live API when native audio plus live image or video input matters and a preview Google API fits your release policy.
  • Choose Inworld Realtime API when you want an OpenAI-shaped event model with separate speech, language-model, and voice choices.
  • Choose xAI Speech to Speech when Grok, built-in search tools, and a relatively small OpenAI-protocol migration are the priority.
  • Choose Hume EVI when vocal expression, prosody-aware turn timing, and response style are central to the product.

Vapi, Retell AI, and ElevenLabs Agents also deserve consideration if you want a hosted orchestration platform. LiveKit Agents, Pipecat, or a custom stack fit teams that want source-level control. Those products sit at different layers, so treating all of them as direct API substitutes creates the wrong comparison.

First decide which layer you are replacing

The OpenAI Realtime API is a stateful model-session API. It accepts live audio, maintains conversation state, calls tools, and returns audio over WebRTC or WebSocket. Session Initiation Protocol (SIP) is available for phone connections. Your application still owns the wider workflow, production policy, monitoring, recovery, and business state.

OpenAI's current audio and voice guidance recommends GPT-Live as the starting point for a new conversational voice application. It keeps Realtime as the option for teams that need its session and tool model. If your only issue is which OpenAI voice architecture to use, that is an internal architecture change rather than a vendor migration. Our OpenAI Realtime API guide covers the implementation details.

External alternatives fall into three groups:

  1. Direct real-time voice APIs expose a live speech session. Gemini Live, Inworld, xAI, and Hume belong here, although their internal architectures and protocols differ.
  2. Managed runtimes and agent platforms own more of the live path and operating surface. Dasha, Vapi, Retell AI, and ElevenLabs Agents belong here.
  3. Frameworks and custom stacks give your team code-level control over media, models, and orchestration. LiveKit Agents and Pipecat belong here, with optional managed-cloud services.

The boundary changes the migration. Replacing OpenAI with another direct API means remapping events, audio formats, tools, session state, and model behavior. Moving to a managed runtime also moves operating responsibilities. Moving to a framework gives your team more of them.

OpenAI Realtime API alternatives compared

OptionProduct layer and architectureConnections and channelsModel and voice controlWhat your team still ownsMain tradeoff
DashaManaged production runtime and agent platformSIP telephony; WebRTC browser voice with WebSocket signalingSupported model, speech, voice, and tool configurationBusiness rules, tool authorization, data, carrier account, compliance, acceptanceBroader operating surface, but no drop-in OpenAI event endpoint or self-hosted open-source runtime
Gemini Live APINative multimodal model APIStateful WebSocket for audio, video, images, text, and audio outputGoogle Live models and voicesChannel integration, telephony, workflow state, observability, recoveryStrong voice-plus-vision path; Live API remains preview and uses Google's protocol
Inworld Realtime APIModular speech-to-text, language model, and text-to-speech pipeline behind one session APIWebSocket is generally available; WebRTC and SIP are early accessMultiple language models and Inworld voicesBusiness workflow, application state, carrier integration where native SIP access is unavailable, production operationsProvider choice and familiar events; migration still changes authentication and session fields, and maturity differs by transport
xAI Speech to SpeechGrok speech-to-speech APIWebSocket, WebRTC examples, and SIP telephonyGrok voice models, built-in or custom voicesWorkflow policy, application state, monitoring, and recoveryClose OpenAI protocol shape and rich built-in tools; reasoning stays tied to Grok
Hume EVIProsody-aware speech-language and voice interfaceReal-time WebSocket; Twilio integration is availableHume voices plus supported or custom language modelsChannel and application integration, operating controls, acceptanceStrong expression and timing controls; proprietary protocol, session limits, and supplemental-model conditions

Pricing also follows the layer. Direct APIs may bill tokens, session time, messages, or separate speech and model components. Managed platforms add a runtime charge and may leave carrier or provider usage outside it. Frameworks replace some vendor charges with infrastructure and engineering. Compare cost per successful task, including retries and failed calls, rather than a headline rate.

Dasha: managed runtime plus production operations

Dasha voice AI backend evaluation framework for agent workflows, load testing, and customer configuration

Best for: Technical teams that want to delegate the live runtime and its operational surface while keeping control of workflows, tools, and business data.

We built Dasha as a managed production platform rather than an OpenAI-compatible model endpoint. The current product combines a real-time voice runtime, REST APIs, a web application, SIP and browser channels, webhook and Model Context Protocol (MCP) tools, testing, call history, activity logs, and completed-call inspection. Browser voice uses WebRTC with WebSocket signaling.

That larger boundary is the reason to choose Dasha. Your team can focus on agent behavior and business systems while we run the live conversational path and provide operating surfaces around it. You still own authoritative data, tool permissions, carrier accounts and phone configuration, compliance decisions, and production acceptance. Our voice AI infrastructure guide maps that responsibility split in detail.

Tradeoff: Dasha is a managed runtime. It does not expose the OpenAI Realtime event contract as a drop-in endpoint, and it does not fit a team that requires source-level ownership of the media runtime. Migration means recreating the agent configuration and tools, connecting phone or web channels, and rerunning production tests. It can remove application work, but it is a larger architectural change than swapping one WebSocket host.

Gemini Live API: native audio and vision in Google's stack

Gemini Live API technical specifications for streaming modalities and WebSocket sessions

Best for: Products where live camera, screen, or image context belongs in the same session as voice.

Google's Gemini Live API streams audio, video, images, and text into Gemini and returns audio, text, and function-call events. It is a native multimodal path, so audio remains part of the model interaction instead of being forced through a separate transcript-first pipeline. Barge-in, tool use, and transcriptions are available within the Live session.

The API uses a stateful WebSocket and Google's BidiGenerateContent messages. That makes it a fresh integration for an OpenAI client. You need to remap session setup, content parts, tool calls, turn events, authentication, and reconnection behavior. Telephony and the wider agent runtime remain your responsibility.

Tradeoff: Gemini Live offers a broader native input mix than a voice-only API. The WebSocket reference still labels the Live API as preview and uses a v1beta endpoint. Teams with a general-availability requirement need to account for that status. Teams that need OpenAI event compatibility or independent model and voice providers should choose another path.

Inworld Realtime API: an OpenAI-shaped modular pipeline

Inworld migration table comparing Realtime session endpoints and configuration fields

Best for: Teams that want to preserve much of an OpenAI-style client while gaining language-model and voice choice.

Inworld's Realtime API uses an OpenAI-shaped session and event model around a modular speech-to-text, language-model, and text-to-speech pipeline. It preserves familiar structural events such as session.update, conversation.item.create, and response.create, while placing the language model and speech output model in separate configuration fields.

That distinction matters. A modular pipeline can route the reasoning step to different supported models and expose independent voice configuration. It also creates explicit component boundaries that OpenAI's native speech-to-speech model does not have. The components are streamed and coordinated behind the API, so this is different from a naive sequential pipeline.

Tradeoff: Protocol familiarity reduces client work, but compatibility does not mean zero-change. Inworld's migration guide shows changes to the endpoint, authentication, model placement, voice configuration, and voice activity detection. The overall Realtime API carries a research-preview label, while voice-agent transport details list WebSocket as generally available and WebRTC and SIP as early access. Treat transport and support tier as separate release gates. If native SIP access is unavailable, your team also needs a carrier or media bridge and owns its lifecycle.

xAI Speech to Speech: Grok with OpenAI protocol compatibility

xAI Speech to Speech WebSocket quickstart in the developer documentation

Best for: Teams that want Grok voice reasoning, server-side search tools, and a short migration from an OpenAI-shaped event loop.

xAI's Speech to Speech API streams audio and text bidirectionally and supports function tools, web search, X search, file search, and remote MCP servers. Its documentation covers WebSocket clients, WebRTC examples, and SIP phone calls. The API also supports conversation resumption and production model pinning.

xAI documents an OpenAI migration based on changing the endpoint, key, and model. The same documentation lists event-name differences, events with transport-specific behavior, and unsupported events. That makes it a small protocol migration compared with Gemini or Hume, while still requiring a compatibility test against the exact events your client consumes.

Tradeoff: The language and reasoning layer stays with Grok. Duration-based audio billing also includes separate rules for text events, so cost modeling must reflect your actual interaction pattern. The current model page lists a specific serving region and default session limits. Review capacity and data-location requirements before treating a working development connection as a production approval.

Hume EVI: prosody-aware voice interaction

Hume EVI documentation showing speech-to-speech capabilities and version comparison

Best for: Coaching, companion, accessibility, game, and support experiences where vocal expression and response style affect the product outcome.

Hume's Empathic Voice Interface runs a real-time WebSocket session and uses vocal prosody, including rhythm, tune, and timbre, to guide turn timing and response delivery. It supports configurable voices, prompts, turn detection, interruptions, versioned configurations, webhooks, and supported or custom language models.

Prosody is useful interaction evidence. It is not ground truth about a person's internal emotional state. Evaluate whether the resulting timing and response style improve the task rather than treating an expression label as a fact about the user.

Tradeoff: EVI uses its own chat protocol and configuration model, so an OpenAI client needs a new integration. Hume's tool documentation ties tool use to supported supplemental or custom language models. Published EVI sessions also have a 30-minute maximum. Hume is a focused voice-interface choice rather than a complete production operations layer.

When an orchestration platform or framework is the better alternative

A direct speech API is only one possible answer. Use another architecture when the requirement is broader provider choice, a ready-made phone-agent surface, or source-level media control.

  • Vapi is a hosted orchestration option with a broad model provider catalog, bring-your-own provider credentials, SIP, simulations, and text-layer evals. It fits teams that want to swap speech and model components without running the orchestration service. Your production system still depends on Vapi plus the selected providers. Our OpenAI Realtime versus Vapi guide goes deeper into that boundary.
  • Retell AI combines hosted voice and chat agents with managed numbers, SIP, web calls, tools, testing, call records, and monitoring. It is a platform migration rather than an OpenAI-protocol swap. It fits phone-heavy teams that value an integrated hosted surface.
  • ElevenLabs Agents combines ElevenLabs speech with telephony, web and mobile clients, tools, conversation history, and agent testing. It fits teams that prioritize the ElevenLabs voice ecosystem. The core speech path remains tied to that ecosystem even when the language model is configurable.
  • LiveKit Agents and Pipecat are code-first frameworks with provider integrations and optional managed deployment. LiveKit Agents includes WebRTC rooms, SIP, agent servers, and Cloud deployment. Pipecat provides real-time pipeline components, transports, and Pipecat Cloud. They fit teams that need source-level media and orchestration control and can own the remaining deployment, integration, observability, and incident work.

For a wider platform catalog, use our voice API comparison. For the media decision inside any option, compare WebRTC and WebSocket separately from the model or runtime.

A practical seven-step pilot

Run every finalist against the same production-shaped workload. A pleasant demo call is a starting point, not a release gate.

  1. Define one complete task. Pick a booking, support, qualification, or account workflow with a measurable outcome, one read tool, one write tool, and a human handoff. Record the authoritative success state in your business system.
  2. Fix the comparison boundary. Use the same caller audio, phone route or browser client, tool backend, data, consent rules, and success criteria. Document which layers each vendor replaces so the pilot does not credit a direct API for work your application performs.
  3. Pin every configuration. Record the model, voice, prompt, tool schema, turn settings, transport, region, and software version. Keep these settings with every result.
  4. Run a conversation matrix. Repeat clean and noisy audio, narrowband phone audio, accents relevant to your users, long pauses, backchannels, corrections, spelled identifiers, mid-sentence interruptions, and overlapping speech.
  5. Inject failures. Delay and fail a tool, return malformed data, disconnect the client, exhaust a quota, trigger a transfer failure, and reconnect a session. Confirm that the agent never reports an action as complete before the system of record commits it.
  6. Measure outcomes and operations. Track task success, critical-entity accuracy, median (P50) and 95th-percentile (P95) time from user endpoint to first useful audio, false interruptions, recovery success, tool correctness, operator time to diagnose a failed call, and cost per successful task. Include carrier, model, speech, platform, storage, monitoring, and engineering costs.
  7. Apply hard gates first. Eliminate options that fail channel, regional, security, recovery, or accuracy requirements. Use weighted preferences such as voice style or migration effort only among the options that pass.

Our voice agent testing guide provides a fuller regression plan. The same test set should run before launch and after every model, prompt, voice, tool, or turn-policy change.

Choose the responsibility boundary before the vendor

Choose Gemini, Inworld, xAI, or Hume when the model-session layer is the product decision and your team is ready to operate the surrounding agent system. Choose Vapi, Retell AI, or ElevenLabs Agents when hosted orchestration and their supported operating surfaces match your requirements. Choose LiveKit, Pipecat, or a custom stack when low-level control justifies the engineering burden.

Choose Dasha when you want the alternative to take on more than the speech session. Start a Dasha evaluation with one real call path, one business tool, one interruption, and one forced tool failure. That pilot will show whether our managed runtime is the right boundary for your product.

Share

Subscribe

Sign up to our e-mail list to get the best of the Dasha blog sent directly to your inbox.

We use cookies for functional and analytical purposes. Please refer to our Privacy Policy for details.