8 Amazon Polly alternatives for TTS and voice AI in 2026

A technical decision-maker compares cloud text-to-speech, audio editing, and voice-agent options.
A technical decision-maker compares cloud text-to-speech, audio editing, and voice-agent options.

Replacing Amazon Polly is easy only when the job begins and ends with text-to-speech. A production voice product may also need speech recognition, turn-taking, an LLM, telephony, tool calls, monitoring, and evaluation. The useful question is which layer you want to replace. That choice narrows the field and makes cost, latency, and migration comparisons meaningful.

What an Amazon Polly alternative actually replaces

Amazon Polly is a managed text-to-speech (TTS) service. It accepts text and returns audio. It also supports streaming, Speech Synthesis Markup Language (SSML), custom lexicons, and speech marks for synchronizing words or visemes with audio. Polly's feature list documents those capabilities.

Alternatives fall into four product layers:

  1. Hosted TTS APIs such as ElevenLabs TTS, Google Cloud Text-to-Speech, Azure Speech, Deepgram TTS, and Cartesia TTS replace the synthesis endpoint.
  2. Creator suites such as Murf Studio pair a separate studio product with a synthesis API.
  3. Self-hosted models such as Chatterbox models give you model code and weights while leaving inference, scaling, and monitoring to your team.
  4. Voice AI runtimes such as the Dasha runtime replace more of the conversational stack, including call execution, orchestration, testing, and operations.

Dasha is our recommended choice when Polly sits inside a real-time voice agent. It is the wrong layer for batch narration or a product that only needs audio files. For those workloads, a hosted TTS API or self-hosted model is a closer substitute.

When keeping Amazon Polly makes sense

Polly remains a sensible choice when you already run on AWS, rely on its speech marks or custom lexicons, and prefer a small migration surface over new voice features. Published US rates are $4 per one million characters for Standard voices, $16 for Neural, $30 for Generative, and $100 for Long-Form. Standard voices are still among the lowest-priced managed options in this comparison. AWS publishes each rate by engine.

The free boundary depends on when the AWS account was created. Accounts created on or after July 15, 2025 use the credit-based program: $100 at sign-up plus up to $100 more for completing activities. The Free plan ends after six months or when credits run out, whichever comes first. AWS explains the current program. Credits must be used within 12 months of account creation. Polly's pricing page states that expiry. For eligible accounts under the earlier service-specific program, the same page lists monthly allowances for the first 12 months: five million Standard, one million Neural, 500,000 Long-Form, and 100,000 Generative characters.

Polly can also return an audio stream. A switch motivated by the belief that Polly can only generate files solves a problem that does not exist. Consider moving when you need a different voice catalog, self-service voice cloning, a creator workflow, self-hosting, or a complete production runtime around the voice.

Amazon Polly alternatives at a glance

Every character-based TTS rate below is normalized to one million characters. Subscription credits, generated-audio minutes, telephony minutes, and agent-runtime minutes remain separate because they measure different products.

AlternativeProduct layerFitsAPI and streamingPublic starting priceFree or self-hosted boundary
DashaManaged voice AI runtimeProduction conversational voice productsREST APIs and managed real-time runtimeGrowth from $0.08 per agent-runtime minute, excluding VoIP and LLM tokens (Dasha pricing)Developer plan includes 1,000 agent-runtime minutes and one concurrent call (Dasha pricing)
ElevenLabsHosted TTS API and creator toolsExpressive narration, localization, and cloned voicesStreaming TTS across current v2 and v3 families (TTS models)$50 per one million characters for v3 Conversational and Flash/Turbo; $100 for v3 and v2 Multilingual (API rates)Free/PAYG includes 20,000 characters for v3 Conversational or Flash/Turbo, or 10,000 for v3 or v2 Multilingual; commercial use requires a paid plan (commercial terms)
Google Cloud TTSHosted TTS APIA hyperscaler alternative with several price and control tiersStreaming on selected models (voice types)$4 per one million characters for Standard and WaveNet; $16 for Neural2; $30 for Chirp 3 HD (Google pricing)Monthly allowances are four million Standard or WaveNet characters and one million Neural2 or Chirp 3 HD characters (Google pricing)
Azure SpeechHosted speech suiteMicrosoft cloud stacks and governed custom-voice projectsReal-time neural synthesis APIs (Azure overview)Paid TTS is billed per character; rates vary by region and voice type (Azure pricing)500,000 neural characters free each month (Azure pricing)
DeepgramHosted TTS API and separate agent APIReal-time voice applications that may also use Deepgram speech recognitionStreaming TTS and a WebSocket Voice Agent API (agent guide)Aura-1 $15 and Aura-2 $30 per one million characters; Flux TTS $45 beginning September 13, 2026 (Deepgram pricing)$200 introductory credit; Flux TTS is free through September 12, 2026, with up to 45 global streams or five in EU/AU (Deepgram pricing)
CartesiaHosted TTS API and separate agent runtimeStreaming TTS with a credit-based planWebSocket TTS; separate Line agent product (Cartesia pricing)Pro is a $5 monthly subscription with 100,000 credits; Line is $0.06 per agent-runtime minute, plus $0.014 per telephony minute for a Cartesia number (Cartesia pricing)Free includes 20,000 monthly credits; a commercial-use license starts on Pro (Cartesia pricing)
MurfTTS API and separate creator studioReal-time synthesis plus batch media voiceoverFalcon 2 streams; Gen2 uses a non-streaming endpoint (API models)Falcon $10 and Gen2 $30 per one million characters (API pricing)API Free Trial includes $10 in monthly credit, five concurrent requests, and no card requirement (API pricing)
ChatterboxOpen-source model familyTeams that need control over deployment and inference costYou build and operate the serving layerMIT-licensed software; infrastructure costs remain (MIT license)Self-hosted on your own CPU or GPU capacity (official repository)

1. Dasha: for replacing the voice-agent stack around Polly

Dasha voice AI backend product page

Fits: Technical teams building production conversational voice products.

Layer and API: Dasha is a managed production runtime exposed through REST APIs and a web application. It covers telephony, tools, testing, monitoring, and call execution, replacing more operating surface than a TTS endpoint.

Price and free boundary: The Developer plan includes 1,000 agent-runtime minutes and one concurrent call. Growth starts at $0.08 per agent-runtime minute, excluding VoIP and LLM tokens. Compare our current pricing with the full Polly stack, including speech recognition, telephony, hosting, and orchestration.

Voice and deployment: Dasha is a managed cloud service for real-time conversation and production operations.

Main tradeoff: Dasha is a poor fit for audiobooks, accessibility narration, and simple text-to-audio jobs. It fits when orchestration and operations are also in scope.

Switching from Polly: Move call control, prompts, tools, telephony, tests, and monitoring along with voice generation. Our runtime and API stack comparison explains the scope.

2. ElevenLabs: for expressive speech and voice creation

ElevenLabs text-to-speech interface

Fits: Narration, dubbing, media localization, and products that need cloned or designed voices.

Layer and API: ElevenLabs offers hosted streaming TTS alongside creator and agent products. Eleven v3 emphasizes expressive output, v3 Conversational targets real-time speech, and Flash v2.5 is the lower-cost API model.

Price and free boundary: Current API list rates are $100 per one million characters for Eleven v3 and v2 Multilingual, and $50 for v3 Conversational and Flash/Turbo. The Free/PAYG tier includes 10,000 characters for v3 or v2 Multilingual and 20,000 for v3 Conversational or Flash/Turbo. ElevenLabs limits commercial usage rights to paid plans.

Voice and deployment: The managed cloud service includes voice libraries, voice design, and instant and professional cloning, with model-specific voice settings.

Main tradeoff: These models cost more per character than Polly Standard, and output is nondeterministic. Repeatable pronunciation and timing require a regression set.

Switching from Polly: Rewrite SSML-heavy content around prompting and voice controls. Migration acceptance criteria should cover pronunciation, chunking, concurrency, formats, and commercial rights. Our ElevenLabs alternatives guide covers adjacent products.

3. Google Cloud Text-to-Speech: for a close hyperscaler substitute

Google Cloud Text-to-Speech product page

Fits: Teams that want a managed cloud TTS API with price tiers that map closely to Polly.

Layer and API: Google Cloud offers Standard, WaveNet, Neural2, Chirp 3 HD, Studio, custom voice, and Gemini TTS options. Chirp 3 HD supports low-latency text streaming. Capabilities vary by model.

Price and free boundary: Google's TTS rates are $4 per one million characters for Standard and WaveNet after a four-million-character monthly allowance, $16 for Neural2 after one million, and $30 for Chirp 3 HD after one million. Instant Custom Voice is $60 per one million characters with no free allowance.

Voice and deployment: It is a managed Google Cloud API. Voice selection and controls depend on the model. Chirp 3 HD, for example, does not accept SSML or speaking-rate and pitch parameters. Google documents those limits.

Main tradeoff: Low-priced and newer models expose different controls, so compare the voice tier your workload needs.

Switching from Polly: Map voice IDs, authentication, request limits, codecs, and SSML behavior. Audit any dependency on Polly speech marks or custom lexicons before cutting over.

4. Azure Speech: for Microsoft cloud and custom-voice governance

Azure Speech in Foundry Tools product page

Fits: Azure-centered applications and organizations prepared for a governed custom-voice process.

Layer and API: Azure Speech provides neural TTS APIs plus personal voice, professional custom voice, and avatars. Real-time synthesis is available.

Price and free boundary: The free tier includes 500,000 neural characters each month. Paid text-to-speech is billed by character, with rates that vary by region and voice type. Custom voice training and endpoint hosting are separate charges.

Voice and deployment: Professional custom voice requires an approved use case, recorded consent from the voice talent, and at least 300 utterances for training. Access is limited rather than an instant self-serve feature.

Main tradeoff: Custom voice adds approval, dataset, training, and endpoint work.

Switching from Polly: Convert voice and SSML settings, with pronunciation, regions, quotas, and formats in the acceptance criteria. Treat custom voice as its own project rather than part of an endpoint swap.

5. Deepgram: when TTS and speech recognition share a vendor

Deepgram text-to-speech product page

Fits: Real-time applications that already use Deepgram speech-to-text or want a streaming speech stack from one provider.

Layer and API: Deepgram sells standalone Aura and Flux TTS. Its separate Voice Agent API combines speech recognition, LLM integration, and synthesis over one WebSocket.

Price and free boundary: Current pay-as-you-go pricing is $15 per one million characters for Aura-1 and $30 for Aura-2. Flux TTS is free through September 12, 2026, with up to 45 concurrent global streams or five in the EU/AU. On September 13, its pay-as-you-go rate becomes $45 per one million characters. New pay-as-you-go accounts include a $200 credit.

Voice and deployment: TTS is a managed streaming API. The Voice Agent API is billed separately by WebSocket connection minute, with the configured LLM and TTS affecting the total. Deepgram's rate table separates those connection-minute tiers from character-based TTS.

Main tradeoff: Standalone TTS prices are higher than Polly Standard. Moving both input and output speech to one provider also raises the cost of a later speech-vendor migration.

Switching from Polly: Standalone TTS is a contained move. The Voice Agent API also changes turn handling, LLM integration, WebSocket state, and observability.

6. Cartesia: for credit-based streaming TTS

Cartesia Sonic text-to-speech product page

Fits: Real-time applications that want streaming TTS and voice cloning under small self-serve plans.

Layer and API: Cartesia offers Sonic TTS through WebSocket and other synthesis endpoints. Line pricing treats the voice-agent product separately.

Price and free boundary: The Free plan includes 20,000 monthly credits. Pro is a $5 monthly subscription with 100,000 credits, a commercial-use license, and instant voice cloning. Line usage is a separate $0.06 per agent-runtime minute, plus $0.014 per telephony minute when using a Cartesia phone number.

Voice and deployment: The managed service adds professional cloning and higher concurrency on higher plans. Credit use and included generated-audio minutes depend on the selected model and plan. Cartesia's plan table publishes those allowances separately from Line minutes.

Main tradeoff: Credits complicate direct per-character comparisons. Production requests should also pin a supported model snapshot instead of a moving alias. Cartesia's model documentation explains the update behavior.

Switching from Polly: Replace authentication, voice IDs, streaming, and audio configuration. Rebuild Polly-specific SSML, lexicon, or speech-mark behavior and include plan concurrency in the pilot.

7. Murf: for a creator workflow with an API path

Murf text-to-speech API product page

Fits: Media teams and developers who want a studio product and a separate TTS API from the same provider.

Layer and API: Murf's product pricing separates Studio and API plans. In the API, Falcon 2 handles streaming, while Gen2 serves studio-quality synthesis through a non-streaming endpoint. The older Gen2 streaming model is deprecated.

Price and free boundary: Current API pricing lists Falcon at $10 per one million characters and Gen2 at $30. The API Free Trial includes $10 in credit every month, five concurrent requests, and no credit card requirement. Pay-as-you-go credit is purchased separately and does not expire.

Voice and deployment: Murf provides both models through a managed API. Falcon 2 is the real-time path; Gen2 is the non-streaming path for media synthesis.

Main tradeoff: One model does not cover both modes. A workload that includes live speech and batch media needs separate Falcon 2 and Gen2 acceptance tests.

Switching from Polly: Separate batch media from real-time traffic. Migrate authentication, voice IDs, audio formats, SSML-dependent copy, and endpoint behavior model by model.

8. Chatterbox: a free, self-hosted Amazon Polly alternative

Chatterbox open-source repository on GitHub

Fits: Teams that need deployment control and can operate speech inference.

Layer and API: Chatterbox is an MIT-licensed model family. Chatterbox-Turbo is a 350-million-parameter English model, Nano is a 110-million-parameter option for CPU and on-device inference, and Multilingual V3 is a 500-million-parameter model covering 23-plus languages and cross-language cloning. The official repository documents each model.

Price and free boundary: The code is free to use under the MIT license. Compute, storage, autoscaling, monitoring, model updates, and engineering time remain your costs. “Free alternative” describes the software license, not production operation.

Voice and deployment: You run the Chatterbox models on CPU or GPU infrastructure. Cloning uses reference audio, so you still need consent and voice-rights controls.

Main tradeoff: Your team owns throughput, cold starts, recovery, patches, and quality regressions. That burden can outweigh savings at low volume.

Switching from Polly: Build serving, queues, caching, autoscaling, rate limits, and observability. Keep Polly as fallback until the new path demonstrates load capacity and failure recovery.

How to choose the right replacement

  • Production voice agents: Start with Dasha for a managed runtime. Compare component APIs when your team wants to assemble the stack. Our voice API guide explains that approach.
  • Direct cloud TTS: Shortlist Google Cloud TTS and Azure Speech for a closer operating model.
  • Expressive media: Compare ElevenLabs' voice models with Murf's separate Studio and API products.
  • Streaming custom stacks: Compare Deepgram and Cartesia on concurrency and adjacent product costs.
  • Self-hosting: Use Chatterbox when you can operate inference and need deployment control.

A migration pilot that catches expensive surprises

  1. Normalize cost by layer. Use input characters for TTS. Use connected agent-runtime minutes plus telephony, LLM, recognition, and support for a runtime.
  2. Build a representative corpus. Include names, numbers, abbreviations, multilingual and long-form text, plus short phrases from live calls.
  3. Measure user-visible behavior. Record time to first audio, synthesis time, concurrency, errors, pronunciation, and artifacts. Add interruption tests for agents.
  4. Inventory dependencies. List voice IDs, SSML, lexicons, speech marks, codecs, sample rates, cache keys, regions, quotas, and IAM policies.
  5. Run both paths. Shadow batch jobs or canary traffic. Keep rollback until cost and quality remain stable under load.

Frequently asked questions

Is there a free alternative to Amazon Polly?

Chatterbox is MIT-licensed software you can self-host, with compute and operations still payable. Hosted options provide different allowances or credits with commercial-use and time limits.

Is Amazon Polly cheaper than its alternatives?

Polly Standard costs $4 per one million characters. Google Standard and WaveNet match that rate after their monthly allowance. Include inference operations when comparing self-hosting and account for workflow work that another product may remove. AWS pricing and Google Cloud publish the respective rates.

Why is ElevenLabs more expensive than Amazon Polly?

ElevenLabs charges more per character across its current TTS families. v3 Conversational and Flash/Turbo list at $50 per one million characters, while v3 and v2 Multilingual list at $100. Polly Standard is $4 per one million and Neural is $16. Compare the exact models that clear your quality, control, and latency requirements. The ElevenLabs API rates and Polly price sheet show the current gap.

What is the difference between Amazon Polly and Amazon Lex?

Polly converts text into speech. Amazon Lex handles conversational input, intent recognition, slots, and dialog. An application can use Lex to understand and Polly to speak.

Can Dasha replace Amazon Polly?

Yes, when the goal is to replace the voice-agent runtime around Polly. Dasha covers real-time execution and operations. Use a TTS API for narration, accessibility audio, or file generation.

If your Polly workload is part of a live conversational product, start with Dasha and evaluate the complete call path rather than one synthesis request.

Related Posts

We use cookies for functional and analytical purposes. Please refer to our Privacy Policy for details.