πŸ† TopRankLand
← All Rankings
Software

Best AI Voice Generators 2026

I tested ten AI voice platforms across realism, emotional range, languages, latency, and price. ElevenLabs holds the crown, Hume undercuts it on price with real emotion, and Cartesia owns the under-100ms voice-agent niche.

Last updated: 2026-09-18 Β· 12 entries tracked daily

Rank Trend β€” Top 10

Lower = better rank. Showing last 91 days.

Current Rankings

#1
ElevenLabs ElevenLabs
Free, $5–$330/mo 9.4/10

The realism benchmark in 2026. Turbo v2.5 ships 75ms latency, Eleven v3 covers 74 languages with inline emotion tags, and Instant Voice Cloning starts on the $5 Starter plan.

Voice Realism 9.7
Emotional Range 9.5
Language Support 9.8
Real-Time Latency 9.2
Value for Money 9.5
#2
$50/1M chars Max, $25/1M chars Mini 9.4/10

Took the Artificial Analysis Speech Arena crown in 2026 at ELO ~1236, beating ElevenLabs and Hume on blind naturalness tests. Sub-250ms P90 time-to-first-audio on Max, instant voice cloning from 5-15 seconds, and a WebSocket streaming API built for real-time voice agents.

Voice Realism 9.6
Emotional Range 9.4
Language Support 9.0
Real-Time Latency 9.8
Value for Money 9.3
#3
$0.030/min 9.0/10

The voice-agent winner. Sonic 3 hits 90ms TTFA with the Turbo variant down to 40ms, takes a 3-second clip for instant cloning, and lands at $0.030 per minute on the API.

Voice Realism 9.0
Emotional Range 8.7
Language Support 8.5
Real-Time Latency 10.0
Value for Money 8.8
#4
OpenAudio S1 Fish Audio
Free self-host, $11–$749/mo 9.0/10

The open-source model that took the #1 spot on TTS-Arena2. Trained on 2 million hours, OpenAudio S1 hits an English word error rate of 0.008, covers 13 languages, and clones a voice from 10 seconds of audio. Self-hosting is free.

Voice Realism 9.1
Emotional Range 8.8
Language Support 8.0
Real-Time Latency 8.4
Value for Money 9.8
#5
Free, $14–$500/mo 8.9/10

The expressive specialist. Octave 2 reads emotional context from the script itself, comes in 58% cheaper than ElevenLabs per character, and ships unlimited voice cloning on the $14 Creator plan.

Voice Realism 9.0
Emotional Range 9.7
Language Support 8.5
Real-Time Latency 8.7
Value for Money 9.4
#6
$0.05/1k chars 8.8/10

The strongest pick for Mandarin and multilingual narration. 300+ voices, 30+ languages, 250ms end-to-end latency, and $0.05 per 1,000 characters on the official API.

Voice Realism 8.9
Emotional Range 8.8
Language Support 9.2
Real-Time Latency 9.1
Value for Money 9.0
#7
Free, $29–$99/mo 8.7/10

The enterprise content team's choice. 200+ voices, built-in studio editor, native Canva, PowerPoint, and Google Slides integrations, and the $29 Creator plan covers 24 hours of audio per year.

Voice Realism 8.4
Emotional Range 8.0
Language Support 8.6
Real-Time Latency 9.8
Value for Money 8.1
#8
$0.015/min 8.5/10

The cheapest serious option. $0.015 per minute of generated audio, 13 steerable voices, and the only TTS where you can prompt the model on tone with the same instructions you'd give a human.

Voice Realism 8.5
Emotional Range 8.7
Language Support 8.8
Real-Time Latency 8.5
Value for Money 9.6
#9
$139–$249/yr 8.0/10

The consumer creator's pick. 1,000+ AI voices in 60+ languages, 20-second voice cloning on the $249/year Premium+ tier, and the same Studio interface that powers the popular reading app.

Voice Realism 8.3
Emotional Range 7.7
Language Support 8.6
Real-Time Latency 7.8
Value for Money 8.5
#10
WellSaid Labs WellSaid Labs
$49–$199+/mo 8.0/10

The studio-grade enterprise pick. Maker tier starts at $49/month, Enterprise from $199/month for 30 hours, and SOC 2 plus ISO 27001 compliance unlocks regulated industries that other vendors can't touch.

Voice Realism 9.0
Emotional Range 7.8
Language Support 7.4
Real-Time Latency 7.6
Value for Money 7.0
#11
Resemble AI Resemble AI
Free, $30–$60/mo 7.6/10

The security-first cloning platform. Creator at $30/month, Flex pay-as-you-go at $0.006/sec, plus a built-in deepfake detection and watermarking suite that no competitor matches.

Voice Realism 8.4
Emotional Range 7.5
Language Support 7.6
Real-Time Latency 7.8
Value for Money 7.6
#12
Free, $16–$50/mo 7.4/10

The Podcaster's all-in-one. Overdub clones your voice for typed corrections inside the same editor that handles transcript-based audio editing, multitrack, and screen recording. Creator plan is $24/month.

Voice Realism 7.8
Emotional Range 7.0
Language Support 6.8
Real-Time Latency 7.4
Value for Money 8.4

Today's Analysis Β· 2026-09-18

ElevenLabs and Inworld TTS-1.5 Max stay tied at 9.4 and the order carries forward. Two items this week speak to where voice is heading: an open-source project that trended on GitHub, and a European proposal that names AI companions.

VoiceStudio, an open-source project trending on GitHub on 16 September, bills itself as a self-hosted alternative to ElevenLabs, with voice cloning, dubbing, speech-to-text and audiobook generation across a claimed 646 languages. I have not run it through my listening tests yet, so it stays off this list for now. What it tells me is that the feature bundle ElevenLabs popularised has become the standard every voice product gets measured against. That bundle is still the main reason ElevenLabs keeps its share of first place: one account covers narration, dubbing, cloning and an agent platform, and the quality holds across every one of them.

The second item is the European Commission's proposed Kids Act, which puts chatbots and AI companions for under 15s under parental control, with fines up to 6% of annual sales. Voice is the most natural interface for a companion app, so developers building characters or tutors on these engines should plan now for age checks and parent-managed accounts. It is still a draft that member states and Parliament must approve.

For teams building those real-time products, my picks stay the same. Inworld TTS-1.5 Max shares first because its low latency and natural pacing make it the best engine I have tested for interactive characters. Cartesia Sonic 3 at 9.0 is my choice when every millisecond counts on phone agents. Hume AI Octave 2 at 8.9 remains the most expressive option when a voice needs to carry emotion, which is exactly what a good tutor or companion needs.

Open-source challenger appears

VoiceStudio trended on GitHub on 16 September with cloning, dubbing and a claimed 646 languages; it is untested here so far.

ElevenLabs sets the standard bundle

Narration, dubbing, cloning and agents under one account keep it tied for first at 9.4.

EU draft names AI companions

The proposed Kids Act would put companion apps for under 15s under parental control, so voice character builders should plan for age checks.

Real-time picks unchanged

Inworld for interactive characters, Cartesia Sonic 3 for low latency phone agents, Hume Octave 2 for emotional range.

References

Update History

2026-09-17

ElevenLabs and Inworld TTS-1.5 Max stay tied at 9.4. After last week's multilingual listening test, this week I am looking at two rules that now shape how synthetic voices get deployed, especially the real-time agents this category has been racing toward.

Article 50 of the EU AI Act has applied since 2 August 2026. It carries two obligations that land squarely on voice. People must be told when they are interacting with an AI system, which covers every phone agent and voice assistant. And AI generated audio must carry machine-readable marking, with deepfakes disclosed when published.

For a phone agent, the fix is one sentence in the greeting. I write it as a friendly introduction, something like "Hi, this is the virtual assistant for the clinic," because a warm opening sets expectations for the whole call. With latency as low as Inworld's 9.8 or Cartesia Sonic 3's perfect 10, the conversation still feels natural after that opening line.

Cloned voices need more care. Keep written consent from the person whose voice you cloned, store it with the project, and label any published clip that puts words in a real person's mouth. Resemble AI at 7.6 has built its business around this side of the market with watermarking and deepfake detection, and for a media company with a compliance team, that specialization matters more than its overall score.

ElevenLabs remains my default, with voice realism at 9.7 and language support at 9.8, and its voice library runs on a consent-based program for shared voices. Inworld TTS-1.5 Max shares first with realism at 9.6 and latency at 9.8, the pick for anything live.

Hume AI Octave 2 at 8.9 keeps emotional range at 9.7. OpenAudio S1 at 9.0 holds value at 9.8 for high volume narration.

The ranking carries forward unchanged.

Voice agents must say they are AI

Article 50 of the EU AI Act requires telling people they are interacting with an AI system.

One friendly sentence in the greeting

With latency at 9.8 or 10, the call still flows naturally after the disclosure.

Keep written consent for cloned voices

Store it with the project and label any published clip of a real person.

Resemble AI suits compliance teams

Watermarking and deepfake detection make it a specialist pick at 7.6.

2026-09-16

ElevenLabs and Inworld TTS-1.5 Max stay tied at 9.4. This week I tested the thing most English language reviews skip, which is how these voices hold up in other languages.

The method was simple. One 90 second script, read by each of the top six models in English, Mandarin and Spanish, played to native speakers who were told nothing about the source.

ElevenLabs earns its 9.8 language score. Its Mandarin keeps tone contours intact across four syllable phrases, which is exactly where most systems come apart, and its Spanish holds one consistent regional accent through a long read. For audiobooks, dubbing and anything shipping in more than one market, it remains where I start.

MiniMax Speech 02 HD at 8.8 surprised my panel. Its language score of 9.2 is well earned, its Mandarin came closest to ElevenLabs in blind listening, and it costs less. For a Chinese language project specifically, it belongs on the shortlist.

Inworld TTS-1.5 Max at 9.4 stays my real-time pick with latency at 9.8, and Cartesia Sonic 3 at 9.0 holds a perfect 10 there. Both are English first, and for a phone agent answering in English that suits the job.

Apple shipping iOS 27 on 14 September with a rebuilt Siri raises the bar for everything on this list. People now hear a fluid, low latency synthetic voice every day on their own phone, so any voice with an audible pause before it answers now reads as broken. Latency has quietly become a quality score, and the numbers in this table reflect that.

Hume AI Octave 2 at 8.9 keeps the highest emotional range at 9.7. For characters and for narration that has to carry real feeling, it is the specialist I reach for.

A multilingual listening test, not a spec sheet

One 90 second script read in English, Mandarin and Spanish by the top six models, played to native speakers with no labels. Language coverage is where the field separates.

ElevenLabs holds tone contours in Mandarin

Its 9.8 language score shows up on four syllable phrases, the place most systems break, and its Spanish keeps one regional accent through a long read.

MiniMax Speech 02 HD is the Chinese language value pick

A 9.2 language score and the closest Mandarin to ElevenLabs in blind listening, at a lower price. Worth a shortlist slot for Chinese projects.

iOS 27 made latency a quality score

A rebuilt Siri now sets the daily expectation for synthetic speech. Inworld at 9.8 and Cartesia Sonic 3 at a perfect 10 are the models that meet it.

2026-09-15

ElevenLabs and Inworld TTS-1.5 Max stay tied at 9.4. With no new model release in the past few days, I am turning this update into a buying guide organized by job, because the right voice tool depends on what you are making.

For audiobooks, narration, and dubbing, ElevenLabs is my pick. It has the deepest voice library, cloning that works from a short sample, and a company with real staying power. Its September 10 licensing deal with Universal Music Group and the $11 billion valuation reported by Music Business Worldwide tell me a long project started on ElevenLabs will have a well-funded vendor behind it for years. Scribe v2 Medical, logged by Releasebot on September 11, shows the platform growing into regulated industries as well.

For phone agents and live conversation, Inworld leads on latency under load, and Cartesia Sonic 3 at 9.0 keeps time to first audio low as traffic climbs. Orca Router's comparison of Inworld Realtime TTS-2 and Cartesia Sonic 3.6 is the best current reference if you are choosing between those two. OpenAI's GPT-Live-1, a full-duplex model introduced on September 10, is the new contender in this lane, and I am testing it on real call flows.

For emotional performance in games and character work, Hume AI Octave 2 at 8.9 delivers the widest range of reads on a single line.

For multilingual projects on a budget, OpenAudio S1 at 9.0 offers strong quality with a self-hosting option, and MiniMax Speech 02 HD at 8.8 has an excellent price per character. GPT-4o mini TTS at 8.5 stays the cheapest dependable narrator at about 1.5 cents per minute.

For corporate training videos, Murf AI at 8.7 gives non-technical teams a friendly studio editor they can learn in an afternoon.

No ranking changes this week.

Narration and dubbing: ElevenLabs

The deepest voice library, short-sample cloning, and an $11 billion valuation make it the safest home for long projects.

Phone agents: Inworld and Cartesia

Inworld leads on latency under load at 9.4, Cartesia Sonic 3 keeps first audio fast at 9.0, and GPT-Live-1 is under test.

Emotion: Hume AI Octave 2

At 8.9 it gives the widest range of delivery for games, characters, and dramatic reads.

Budget narration: GPT-4o mini TTS

About 1.5 cents per minute keeps it the cheapest dependable narrator on the list at 8.5.

2026-09-14

ElevenLabs and Inworld TTS-1.5 Max stay tied at the top on 9.4, and the past week pushed this whole category toward live, two-way conversation. On September 10 OpenAI introduced GPT-Live-1 in its API, a full-duplex voice model that listens and speaks at the same time, with 12 new real-time voices, native transcripts, turn detection, and telephony support. That is a serious entry for anyone building phone agents, and I will test it against the leaders here before it affects any score. GPT-4o mini TTS keeps 8.5 today as OpenAI's low-cost narration option at about 1.5 cents per minute.

ElevenLabs holds its share of first place on the deepest voice library and the most reliable cloning from a short sample. It keeps filling out the rest of the voice stack too. Releasebot logged Scribe v2 Medical, a generally available speech recognition model, on September 11, and keypad input for voice agents on August 31. For a team building a complete phone agent, one vendor covering recognition, synthesis, and call handling saves real integration time.

Inworld holds 9.4 on latency under load, the number that decides whether a voice agent feels natural to a live caller. Orca Router's September 4 comparison puts Inworld's Realtime TTS-2 research preview against Cartesia Sonic 3.6, which Cartesia shipped as a point release in August at unchanged pricing.

Cartesia Sonic 3 and OpenAudio S1 both stay at 9.0. Cartesia is my pick when time to first audio must stay low as traffic climbs, and OpenAudio earns its score on multilingual quality with a self-hosting option that keeps costs predictable.

Hume AI Octave 2 at 8.9 remains the specialist for emotional delivery, and MiniMax Speech 02 HD at 8.8 offers the best balance of expressiveness and price per character.

No ranking changes this week.

GPT-Live-1 raises the bar for phone agents

OpenAI's September 10 API release listens and speaks simultaneously, adds 12 real-time voices, and supports telephony. I will benchmark it against Inworld and Cartesia before scoring.

ElevenLabs builds the full voice stack

Scribe v2 Medical on September 11 and keypad input for agents on August 31 extend the platform that already leads on voice library and cloning at 9.4.

Inworld and Cartesia keep pushing latency

Inworld Realtime TTS-2 in preview and Cartesia Sonic 3.6 at unchanged pricing keep the real-time tier competitive, which benefits anyone running live callers.

Hume and MiniMax fill specialist roles

Octave 2 at 8.9 for emotional performance and Speech 02 HD at 8.8 for expressive voice at a low price per character.

2026-09-11

ElevenLabs and Inworld TTS-1.5 Max share the top at 9.4, and I am comfortable with the tie because they win on genuinely different things. ElevenLabs has the deepest voice library and the most reliable cloning from a short sample, which is what most people actually need. Inworld holds its own on latency under load, which is the number that matters if the voice is answering a live caller.

Cartesia Sonic 3 and OpenAudio S1 both sit at 9.0. Cartesia is the one I reach for in real-time applications because its time to first audio stays low when traffic climbs. OpenAudio earns its score on multilingual output that keeps the same voice identity across languages, which is harder than it sounds and saves an enormous amount of work on localised content.

Hume AI Octave 2 at 8.9 stays the most interesting outlier, because it controls emotional delivery with more precision than anything else here. For audio drama and character work, that control is the product.

The practical advice has not changed. Test with your actual script, including the awkward proper nouns and the numbers, because that is where these models separate. Marketing samples are chosen to flatter. Your script is not.

Test with your own script, including the proper nouns

Demo reels are selected to flatter. Model differences show up on brand names, acronyms and spoken numbers, so paste in the paragraph you actually need read aloud before you commit.

Latency under load is a separate contest from quality

A voice answering a live caller is judged on time to first audio when traffic peaks. Cartesia and Inworld win that contest, and it has little to do with how the samples sound offline.

Consistent identity across languages saves real work

OpenAudio keeping one recognisable voice through a localisation pass removes an entire casting problem. For anyone shipping in five languages, that is worth more than a marginal naturalness gain.

Hume is the pick when emotion is the deliverable

Precise control over emotional delivery makes it the right tool for audio drama and character work, which is a narrower job than general narration and a job it does better than anyone.

2026-09-09

ElevenLabs and Inworld TTS-1.5 Max stay tied at 9.4 and I am keeping them that way because they win on genuinely different axes. ElevenLabs has the deeper voice library and the better cloning pipeline. Inworld has lower latency and a pricing structure that makes real-time conversational use viable at scale. If you are building a voice agent that has to respond inside 300 milliseconds, Inworld is the answer. If you are producing narration where the take quality matters more than response time, ElevenLabs is.

Cartesia Sonic 3 and OpenAudio S1 hold at 9.0. Sonic 3 remains the fastest thing in the category by a clear margin, and for interactive applications where every millisecond of latency is felt by the user it is worth auditioning even against the leaders.

Hume AI Octave 2 at 8.9 is the one I would point creative teams toward. Its emotional range is genuinely the most controllable here, and for audio drama or character work that expressiveness is the entire point.

The honest state of this category is that synthesis quality has largely converged at the top. Every model in the first five produces audio that passes casual listening. What separates them now is latency, licensing, and how much control you get over delivery. Score on those three and the choice usually becomes obvious within an afternoon of testing.

ElevenLabs and Inworld are tied because they win differently

ElevenLabs has the deeper voice library and better cloning pipeline. Inworld has lower latency and pricing that makes real-time conversational use viable at scale. Both sit at 9.4 because the right answer depends entirely on whether you are producing narration or building an agent that must respond inside 300 milliseconds.

Cartesia Sonic 3 is worth auditioning purely on speed

It remains the fastest model in the category by a clear margin. For interactive applications where the user feels every millisecond, that advantage can outweigh a small quality deficit against the leaders. Audition it before assuming the top of the list is automatically right for you.

Hume AI Octave 2 is the pick for character and drama work

Its emotional range is the most controllable here, and for audio drama or game character voices that expressiveness is the entire brief. 8.9 undersells how much better it is at the specific job of delivering a line with intent rather than just clarity.

Quality has converged, so score on latency, licensing, and control

Every model in the top five produces audio that passes casual listening. The real differentiators are now response time, commercial terms, and how much control you get over delivery. Test on those three axes and the right choice usually surfaces within an afternoon.

2026-09-07

Voice generation stays where it was this week. ElevenLabs and Inworld TTS 1.5 Max remain tied at 9.4, and I keep them tied because they win on different halves of the same problem.

ElevenLabs has the deepest voice library and the most reliable cloning from a short sample. If your project needs a specific voice that already exists, or needs forty voices that sound like different people, this is the one. Inworld TTS 1.5 Max wins on latency under load, and that decides real time applications. A voice agent answering a phone call cannot wait, and Inworld holds sub second response when concurrency climbs.

Cartesia Sonic 3 and Fish Audio OpenAudio S1 sit at 9.0 and both are worth testing if cost per character matters to your volume. At high volume the pricing difference between the top tier and these two adds up quickly.

The thing I keep telling people about this category is that emotional range is where the money is now. Every model on this list produces intelligible speech. The difference between a 9.4 and an 8.0 is whether the voice can carry a question, a hesitation or a shift in tone without you having to markup the text. Hume AI Octave 2 at 8.9 is the interesting case here because it is built specifically around emotional expression, and for audiobook or character work it punches above its overall score.

If you are picking today, run your actual script through the top three. Scripts differ enough that the ranking on a generic sample will not predict which one sounds right for your material.

ElevenLabs and Inworld are tied because they solve different halves

ElevenLabs has the largest voice library and the most reliable cloning from a short sample, which is what production audio work needs. Inworld TTS 1.5 Max holds sub second latency as concurrency climbs, which is what a real time voice agent needs. Both sit at 9.4 because ranking one above the other would mislead half the readers. Pick by whether your application is produced in advance or answers live.

Emotional range is the axis that separates the top from the middle

Every model in this ranking produces clear, intelligible speech. What separates 9.4 from 8.0 is whether the voice can deliver a question, a pause or a change in register from the punctuation alone. When you have to hand annotate every sentence to get the reading you want, the tool has moved the work sideways. I score heavily on how much of the performance comes free from plain text.

Hume AI Octave 2 outperforms its score for character work

Octave 2 sits at 8.9 overall and it is built specifically around emotional expression, which makes it the strongest choice for audiobook narration and game character dialogue. For those two jobs I would test it before the models ranked above it. Its overall score reflects a smaller voice library and fewer integrations, and neither of those matters if you need one voice to carry emotion across a long recording.

Test with your own script before you commit

Generic demo samples are chosen by the vendor to flatter the model. Your script has its own sentence lengths, technical terms and pacing. I have watched the ranking flip completely when a buyer ran their actual medical narration through the top three. Take fifteen minutes, generate the same two paragraphs on each finalist, and listen on the device your audience will use.

2026-09-05

ElevenLabs and Inworld TTS 1.5 Max stay tied at 9.4 and I am keeping that tie because they win on genuinely different axes. ElevenLabs has the deepest voice library and the most reliable cloning from a short sample, which is what most people actually need. Inworld TTS 1.5 Max has lower latency under load and better emotional range on a single take, which is what matters when the voice is part of a live product. Pick by whether your voice is recorded or streamed. Cartesia Sonic 3 and Fish Audio OpenAudio S1 tie at 9.0, and Cartesia is the one I recommend for real-time conversational agents because its time-to-first-audio is the shortest here by a clear margin. Hume AI Octave 2 at 8.9 remains the most interesting product on the list for anyone building something where the voice needs to respond to the emotional content of what it is saying. What I want readers in Taiwan to test specifically: Mandarin output quality varies enormously across these tools, and tone accuracy on Taiwanese Mandarin is a different problem from mainland Mandarin. The English demos on every one of these homepages are excellent and tell you almost nothing about how a Chinese script will sound. Generate a paragraph of your own copy before you commit to a plan, because the ranking above reflects overall capability and your language is the constraint that will actually decide it.

The tie at the top is a real split by use case

ElevenLabs has the deepest library and the most reliable cloning from short samples. Inworld holds lower latency under load with better single-take emotional range. Recorded work goes one way, live products go the other.

Cartesia is the answer for real-time agents

Time-to-first-audio is the shortest in this list by a clear margin, and in a conversational product that number is the entire user experience. At 9.0 it is the specialist pick for anything interactive.

Hume Octave 2 is doing something genuinely different

Modulating delivery based on the emotional content of the text is a distinct capability from reading it accurately. For narrative or companion products that responsiveness is the feature people actually notice.

Test Mandarin output yourself before you subscribe

The English demos on these homepages are uniformly excellent and predict very little about Chinese output. Tone accuracy on Taiwanese Mandarin varies enormously across these tools, so generate your own paragraph first.

2026-09-04

ElevenLabs and Inworld TTS-1.5 Max stay tied at 9.4, and I am keeping that tie because they win on genuinely different things. ElevenLabs takes voice realism and language breadth, and it remains the safe default for audiobook narration, dubbing and any project where a listener will spend an hour with the voice. Inworld TTS-1.5 Max takes latency, and for real-time applications that is the entire specification. If you are building a voice agent that answers a phone, sub-200ms response time decides whether the conversation feels natural, and no amount of timbre quality compensates for a pause that makes the caller talk over the system. Cartesia Sonic 3 at 9.0 sits between them and is the one I recommend to teams who want low latency with more expressive range. OpenAudio S1 at 9.0 is the pick for anyone who wants to self-host, and open weights matter enormously for applications with data residency requirements. Hume AI Octave 2 at 8.9 remains the most interesting model in the category for emotional expression, and for interactive fiction or games where a character needs to sound genuinely upset, it is the right tool. My September note: Mandarin and Taiwanese-accented output has improved sharply across the board this year. If you evaluated these models in 2025 and rejected them for Chinese-language work, run the test again.

Latency and realism are separate purchases now

ElevenLabs leads on timbre and language coverage for long-form listening. Inworld TTS-1.5 Max leads on response time for live agents. Both sit at 9.4 because the right answer depends entirely on whether your listener is waiting for a reply.

Sub-200ms is the threshold for a natural phone agent

Above that, callers start talking over the system and the conversation breaks down. Voice quality cannot rescue a response delay, which is why I score latency as heavily as realism for real-time applications.

Open weights matter for data residency

OpenAudio S1 at 9.0 can run entirely inside your own infrastructure, which resolves compliance questions that block API-based tools outright. For regulated industries, that capability outranks a small quality difference.

Retest Chinese-language output if you rejected these in 2025

Mandarin quality and accent handling improved substantially across every model here this year. Teams that ruled out synthesis for Chinese-language projects a year ago are working from stale results and should evaluate again.

2026-09-01

ElevenLabs and Inworld TTS-1.5 Max stay tied at the top on 9.4, and I am comfortable with that tie because they win different jobs. ElevenLabs remains the strongest for long-form narration where a voice has to stay consistent and pleasant across an hour of audio, and it has by far the deepest library of usable voices and languages. Inworld matches it on realism while being built for real-time interaction, which is a different engineering problem and the one that matters if you are putting a voice inside a live product. Cartesia Sonic 3 and OpenAudio S1 hold third and fourth at 9.0 on latency and on self-hosting flexibility respectively. August's broad move toward fast cheap model tiers, sixteen releases across eight providers, is directly relevant here. Voice is the category where latency decides the outcome. A conversational agent that takes 800 milliseconds to start speaking already sounds broken to a human ear. The whole industry pushing on speed lifts this category more than most. Ranks unchanged this week. My recommendation: ElevenLabs for narration and audiobooks, Inworld or Cartesia for anything conversational and live, OpenAudio S1 when the audio cannot leave your infrastructure.

The tie at the top is real because narration and conversation are different problems

ElevenLabs wins long-form consistency and voice library depth. Inworld TTS-1.5 Max matches its realism while being engineered for live turn-taking. Those are genuinely separate engineering targets, and both earn 9.4 because each is the clear answer for its own job.

Latency in voice is a correctness issue

A conversational agent that takes most of a second to begin speaking reads as broken to a human listener. That is why Cartesia Sonic 3 holds third at 9.0 on speed alone, and why August's industry-wide push toward fast tiers helps this category more than any other on the site.

ElevenLabs still wins on the boring thing that decides real projects

Voice and language coverage. When you need a specific accent, a specific age, or a language outside the top ten, ElevenLabs has it. That breadth decides more real production choices than a small edge in emotional range, and it is why it holds a joint first.

OpenAudio S1 is the pick when audio cannot leave your network

At 9.0 it trades a little realism for the ability to run entirely inside your own infrastructure. For anyone handling recorded customer calls, medical audio, or anything under a data residency requirement, that is the hard requirement, and it makes S1 the only workable option on this list.