Alibaba's Qwen-Audio-3.0-TTS Plus arrived in September 2026 and sits second on the Artificial Analysis text-to-speech provider-voice arena at 1260 Elo.
Eleven v3, from ElevenLabs, sits fourteenth at 1168. It bills $100.00 per million characters. Qwen bills $27.60.
The top of the board
Cartesia's Sonic 3.6 holds first place at 1276, sixteen points above Qwen, at $49.00 per million characters — still half ElevenLabs' rate.
Read the confidence intervals before reading the ranks. Sonic 3.6 and Qwen are ±16 apart on a 16-point gap: on this evidence they are not separable, and anyone claiming a clear winner between those two is over-reading the board.
Languages are the real differentiator
Qwen-Audio-3.0-TTS Plus covers 16 languages — Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai and Vietnamese — plus 20 Chinese dialect regions.
MiniMax Speech 2.8 HD goes further: 32 languages, including Cantonese, Ukrainian, Hindi, Czech, Finnish, Greek, Romanian and Afrikaans. It ranks sixteenth at 1166 Elo. If your requirement is coverage rather than naturalness in one language, that ordering inverts.
What the arena does not tell you
Three things, and they decide most real deployments.
Voice cloning and licensing. The arena tests provider-supplied voices. It says nothing about whether you may clone a voice, or under what terms.
Streaming latency. These scores come from generated samples, not from time-to-first-audio. A model that sounds marginally better but starts half a second later is worse for a live agent.
Per-character billing is not per-token billing. The rates above are the arena's own per-character figures, which is why they compare. Your provider may meter tokens, and the conversion is not one to one.
These are the first two speech-synthesis models CompareLLM tracks; the category previously held only music generation and speech recognition, which are scored on different boards and are not comparable to these. Elo, voice counts and prices are from the Artificial Analysis text-to-speech arena, read 21 September 2026.

