AI & Machine Learning

Alibaba’s Qwen Audio 3.0 TTS Tops the Speech Leaderboard

Alibaba released Qwen Audio 3.0 TTS, a speech model that ranks first on an independent text-to-speech leaderboard, clones voices across 16 languages and costs a third of ElevenLabs.

By Daniel Mercer Edited by Maria Konash Published:
Alibaba’s Qwen Audio 3.0 TTS Tops the Speech Leaderboard
Alibaba released Qwen Audio 3.0 TTS, a speech model ranking first on an independent leaderboard with voice cloning across 16 languages. Image: Jumping Jax / Unsplash

Alibaba’s Tongyi Lab released Qwen Audio 3.0 TTS, a text-to-speech model that now ranks first on the independent Artificial Analysis speech leaderboard while charging roughly a third of what ElevenLabs and MiniMax do.

The model ships in two tiers: Flash, tuned for real-time interaction with first-packet latency around 300 milliseconds, and Plus, tuned for high-quality output. It supports 16 languages including Russian, Arabic, Japanese and less commonly covered ones like Tagalog, Malay, Thai and Vietnamese, plus 20 Chinese dialect regions.

Plus is priced at about $27.59 per million characters. Notably, despite the Qwen name, this model is closed and available only as a hosted API through Alibaba Cloud Model Studio, with no downloadable weights, a departure from Alibaba’s earlier open-weight speech releases.

The capability set is aimed at production use rather than demos. The model can clone a voice from an audio sample and preserve that timbre when generating speech in a different language, and Alibaba reports speaker similarity averaging 82.75 across all 16 languages, its own figure.

Users can steer delivery with plain instructions, asking it to read slowly like a bedtime story or in a particular manner, and can apply 86 inline tags for fine control over individual phrases, including cues like whisper, angry, breaths and laughs. It handles noisy, reverberant or poorly recorded reference audio, generates up to three minutes of speech in a single pass, and outputs at sample rates up to 48 kHz across PCM, WAV, MP3 and Opus formats.

The leaderboard claim deserves a precise reading. Artificial Analysis places Qwen-Audio-3.0-TTS-Plus first with an Elo of about 1,236, but that sits only two points above SpeechifyAI’s Simba 3.2 at 1,234, within overlapping confidence intervals, making it effectively a statistical tie at the top rather than a decisive win. It does rank clearly ahead of Google’s Gemini 3.1 Flash TTS at 1,214 and Sonic 3.5 at 1,207.

Quality Bought With Speed

The model’s genuine trade-off is throughput. Qwen generates roughly 16 characters per second, well below Simba 3.2 at 30.2, Gemini 3.1 Flash TTS at 27, and Sonic 3.5 at 120. That makes it a poor fit for high-volume batch synthesis, where a faster model will finish far sooner even at lower quality scores, while suiting work where naturalness matters more than speed, such as audiobooks, dubbing or character voices.

The combination on offer, top-tier perceived quality at roughly a third of the leading Western price, is the real competitive story, and it mirrors a broader pattern in which Chinese labs are pressuring US incumbents on price rather than out-innovating them outright. The release continues a striking run for Alibaba, which has shipped leading models in language, image and video generation in recent months.

The Voice-Cloning Question

Capable cross-lingual voice cloning from short samples raises consent and misuse concerns that the release does not fully address. A system that can reproduce a person’s voice from imperfect audio and then speak in languages they never learned is powerful for accessibility, localization and content production, and equally useful for impersonation, fraud and synthetic audio of public figures.

Alibaba has not detailed the consent verification or watermarking safeguards attached to the hosted service, which is where the closed distribution model cuts both ways: keeping the weights private gives Alibaba a control point to enforce policy, but it also means outsiders cannot audit what protections exist.

As voice synthesis approaches the point where cloned speech is difficult to distinguish from a recording, the governance around these tools is becoming as consequential as their benchmark scores, and it remains the least documented part of this launch.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, News