AI · 24 September 2026
Qwen-Audio 3.1: Alibaba Cuts AI Voice Pricing by 95%
Alibaba's Qwen team launched five new audio AI models under Qwen-Audio-3.1, adding speaker separation, emotion and noise detection, while cutting usage prices by up to 95 percent.
What happened
Alibaba's Qwen team has released Qwen-Audio-3.1, a new family of five AI audio models covering speech recognition, text-to-speech synthesis and real-time voice interaction, alongside a steep cut in usage pricing of up to 95 percent.
The core automatic speech recognition (ASR) model improves accuracy across multiple languages and dialects, and automatically strips out filler words and repeated phrases from transcripts. A more advanced variant, ASR-Next, goes further by identifying multiple speakers in a single recording with timestamps, and by detecting emotional tone, background ambience and machine or mechanical noise. The lineup also includes text-to-speech models capable of multilingual synthesis, rounding out a suite designed to handle recognition, generation and live interaction in one release.
According to The Decoder, the release positions Qwen-Audio-3.1 as a broader, cheaper alternative to existing audio AI tooling, with Alibaba framing the price reduction as a way to widen access to the technology for developers and enterprises building voice-based products.
Why it matters
Voice interfaces have long been constrained by the cost and fragility of transcription and synthesis, particularly in multilingual or noisy real-world settings such as call centres, in-car systems or multi-speaker meetings. A model that can reliably separate speakers, read emotional and environmental cues, and clean up messy speech automatically removes several of the manual post-processing steps that have made production-grade voice AI expensive to deploy at scale.
The pricing move is arguably as significant as the technical upgrade. A near order-of-magnitude cost reduction changes the calculus for organisations deciding whether voice AI is viable for high-volume use cases — customer service transcription, compliance monitoring, multilingual support — rather than limited pilots. It also intensifies competitive pressure on other audio AI providers to match both capability and price.
By the numbers
- Up to 95 percent reduction in Qwen's AI audio processing prices with the new release.
- Five new models released, spanning speech recognition, text-to-speech and real-time interaction.
The Renascence take
The headline is the price cut, but the more consequential detail is what the models now notice on their own — who is speaking, how they sound, and what is happening around them — without a human having to tag it afterwards.
Most organisations still treat voice data as something to transcribe and archive, not something to interpret. A model that flags frustration, overlapping speakers or background noise in real time turns every call or meeting into a live signal for service design, not just a compliance record. The operators who benefit won't be the ones who adopt cheaper transcription — they'll be the ones who rebuild service workflows around the fact that emotional and contextual cues are now available at the point of interaction, not three weeks later in a QA review.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
FAQ
Questions we get on this topic
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.
