On September 23, Alibaba's Qwen team officially released the Qwen-Audio-3.1 series of large speech models. The update spans three major areas—speech recognition, speech synthesis, and real-time voice interaction—and adds the audio creation model Qwen-Audio-3.1-TTS-Next and the audio understanding model Qwen-Audio-3.1-ASR-Next.
Five models were released in total, forming a complete audio capability system covering "understanding—generation—interaction—creation." The team also announced across-the-board price cuts for the Qwen-Audio series, with TTS prices down about 70%, Realtime down about 85%, and ASR down as much as 95%.
Qwen-Audio-3.1-ASR: Enhanced Contextual Understanding and Transcription
Qwen-Audio-3.1-ASR is a new-generation speech recognition model that further strengthens multilingual, Chinese dialect, and contextual understanding capabilities, and adds native transcription refinement.
The model can automatically remove filler words and repeated expressions, handle self-corrections and colloquial phrasing, and reorganize meaning based on context, making transcripts smoother and clearer.
In multi-person meetings, interviews, and conversations, the model can simultaneously output speaker labels, timestamps, and text content, recording "who said what and when." Even with brief interruptions and overlapping speech, it works to distinguish different speakers and produce structured transcripts.
Support for 30 Languages and 16 Chinese Dialects
Qwen-Audio-3.1-ASR supports 30 languages and 16 Chinese dialects, making it useful for multilingual communication, dialect conversations, and cross-language content processing.
In professional settings, the model supports industry terms, specialized entities, and tiered hotword recognition, and can also draw on historical context within long audio to improve consistency in recognizing names, abbreviations, and technical terminology.
In streaming recognition mode, the model transcribes as you speak, with a first-word response time of about 160 milliseconds, suitable for meeting notes, live captions, and voice input.
Dialect Recognition Benchmark Performance
In public dialect transcription and translation benchmarks based on the KeSpeech and WSYue datasets, Qwen-Audio-3.1-ASR achieved an average character error rate (CER) of 4.55% and delivered the best results in 6 of 11 subsets.
In internal tests covering 16 Chinese dialects, the model achieved the best results on 11 dialect subsets, with an average CER of 10.38%, showing especially notable gains on harder dialects such as Wenzhou dialect.
In tests translating dialect speech into Mandarin, the model achieved an average semantic sentence accuracy of 82.10% and delivered the best results in 10 of 11 test subsets.
Qwen-Audio-3.1-ASR-Next: Expanding from Speech Recognition to Full Audio Understanding
Qwen-Audio-3.1-ASR-Next is built on a new-generation architecture, with a focus on improving the ability to understand complete audio content.
It can not only recognize spoken words but also understand human emotions, environmental sounds, and mechanical sounds, and carry out sound description, event localization, audio question answering, and complex reasoning.
For example, given audio that contains dialogue, background music, and ambient sound at the same time, the model can first output the dialogue, then identify the other sounds in the audio, determine when the related events occurred, and answer questions about the entire audio clip.
This means the object handled by speech recognition models has expanded from "the words in speech" to complete audio information that includes tone, environment, and events.
In multi-speaker recognition tests, Qwen-Audio-3.1-ASR-Flash-Next and Qwen-Audio-3.1-ASR-Flash achieved the best results on five and three of eight metrics across four test sets, respectively, outperforming the earlier Fun-ASR cascade system overall.
Qwen-Audio-3.1-TTS: Multilingual and Dialect Speech Synthesis
Qwen-Audio-3.1-TTS is a new-generation speech synthesis model that supports synthesis in multiple languages and Chinese dialects.
The model places particular emphasis on the naturalness of transferring the same voice across different languages. Users can deliver multilingual output with a single voice, without reconfiguring the timbre for each language.
Beyond ensuring accurate pronunciation, Qwen-Audio-3.1-TTS also supports controlling emotion, speaking rate, and delivery style through instructions, allowing speech to better match the content and setting.
The model can be applied to multilingual customer service, smart assistants, educational content, audio content, and real-time translation.
Qwen-Audio-3.1-TTS-Next: Unified Generation of Voice, Sound Effects, and Ambient Sound
Qwen-Audio-3.1-TTS-Next is the new audio creation model in this release.
Traditional audio production typically requires hosts, voice actors, sound designers, mixing engineers, and editors to handle dialogue, sound effects, background sound, and post-production separately. TTS-Next aims to integrate these steps into a single model, generating complete audio content in a unified way.
Unified Generation of Multiple Sound Types
TTS-Next can generate complete audio that includes character dialogue, background music, action sounds, and ambient sound.
For example, in a podcast setting the model can generate the dialogue of multiple speakers along with background sound effects at the same time; in film or game settings, it can generate lines, footsteps, ambient sound, and a soundtrack based on a script.
Maintaining Consistent Voices and Emotions Across Multiple Characters
The model supports multi-speaker voice cloning and multi-turn dialogue generation, keeping the vocal characteristics of different characters consistent throughout continuous content.
When generating dialogue from a script, the model can handle pauses, turn-taking, back-channeling, and emotional shifts, making multi-character conversations more coherent and reducing the problem of characters "changing voices" during continuous generation.
Blending Voice, Sound Effects, and Spatial Environment
TTS-Next can process voice, sound effects, and ambient sound at the same time, conveying the texture of sounds, changes in distance, and spatial depth.
For example, in an audiobook setting the model can generate narration, character dialogue, city ambience, wind, and the clink of a soda can, giving the audio a fuller sense of scene.
The model also supports controlling tone, speaking rate, volume, music style, instrumentation, and the ordering of sound events through natural language, and supports fine-grained timestamp control, voice cloning, and 48kHz audio output.
These capabilities can be used in professional audio creation such as podcasts, audiobooks, film and TV, games, advertising, and short videos.
Qwen-Audio-3.1-Realtime: Full-Duplex Real-Time Voice Interaction
Real-time voice interaction is not simply a matter of chaining together speech recognition, a language model, and speech synthesis. In real conversations, users may interrupt while the other party is speaking, or add new information before a sentence is finished.
Qwen-Audio-3.1-Realtime supports full-duplex real-time interaction, listening to the user's voice and responding at the same time, allowing users to interrupt, break in, and add conditions at any time, making the interaction closer to a call with a real person.
Understanding Emotion and Communicative Intent
What the model understands is not just the spoken words, but also the emotion and communicative intent behind the voice.
When the system detects that a user sounds down, it can slow its speaking rate and adjust its wording accordingly; faced with states such as nervousness, happiness, disappointment, or hesitation, it can also choose more appropriate phrasing based on context.
In role-play settings, a character's persona is reflected not only in text prompts but also through continuous vocal expression.
Support for Real-Time Multilingual Switching
During real-time conversations, users can switch between languages. The model can retain prior context, tone, and conversational rhythm, without having to restart the conversation because of a language change.
Such capabilities can be applied to multilingual customer service, cross-language meetings, real-time translation, and internationalized smart hardware.
Strengthening Safety and Factual Reliability
When faced with uncertain information, Qwen-Audio-3.1-Realtime reduces unsupported answers; when faced with sensitive or high-risk content, it strengthens boundary recognition and gives more cautious responses.
The model focuses not only on "whether it can answer," but also on when it should seek further confirmation and when it should not offer unverified conclusions.
Support for Real-Time Agent Tool Calls
When users pose tasks that require external information or coordination with business systems, the model can call APIs, knowledge bases, and business tools during a real-time conversation, and return the results to the user in a natural way.
For example, users can look up information, call business systems, modify task requirements, or check on progress within a voice conversation.
Built for Agents and Smart Hardware Applications
Qwen-Audio-3.1-Realtime is being integrated with agent products such as Qoder and Qwen Office, as well as smart hardware including QwenNote, A2, Eva, Qwen AI glasses, and Leqi AI glasses.
In these settings, users don't need to keep a computer or phone open at all times; they can initiate tasks directly by voice and continue to add conditions, modify requirements, or ask about execution status during the conversation.
For smart glasses, in-vehicle devices, home terminals, and mobile office devices, real-time voice interaction can become the main entry point between users and AI systems.
From a Single Speech Model to a Complete Audio Capability System
The Qwen team says the Qwen-Audio-3.1 series upgrade is not an improvement on a single metric, but rather an effort organized around five technical paths:
- Qwen-Audio-3.1-ASR: improving multilingual and dialect recognition and contextual understanding, with support for native transcription refinement;
- Qwen-Audio-3.1-ASR-Next: expanding the scope of understanding to complete audio information such as human emotions, ambient sounds, and mechanical sounds;
- Qwen-Audio-3.1-TTS: improving multilingual and dialect synthesis, supporting natural transfer of the same voice across languages;
- Qwen-Audio-3.1-TTS-Next: expanding from single-voice synthesis to unified generation of voice, sound effects, and ambient sound;
- Qwen-Audio-3.1-Realtime: supporting full-duplex real-time communication, emotion understanding, multilingual switching, and agent tool calls.
As speech recognition, speech synthesis, audio understanding, and real-time interaction gradually converge, Qwen-Audio 3.1 is expanding voice from a mere input-output method into a comprehensive entry point connecting users, content, applications, and smart hardware.
Comments
00No comments yet. Be the first to weigh in.