Alibaba unveils advanced audio model with real-time interaction capabilities
Alibaba’s Qwen team has launched Qwen-Audio-3.1, a comprehensive audio stack encompassing ASR, TTS, and real-time interaction. The core model, Qwen-Audio-3.1-Realtime, is a full-duplex speech model designed for voice agents that can call tools. Pricing has been reduced by up to 95% on ASR, with discounts of 70% on TTS and 85% on Realtime. The model is available as a managed API via WebSocket on QwenCloud, without open weights.
The system operates using two models with the same Audio Encoder and LLM design. A full-duplex decision model predicts whether to listen, speak, stop, or resume, while a speech-to-text model converts responses into streaming speech. Training involves three layers: Think, Act, and Speak and Coordinate, with Core-Cocktail SFT, Multimodality OPD, and GRPO used to refine the model. The model’s performance has improved across various benchmarks, including reduced filler rates and enhanced multilingual capabilities.
Pricing details were confirmed as of September 28, 2026, with token rates varying due to differing audio tokenization methods across providers. The release includes additional models like Qwen-Audio-3.1-ASR-Flash-Filetrans, which focuses on offline long-audio transcription. Alibaba also encourages engagement through social media and collaboration opportunities for partners.