lupAI
novos-apps

Alibaba unveils Qwen3.8-Omni-Flash multimodal model with enhanced audio-visual capabilities

Alibaba + Qwen3.8-Omni-FlashSource: ITHome, MarkTechPost, AIBase, AIBase18/09/2026, 02:59
Alibaba has launched the Qwen3.8-Omni-Flash multimodal model, capable of processing text, images, audio, and video inputs simultaneously, with support for up to 1 million tokens. The model shows significant improvements over its predecessor, Qwen3.5-Omni-Plus, with an average score increase of over 26% across 30 benchmarks. It excels in areas like audio-video agents, coding, and long-term tasks, with notable gains in specific benchmarks. Alibaba claims its audio capabilities surpass those of Gemini 3.8 Flash, and API costs for audio and video inputs have dropped by over 93% and 98%, respectively. The model also supports extended audio-video processing, with new plugins and tools for real-time interaction and long-form content analysis.
Alibaba unveils Qwen3.8-Omni-Flash multimodal model with enhanced audio-visual capabilities — lupAI