Huawei unveils Spark-Audio-1.0-Preview, a domestic voice foundation model
On September 16, 2026, iFLYTEK announced the release of Spark-Audio-1.0-Preview, a voice foundation model trained entirely on domestic computing power.
This model addresses longstanding issues with traditional audio processing, which relied on cascading systems that converted speech to text before analysis, leading to information loss and fragmented workflows. Spark-Audio-1.0-Preview directly processes audio, understanding speech, environment sounds, and emotions without intermediate steps.
It features a 0.65B dense audio encoder and a 30B-A3B MoE language model, trained on 13 million hours of audio and text data.
The model supports 99 languages and 202 dialects, outperforming competitors like Qwen3.5-omni-flash in multilingual tasks and achieving state-of-the-art results on the Fleurs dataset. iFLYTEK claims the model performs well in noisy and low-volume conditions, and it is now available for public testing with API access planned for the open platform.