Zhipu AI unveils GLM-5.3-FlashX with enhanced speed and efficiency
Zhipu AI has launched the GLM-5.3-FlashX large model on its BigModel platform, with the API now available. The model achieves a maximum output speed of 200 tokens per second, emphasizing intelligence, affordability, and speed. It offers high-throughput, low-latency inference services to enterprises and developers.
The previous version, GLM-5.3-Flash, was previously known as 'Ox Alpha' overseas and was praised for its strong performance and cost-effectiveness. To meet rising demand, Zhipu leveraged a domestic chip computing power base of 100,000 units and increased investment in inference optimization, leading to the release of the enhanced FlashX version.
FlashX maintains the original model's intelligence and pricing while maximizing generation speed, enhancing its commercial appeal in high-concurrency scenarios. Developers can access it via the official API, while non-technical users can try it in the experience center. This upgrade marks progress in domestic large model inference efficiency and accessibility.