lupAI
analises

China's AI models surpass U.S. in weekly token usage, maintaining 21-week lead

DeepSeek V4.1-Flash + Tencent HuanYuan + Zhipu GLM5.3FlashSource: AIBase21/09/2026, 03:14
According to OpenRouter's latest data, global AI large model token usage from September 14 to 20 totaled 129 trillion, a 1.57% increase from the prior week. China's models accounted for 67.46 trillion tokens, up 10.28% weekly, while U.S. models used 14.21 trillion tokens, a 34.7% drop. This marks China's 21st consecutive week of higher token usage than the U.S. In the global rankings, four of the top five models were Chinese. DeepSeek V4.1-Flash led with 15.8 trillion tokens, a 219% weekly increase. Released on September 10, the model features improved coding, agent, and multimodal understanding capabilities. It uses a Causal-Encoder-Decoder architecture, with a global KV Cache reduced to about 1/4 of the previous version. The cost for input cache hit, miss, and output per million tokens during idle time are 0.02 yuan, 1 yuan, and 4 yuan, representing 60%, 33.3%, and 11.1% reductions compared to the prior generation. Artificial Analysis testing found that V4.1-Flash uses approximately 89,000 output tokens per task on average, higher than the 62,000 tokens of V4Flash0731, and is considered relatively verbose in output. Zhipu GLM5.3Flash ranked second with 14.1 trillion tokens, Tencent HuanYuan Hy4preview third with 12.5 trillion tokens, and DeepSeek-V4-Flash-0731 fifth with 9.44 trillion tokens.