Strata engine enables 125b Qwen3.8 model on 12gb GPU with 94 tokens per second performance
A new engine called Strata has been released by developer Niko1221, allowing the 125B parameter Qwen3.8-Flash-Next model to run on consumer-grade GPUs with 12GB or more VRAM. The model, developed by Alibaba's Qwen team in August 2026, is a compressed version of a multimodal MoE model with 125B parameters and a 51B parameter n-gram embedding table.
It supports up to 262K token context length. Strata reduces memory usage by loading the entire MoE model into RAM and only the frequently used experts into VRAM. It also uses lightweight models for speculative decoding to speed up inference.
On an RTX 5070 with 12GB VRAM, the 2-bit quantized version of the model achieves 94 tokens per second, while the 3-bit version reaches 53 tokens per second. The AMD RX 9070 XT with 16GB VRAM also shows performance with the 2-bit version at 60 tokens per second.