PrismML launches bonsai 2 27b with 98.2% performance at 1/9 memory usage
On September 18, 2026, PrismML announced the release of Bonsai 2 27B, a model that achieves 98.2% of the performance of the Qwen3.8 27B base model while using less than 1/9 of the memory. This follows the earlier launch of Bonsai 27B in early 2026, which already demonstrated 95% of the performance of Qwen3.6 27B with reduced memory usage.
The new version, featuring three-value quantization, supports a 262K context window and operates at 143 TPS on an NVIDIA GeForce RTX 5090 and 46.8 TPS on an Apple M5 Max. It also boasts an energy efficiency of 0.714 mWh per token, 40% lower than a full-precision 8B model.
PrismML claims the model can handle complex tasks like programming, document parsing, and multimodal debugging locally, with cloud API calls only when necessary. The model is available under the Apache 2.0 license, supporting both NVIDIA and Apple hardware through CUDA and MLX respectively.