lupAI
novos-desenvolvimentos

Bottlecap AI launches ThinkingCap-Qwen3.8-27B with reduced token usage

BottleCap AISource: MarkTechPost25/09/2026, 10:51
BottleCap AI has introduced ThinkingCap-Qwen3.8-27B, a refined version of the Qwen3.8-27B model, designed to minimize reasoning tokens without compromising core capabilities. The model achieves a 37.2% reduction in average thinking tokens across 12 benchmarks, though macro-average accuracy drops by 0.86pp to 85.79%. Deployable via vLLM or SGLang with FP8, NVFP4, GGUF, and MLX builds, the model is available under a gated repository with commercial use requiring a BottleCap agreement. The focus of the update was on reducing unnecessary token usage while preserving reasoning ability, instruction following, and safety. Key benchmarks show mixed results, with significant token cuts in knowledge and multilingual tasks, such as MMMLU and MMLU-Pro. However, the model maintains performance in long-context retrieval and agentic tasks, with only minor accuracy drops. The most notable trade-off is in AIME 2026, where accuracy falls by 3.85pp for a 30.2% token reduction. BottleCap recommends using the xhigh reasoning-effort setting for optimal accuracy-to-token balance.
Bottlecap AI launches ThinkingCap-Qwen3.8-27B with reduced token usage — lupAI