Fireworks AI introduces firerouter with opus, cutting encoding costs by 57% with minimal accuracy drop
Fireworks AI has launched FireRouter with Opus, a cache-aware routing model designed to optimize the use of the Claude Opus series. The solution, now available via a serverless endpoint, reduces encoding costs by 57% while maintaining 98.1% accuracy, a 1.5 percentage point drop from using Opus alone. FireRouter evaluates model suitability for each task, balances cost and quality, and routes to the most efficient option. The routing pool includes Claude Opus5.5, GLM5.3, and GLM5.3Flash, with plans to expand as new models are released.
Internal testing showed a 57% cost reduction, with session costs dropping from $15.36 to $6.63. Accuracy remained high, with FireRouter scoring 78.7% compared to 80.2% for Opus alone. Cache hit rates also saw a slight decline, with FireRouter achieving 94.2% versus 97.8% for Opus. Fireworks AI noted this trade-off is intentional, as cheaper models can handle routine tasks, significantly lowering overall costs.
The test data was sourced from internal encoding traffic, with sessions randomly assigned to either the FireRouter with Opus group or the control group using Opus alone. The results highlight the effectiveness of cache-aware routing in balancing performance and cost efficiency.