Anthropic adds three new benchmarks to Claude Opus 5
Anthropic has integrated new performance tests into the Claude Opus 5 system card. The benchmarks—Riemann-bench, Chartography, and GDP.pdf—evaluate specific capabilities of the AI assistant across various technical domains.
The addition of these evaluation criteria broadens the assessment framework for measuring model performance across multiple application contexts.