New Developments

Anthropic adds three new benchmarks to Claude Opus 5

Anthropic + Claude Opus 5Source: Edwin Chen (X)24/07/2026, 19:17
Anthropic has integrated new performance tests into the Claude Opus 5 system card. The benchmarks—Riemann-bench, Chartography, and GDP.pdf—evaluate specific capabilities of the AI assistant across various technical domains. The addition of these evaluation criteria broadens the assessment framework for measuring model performance across multiple application contexts.
Anthropic adds three new benchmarks to Claude Opus 5 — lupAI