Samaya releases FrontierFinance benchmark for evaluating AI agents in investment workflows
Samaya unveiled FrontierFinance, a benchmark comprising 220 test cases with 11,543 expert-crafted rubrics covering the full investment cycle: company screening, research, sector analysis, earnings events, and catalyst monitoring.
Unlike existing finance benchmarks focused mainly on data extraction, FrontierFinance targets complex, ambiguous long-term decision scenarios. Samaya's AI system achieved state-of-the-art accuracy of 50.8% with 4x lower computational costs than Claude Fable 5 (49.2%), which placed second. Claude Opus 4.8 scored 45% and GPT 5.5 43.5%.
The benchmark, methodology, and complete evaluation results are now publicly available. Samaya plans to release increasingly difficult versions derived from its internal dataset of approximately 5,000 examples.