New Developments

New benchmarks and alignment startup reshape AI evaluation landscape

Cognition + Sequent + XiaomiSource: Jack Clark - Import AI15/06/2026, 08:30
This week brought significant developments across multiple research fronts in artificial intelligence. The nonprofit research organization Sequent was launched by UK-based researchers and collaborators to develop more reliable alignment techniques for superintelligent systems, arguing that current work in frontier labs remains unprepared to guarantee safety before ASI development. The organization aims to grow to 40-80 employees and raise between $100-150 million initially. Meanwhile, Cognition introduced FrontierCode, an exceptionally challenging coding benchmark designed to assess whether AI models can actually produce production-quality code. Claude Opus 4.8 achieves only 13.4% on the highest difficulty level, suggesting the benchmark will retain long-term value for measuring future progress. Finally, Xiaomi unveiled a 1 trillion parameter model capable of generating 1,000 tokens per second, demonstrating how joint optimization of software and model architecture can achieve remarkable speeds on standard hardware.
New benchmarks and alignment startup reshape AI evaluation landscape — lupAI