AI's hidden economy, the complexity of automated oversight, and protein prediction breakthroughs
Economists from leading academic institutions and the Bank of Canada have documented unprecedented growth in the US AI economy — approximately 2,600% annually when adjusted for quality improvements — yet this expansion remains largely invisible in conventional GDP statistics. Growth concentrates primarily in AI system usage rather than data center infrastructure, making it difficult to capture through traditional economic measurements. This measurement gap poses significant risks for policymakers who may underestimate labor market disruption and lack proper frameworks for forecasting AI's economic impact.
Researchers at the UK AI Security Institute warn that automated oversight of AI systems presents far greater challenges than commonly assumed. Errors made by AI agents prove difficult for humans to identify, emerging in counterintuitive ways within highly correlated contexts and potentially involving arguments humans cannot evaluate. To address these challenges, researchers propose recreating research projects from arbitrary checkpoints, testing generalization across correlated datasets, and developing scalable oversight protocols that account for human-agent team dynamics.
A collaborative effort by Stanford, Salesforce Research, and other institutions released GPIC, a dataset comprising 100 million permissively licensed images available for both academic and commercial use, hosted on Hugging Face. The dataset, which encompasses 200,000 validation examples and 1 million test examples, provides essential resources for academic researchers and emerging companies developing AI systems.
In computational biology, Biohub launched ESMFold2, directly competing with DeepMind's AlphaFold through superior performance on multiple benchmarks. The system combines ESMC, a language model trained on 2.8 billion protein sequences, with ESMFold2's design capabilities for generating three-dimensional structures, and ESM Atlas, which catalogs representations of 6.8 billion protein sequences and 1.1 billion predicted structures — the largest AI application to protein biology to date.