AI agents now capable of performing economically valuable work
OpenAI released a new benchmark measuring AI performance against human experts on realistic tasks requiring four to seven hours to complete. Results show AI models approaching human-level capability across various sectors, with margins varying by industry. Surprisingly, the primary reason for AI underperformance was not hallucinations or errors, but poor result formatting and instruction-following—areas showing rapid improvement.
While AI's ability to execute individual tasks does not translate to immediate job replacement—since jobs involve multiple functions—there are areas where AI already delivers significant value. Scientific research replication exemplifies this potential: a process normally requiring many hours of painstaking work can now be automated. Models such as Claude Sonnet 4.5 have demonstrated capacity to independently analyze complex papers, convert statistical code, and reproduce findings accurately.
AI agents have made dramatic improvements in autonomous work. Small gains in model accuracy translate into exponential growth in task scope they can complete. Latest-generation models are self-correcting, enabling error recovery without complete failure.
A significant risk exists: deploying AI thoughtlessly to generate unnecessary content. The solution involves experts working deliberately with these agents—delegating tasks as first attempt and reviewing results. Research suggests this workflow improves productivity by 40% while reducing costs by 60%, keeping experts in control.