Quality data: the foundation for AI-driven pharmaceutical discovery
Source: MIT Technology Review - AI27/07/2026, 08:40
Drug discovery is a costly and high-risk endeavor. Since the 1950s, development costs have doubled roughly every nine years. Today, bringing a drug to market takes 10 to 15 years, costs between $1 and $2.5 billion, and faces failure rates exceeding 90%. The pharmaceutical industry views artificial intelligence as its primary lever to compress timelines and boost success rates.
One of the most promising applications uses AI to identify molecular candidates against disease targets. Rather than extensive physical screening in laboratories, companies now design candidates computationally and predict their interactions before experimental testing. This eliminates volume limitations. However, laboratory validation remains essential, as models cannot yet reliably predict how compounds will behave. This intensifies pressure on research teams, who must test a growing volume of AI-generated candidates.
The central challenge is data quality. Models trained on limited public datasets experience diminishing returns because researchers access the same information. These datasets also lack the structure and diversity needed to prevent bias. A critical problem is publication bias: researchers share successes but hide failures in lab notebooks, denying models the essential understanding of what doesn't work. Without comprehensive access to failure data, models cannot train adequately.
Data integrity faces additional threats. Generative AI tools have made scientific manipulation trivially easy. Earlier studies found that nearly 4% of biomedical publications contain fabricated images, numbers that could rise dramatically with AI. Emerging solutions employ cryptographic hash algorithms, similar to blockchain technology, to detect manipulated images, allowing publishers to verify authenticity.
The future points toward fully autonomous laboratories operating continuously, cycling through prediction, experimentation, and optimization, feeding results back into AI models. This could significantly improve success rates of candidates entering clinical trials. However, it requires highly integrated infrastructure with interoperable systems, comprehensive structured data, and continuous information flow. Many laboratories lack this integration. Notably, no drug discovered primarily through AI design has yet received full FDA approval, though changes are expected in the next two to three years.