Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

Can AI Discover New Drugs? A High-Level Overview

1Why Drug Discovery Is Hard, and Where AI Fits2How AI Learns From Molecules and Proteins3Finding and Validating a Biological Target4Designing Molecules: Generative AI and Virtual Screening5From Hit to Lead: Optimizing Properties With AI6What AI Still Cannot Do7Judging the Claims: Real Successes, Failures, and Open Questions
What AI Still Cannot Do

The Data Is Not a Neutral Sample

1 / 4
Think of the map as every molecule and every biological condition a model might be asked about. The bright clusters are where measurements actually exist — heavily studied protein families, convenient assays, well-funded areas. The dark areas are everything else. A model trained on the bright clusters reports high accuracy, but that number is an average over the bright regions. Move into the dark and the model is guessing, and it will not tell you it has crossed that line. The point is not that the data is small in absolute terms; it is that the data is uneven, and the unevenness lines up with where new drugs are most needed.
0:00 / 0:00

Models learn from what has been measured, and what has been measured is a small, uneven slice of biology. Three properties of the data matter more than its total size.

First, scarcity. The number of compounds with measured activity against a given protein is often in the hundreds or low thousands, and the number with measured human outcomes is far smaller. Compare that with the billions of molecules that could be made. A model trained on a thin slice is extrapolating over most of the space it is asked to judge.

Second, bias in coverage. Measurements cluster where past research was funded and where assays were convenient. Some protein families, some tissue types, and some patient populations are heavily represented; others are nearly absent. A model trained on the dense regions performs well there and degrades quietly in the sparse ones — and the sparse regions are often exactly where a new drug is needed.

Third, noise. Biological measurements vary between labs, between assay formats, and between days. Two datasets describing the same compound can disagree. A model cannot separate the true signal from that measurement scatter unless the scatter is characterized, and it usually is not.

The practical consequence: a model's reported accuracy is an average over the regions where data is dense. It says little about performance in the regions where data is thin, which is where the interesting predictions live.

Previous1 / 4Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion