Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

Can AI Discover New Drugs? A High-Level Overview

1Why Drug Discovery Is Hard, and Where AI Fits2How AI Learns From Molecules and Proteins3Finding and Validating a Biological Target4Designing Molecules: Generative AI and Virtual Screening5From Hit to Lead: Optimizing Properties With AI6What AI Still Cannot Do7Judging the Claims: Real Successes, Failures, and Open Questions
How AI Learns From Molecules and Proteins

Why the Training Data Sets the Ceiling

4 / 4
The dense cluster in the middle is where the training data lives — compounds and targets the model has actually seen. The sparse edges and the empty regions are everything else. Follow a new input that lands inside the dense cluster: the model has neighbours to compare against, and its prediction is grounded. Now follow one that lands out in the empty region. The model still produces a number, and it produces it with the same confidence as before. Nothing in the output tells you that you have left the region where the model has evidence. That is the failure mode to watch for. It is not that the model breaks; it is that it answers anyway, and the answer looks exactly like a reliable one.
0:00 / 0:00

A model learns patterns from the examples it was trained on, and it has no way to know when a new input falls outside that experience. This is the single most important practical limit on everything described so far.

Consider activity prediction. Suppose a model is trained on measured interactions between compounds and a well-studied family of proteins. It will perform well on new compounds against those proteins. Now give it a protein from a family it has never seen. The model still returns a number — it always returns a number — but that number is an extrapolation from unrelated chemistry, and it can be confidently wrong.

The same failure appears with bias rather than sparsity. If the training set contains mostly compounds from one chemical series, the model learns that series' regularities and treats them as general rules. Novel scaffolds that fall outside the training distribution get scored as if they were familiar, and the error is not flagged.

There is a second, subtler problem: the data records what was measured, not what is true. Negative results are under-reported, so a model trained on published data sees a world where most tested compounds worked. That skews its sense of what is normal.

The practical takeaway is not that these models are useless. It is that a prediction is only as trustworthy as the overlap between the new case and the training data, and that overlap is usually invisible from the output alone.

Previous4 / 4Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion