A model's output is a statement about similarity to past data, not a statement about biology. That gap is why experiments stay in the loop and why AI's practical role is to prioritize experiments rather than replace them.
Three limits that keep experiments essential
- Bias: a model trained on one population or cancer subtype loses accuracy on others, often without any visible warning.
- Uncertainty: scores are probabilities, and a confident-looking number can rest on very little evidence.
- The prediction–reality gap: a statistical pattern in data is not the same as a mechanism in a living tumor, and biology can change underneath it.
What a high score actually tells you
Suppose a model scores a new compound at 0.92 for activity against a target. The honest reading is: this compound resembles the active compounds in the training set more than it resembles the inactive ones. It does not say the compound will bind, that it will reach a tumor, that it will be safe, or that the resemblance is causal. If the training set happened to contain many compounds from one chemical family, the model may simply be recognizing that family. The score earns the compound a place in the test queue — nothing more.
The practical takeaway
Treat AI as a way to spend scarce laboratory capacity well. It is most useful where the candidate space is enormous and experiments are expensive — which is precisely the situation in cancer drug discovery.