A model can be right about everything it was asked and still be wrong about the patient. The distance between a computational result and a clinical outcome is crossed by biology that no dataset fully represents.
Start with what the model actually predicts. It predicts a property of a molecule — binding, solubility, a toxicity flag — measured in a defined assay, often in cells or in animals. That is a real result. But a human body is not an assay. It has variable absorption, competing metabolic pathways, an immune system, other diseases, other drugs, and a genetic background that differs from every training sample.
Each step away from the assay adds a source of failure the model never saw. A compound that binds beautifully in a test tube may never reach the target organ. A toxicity flag that was clean in a cell line may not capture an effect that only appears after months in a whole organism. A result that holds in one population may not hold in another, because the training data underrepresented that group.
The decisive test is therefore not the model's score. It is whether the compound works in humans, and that question can only be answered by running the trial. Everything before that point — every prediction, every ranking, every in-silico result — is a reason to run the experiment, not a substitute for it. This is the same prediction-versus-experiment line from the previous chapter, extended to its final and most consequential step.