Models learn from what has been measured, and what has been measured is a small, uneven slice of biology. Three properties of the data matter more than its total size.
First, scarcity. The number of compounds with measured activity against a given protein is often in the hundreds or low thousands, and the number with measured human outcomes is far smaller. Compare that with the billions of molecules that could be made. A model trained on a thin slice is extrapolating over most of the space it is asked to judge.
Second, bias in coverage. Measurements cluster where past research was funded and where assays were convenient. Some protein families, some tissue types, and some patient populations are heavily represented; others are nearly absent. A model trained on the dense regions performs well there and degrades quietly in the sparse ones — and the sparse regions are often exactly where a new drug is needed.
Third, noise. Biological measurements vary between labs, between assay formats, and between days. Two datasets describing the same compound can disagree. A model cannot separate the true signal from that measurement scatter unless the scatter is characterized, and it usually is not.
The practical consequence: a model's reported accuracy is an average over the regions where data is dense. It says little about performance in the regions where data is thin, which is where the interesting predictions live.