A model learns from data that already exists — molecules that have been made and measured, targets that have been studied. That data is not a random sample of chemistry or biology. It is concentrated where research has already happened, which means the model is strongest on molecules and targets that resemble what it was trained on, and weakest on genuinely new ones.
The animation shows this directly. The model is trained on a limited set of examples, then meets a biological case outside that set. Its prediction is confident but unreliable, because nothing in its training tells it how to handle the new situation. The same problem appears in two forms: gaps, where no relevant data exists at all, and bias, where the data exists but over-represents certain kinds of molecules, targets, or patient populations. Both mean the model's confidence is not the same as its accuracy on new biology.