Both screening and generation depend on a scoring function, and a score is a prediction, not a measurement. When predicted binding is plotted against measured binding for compounds that have actually been tested, the points scatter widely around the ideal line. Some compounds the model loves turn out to be inactive; some it dismisses turn out to bind well.
The uncertainty has a specific cause. A model is reliable where it has seen many similar examples during training, and unreliable where it has not. Truly novel molecules — exactly what generative design produces — sit in the sparse region, so their scores are the least trustworthy precisely when they matter most. A high predicted score therefore means "this is worth testing," not "this works." The only way to convert a prediction into knowledge is to run the experiment.