The model's output is a position on a scale, and positions are shared. Any number of candidates can occupy the same position, and nothing in the model's output distinguishes them from one another. The differences that decide their real fates are differences the model was not given.
A high score is a reason to scrutinize a candidate more closely, not a reason to abandon it. A low score is a reason for ordinary confidence, not a guarantee. In both directions the score shifts attention; it does not settle the question.
The properties a model receives are a compressed description of the drug and its context. Patient population, concurrent treatments, and trial design are examples of real influences that may simply not be represented in that description.