The value of a prediction is not constant across the pipeline. It depends on how many candidates are available to compare and how much other evidence exists for each one.
Early on, a team may hold hundreds of compounds with little more than measured properties to distinguish them. Here the model has the most to work with and the most to contribute: it can rank a large, thinly characterized set and direct scarce laboratory attention toward the candidates most likely to survive. Late in development, the situation reverses. The candidate set has narrowed to a handful, each backed by years of accumulated experimental and clinical evidence, and the model's training history contains few comparable cases. Its contribution shrinks precisely where the cost of a wrong call grows.
This is why the strongest uses of these tools cluster around early triage and portfolio-level ranking, and why they are weakest when asked to make a final call on a single advanced candidate.