The output of a failure-prediction model is usually a number between 0 and 1, or a position on a scale from low to high risk. It is easy to read that number as a probability that this particular drug will fail. That reading is too strong.
The score is built from the historical record. It reflects how often drugs that looked like this one — similar structure, similar assay profile, similar position in the evidence — went on to fail in the past. A score of 0.7 means that among the comparable drugs in the training history, roughly seven in ten failed. It does not mean this drug has a seventy percent chance of failing, because this drug is not a random draw from that history; it is a specific candidate with its own biology.
Two further qualifications follow. First, the score is only as meaningful as the comparison group behind it: if few past drugs resemble this candidate, the score rests on thin evidence and should be treated as uncertain. Second, the scale is not calibrated across different models or different disease areas. A 0.7 from one model and a 0.7 from another are not the same claim, and neither transfers cleanly to a disease area the training data barely covered.
What the score is genuinely good for is ranking. If candidate A scores higher than candidate B on the same model and the same evidence, the model is saying A resembles past failures more than B does. That comparative reading survives the caveats above far better than any absolute interpretation.