A model does not begin with any knowledge about drugs. It begins with a large set of finished drug programs, each described by the same measured properties — chemical structure, assay results, and the other evidence streams — and each carrying a known outcome: reached patients, or abandoned.
The model makes an initial guess about each past drug. That guess is compared with what actually happened, and the internal settings are nudged so the next guess is slightly closer. Repeating this across thousands of examples is the whole of the learning process. Nothing is memorized as a lookup entry; what accumulates is a set of internal settings that, taken together, separate the drugs that failed from the drugs that succeeded.
That distinction matters for what comes next. Because the pattern lives in the settings rather than in a stored list, it can be applied to a candidate that has never existed before. The model is not retrieving a similar past drug and reporting its fate. It is placing the new candidate at a position defined by the pattern it extracted.
The pattern is also only as good as the examples behind it. If the finished programs were mostly from one disease area or one era of chemistry, the settings will encode that slice of history, and the separation they produce will be sharpest there.