The number of chemically plausible small molecules is astronomically large — far more than could ever be made and tested. Generation models work inside that space. A common approach builds a molecule one fragment at a time: the model starts from an atom or a small group, then repeatedly chooses the next fragment and the bond that attaches it, guided by what it learned from structures of known drugs and known inactive compounds. The result is a structure that has never existed, but whose local chemistry resembles things that do.
Generation alone would be useless, because most generated structures are unusable. So a second set of models filters. One scores whether the molecule looks like a drug at all — size, solubility, the balance of water-loving and water-repelling regions, whether it resembles compounds that have survived development. Another estimates whether it could actually be synthesized, since a beautiful molecule that no chemist can make is worthless. A third flags chemical groups associated with toxicity or instability. Some of these filters are learned models; others are explicit rules distilled from decades of medicinal chemistry.
The two halves are usually run as a loop rather than a single pass. The generator proposes, the filters score, the scores steer the next round of proposals toward regions of chemical space that survive filtering, and the cycle repeats. What comes out is not one answer but a ranked shortlist — perhaps a few dozen to a few hundred structures — that a chemistry team can examine.
The shortlist is a prioritization, not a verdict. Every entry is still a hypothesis about a molecule that has never been made, and the filters can only recognize what they have seen before. A genuinely novel scaffold that does not resemble anything in the training data may be scored poorly and discarded, or scored well for the wrong reasons.