The scoring criteria conflict. Improving binding affinity usually degrades drug-likeness or synthesizability, so the output is a ranked set of compromises rather than a single best molecule.
One change, two consequences
Suppose a candidate binds weakly because it is small and not greasy enough. Adding a bulky greasy group can tighten the fit in the pocket and raise the predicted affinity. The same group raises the molecule's greasiness, which works against the drug-likeness rules, and adds a step or two to the synthesis. The affinity score improves and two other criteria get worse. Nothing went wrong — this is the normal shape of the problem.
Why a high score is not a guarantee
Predictions are extrapolations from training data. When a candidate lies outside the chemistry the model has seen, the prediction is unreliable even though it looks precise. Targets also change shape when a molecule binds, and a model trained on static structures may not capture that. And the property predicted may not be the property that decides whether the molecule works in a living system.
The criteria in play
- Predicted binding affinity — how tightly the candidate is expected to hold the target
- Drug-likeness — size, greasiness, and polar group counts consistent with being absorbed
- Safety — absence of structural features linked to reactivity or toxicity
- Synthesizability — whether a chemist can build it in a practical number of steps