Every claim about an AI-discovered drug can be placed on a ladder of evidence, and its weight is set by the rung it occupies rather than by the sophistication of the method behind it. The lowest rung is a computational prediction: a model proposes a molecule or a binding score, and nothing has been made or measured. The next rung is a synthesized compound tested in a laboratory assay, which establishes that the molecule does something measurable in a defined system. Above that sits testing in animals, where absorption, metabolism, and toxicity begin to appear as real obstacles rather than predictions. The fourth rung is a human trial designed primarily for safety and tolerability in a small number of participants. Only the top rung — a controlled human trial measuring whether patients actually benefit — supports the claim that a drug works.
A practical checklist follows from this ladder. First, what was measured, and in what system? Second, what was it compared against — an existing treatment, a conventionally designed molecule, or nothing? Third, how far up the ladder has the program actually climbed, as opposed to how far the announcement implies? Fourth, is the emphasis on the method used or on the outcome achieved? Fifth, what is missing: the number of failed candidates, the elapsed time, the presence of a control group, peer review?
Sorting a claim is not a judgment about whether the science is good. A well-executed computational study can be excellent science and still be weak evidence that a drug will work, because those are different questions.