Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

Can AI Discover New Drugs? A High-Level Overview

1Why Drug Discovery Is Hard, and Where AI Fits2How AI Learns From Molecules and Proteins3Finding and Validating a Biological Target4Designing Molecules: Generative AI and Virtual Screening5From Hit to Lead: Optimizing Properties With AI6What AI Still Cannot Do7Judging the Claims: Real Successes, Failures, and Open Questions
Finding and Validating a Biological Target

Turning Scattered Evidence Into a Ranked List

2 / 4
Look at the three streams on the left. Human genetics carries the most causal weight: if a variant consistently protects people from a disease, that points to the gene being involved, not just correlated. Literature is broad but uneven — heavily studied proteins look important partly because they were studied. Omics shows where a gene is active and how it shifts across tissues, but that is association, not cause. The model's job is to merge these into one ranked list, weighting each stream by how much causal evidence it tends to carry. Watch what the ranking does not do: it does not confirm biology. A target can rank low simply because it has been under-studied, and a high rank means the evidence converged, not that the hypothesis is proven.
0:00 / 0:00

No single source of evidence settles whether a protein causes a disease. Human genetics offers the strongest signal: if people who carry a particular variant of a gene are consistently protected from or predisposed to a disease, that is evidence the gene's protein is causally involved, not merely correlated with it. Literature records decades of experimental observations, but it is uneven — heavily studied proteins accumulate evidence simply because they were studied. Omics measurements (transcriptomics, proteomics, and similar large-scale readouts) show where a gene or protein is active and how it changes across tissues and disease states, but they describe association, not causation.

AI's contribution at this stage is integration. A model takes these heterogeneous inputs — genetic association statistics, literature-derived relationships, expression and abundance measurements, pathway membership — and produces a single ranked list of candidate targets, weighting each source by how much causal evidence it tends to carry. The ranking is a prioritization, not a measurement: it says which hypotheses are worth testing first, given everything currently known.

The output inherits the weaknesses of its inputs. Literature-derived features are biased toward well-studied proteins, so a genuinely important but under-studied target can rank low. Genetic evidence is strongest where large biobanks and genome-wide association studies exist, which skews toward populations that have been sampled. A high rank means the evidence converged; it does not mean the biology is confirmed.

Previous2 / 4Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion