No single source of evidence settles whether a protein causes a disease. Human genetics offers the strongest signal: if people who carry a particular variant of a gene are consistently protected from or predisposed to a disease, that is evidence the gene's protein is causally involved, not merely correlated with it. Literature records decades of experimental observations, but it is uneven — heavily studied proteins accumulate evidence simply because they were studied. Omics measurements (transcriptomics, proteomics, and similar large-scale readouts) show where a gene or protein is active and how it changes across tissues and disease states, but they describe association, not causation.
AI's contribution at this stage is integration. A model takes these heterogeneous inputs — genetic association statistics, literature-derived relationships, expression and abundance measurements, pathway membership — and produces a single ranked list of candidate targets, weighting each source by how much causal evidence it tends to carry. The ranking is a prioritization, not a measurement: it says which hypotheses are worth testing first, given everything currently known.
The output inherits the weaknesses of its inputs. Literature-derived features are biased toward well-studied proteins, so a genuinely important but under-studied target can rank low. Genetic evidence is strongest where large biobanks and genome-wide association studies exist, which skews toward populations that have been sampled. A high rank means the evidence converged; it does not mean the biology is confirmed.