No single kind of information is enough to decide whether a cancer drug candidate is worth pursuing. Four broad categories of data are used together, and they come from very different sources.
Genetic data describes the DNA of tumors and of patients — which genes are altered, and how those alterations differ between people. Protein data describes the molecules that actually carry out cell functions, since genes mostly matter through the proteins they produce. Chemical data describes the candidate compounds themselves: their structure, how they behave, and how they interact with proteins. Patient data describes what happened to real people — diagnoses, treatments, responses, and side effects.
Each category has its own format, its own scale, and its own noise. Genetic and protein data are biological and highly variable; chemical data is structured and precise; patient data is messy and often incomplete. A recurring difficulty in cancer drug discovery is that a decision usually requires evidence from several of these categories at once, and they do not naturally line up with each other.