The front-loaded pattern has a single underlying cause: the kind of data each stage produces.
At the discovery and design stage, the working material is already digital. Protein structures, chemical structures, and measured properties live in databases, some of them public and containing millions of entries. A model can be trained on that material and can be retrained as more is added. The cost of one more data point is close to zero, because the data was collected for other reasons and is simply being reused.
Preclinical testing changes the picture. Here the data comes from experiments in cells and animals, and each experiment has to be designed, run, and paid for. The volume is smaller and the cost per data point is much higher. AI can still help — for example, by predicting which compounds are worth testing first — but it is working with less material and cannot generate more of it on demand.
Clinical trials change it again. The data comes from human volunteers and patients, and it arrives over years, in limited numbers, under strict rules. You cannot run a trial faster by writing better software, because the constraint is the availability of people and the time it takes for a disease to progress. Regulatory approval and post-approval monitoring sit at the far end of the same gradient: the evidence is produced by clinicians and patients, and it accumulates slowly.
So the gradient in the diagram is really a data gradient. Where data is abundant and cheap, AI has room to work. Where data is scarce and expensive, its influence shrinks — not because the later stages matter less, but because there is less for a model to learn from.