Data: fragmented, private, and unrepresentative
The information needed to train and test these tools is scattered across hospitals, companies, and countries, and much of it cannot be freely shared because it belongs to patients. Even when data is pooled, it tends to over-represent people who had access to care in the first place. A model learns the population it was trained on, so gaps in the data become gaps in the predictions.
Ethics: bias becomes a decision about a person
If a tool is used to decide who gets into a trial or who receives a scarce treatment, then the biases in its training data become decisions about real people. The question of who is included in the data and who benefits from the result is not a technical afterthought; it is part of whether the tool should be used at all.
Validation: a prediction is a hypothesis until tested in people
A model can look excellent on historical data and still fail in a real trial, because the past is not the same as the future and a statistical pattern is not the same as a biological cause. This is why the field still measures success in trial phases, not in model accuracy.
These three problems are connected. Weak data produces unreliable predictions, unreliable predictions used on people raise ethical questions, and only real trials can settle whether the prediction was right.