A model's reliability is local to the chemical and biological territory of its training drugs. Performance on a familiar class does not transfer to a genuinely new one, and the model will not signal that it has left its territory.
Two reasons the pattern breaks down
The same measured property can carry a different meaning in a new chemical family, so a feature the model treats as a warning sign may be neutral or even favorable there. Separately, a new class is usually underrepresented in the history, leaving too few comparable past drugs for a stable estimate. Both problems produce a normal-looking score rather than an obvious error.
A concrete illustration
Suppose a model was built mostly from oral small-molecule drugs for metabolic disease, where a certain molecular feature often accompanied liver-related failure. A biologic for an autoimmune condition sits far outside that territory: the feature may not exist in the same form, and the comparison group of similar past drugs may be a handful of entries. The model still returns a number, and that number looks no different from a well-supported one — which is exactly why the territory has to be checked by a person, not inferred from the score.