Two kinds of evidence, two different claims
Retrospective evaluation
- Model is tested on antibodies already characterized in databases
- Measures whether the model recovers known binders
- Cheap, fast, and common
- Cannot show that a genuinely new antibody would work
Prospective evaluation
- Model proposes sequences that have never been made or tested
- Those sequences are expressed and assayed in the lab
- Expensive, slow, and less common
- The only design that supports a real performance claim
AI-designed antibodies have been shown to bind their targets in real experiments, but the strongest claims rest on a small number of prospective studies. Most published comparisons are retrospective, and true head-to-head trials against human-designed antibodies on the same target are rare. The defensible statement is that AI design works for some targets, not that it outperforms human design in general.
A high success rate in a retrospective study is partly a measure of how well the test set resembles the training set. When the same databases supply both, strong scores can reflect memorization of familiar antibody families rather than general design ability.