Three recurring weaknesses in evaluation
- Test-set leakage: because training and test data come from the same public databases, the test set is not truly unseen, and scores partly reflect familiarity.
- Proxy mismatch: benchmarks score predicted binding, while the properties that decide whether an antibody is usable — stability, manufacturability, immunogenicity — are usually not scored at all.
- Simplification: benchmark tasks fix parts of the problem, such as the framework or the target class, so they understate the difficulty of a genuinely novel target.
The gap between a proxy and the real outcome is not a technicality. A predicted binding score is computed from a model of molecular interaction; the measured affinity comes from an assay on a real protein in real buffer. Between the two sit folding, expression, aggregation, and measurement error. A benchmark that reports the first while the field needs the second will systematically overstate how close a design is to a usable drug.