A claim that AI-designed antibodies outperform human-designed ones can be checked against a small set of questions, and the answers determine how much weight the claim deserves.
First, what is the comparison baseline? A claim of superiority needs a human-designed or human-discovered antibody evaluated on the same target under the same conditions. If the baseline is a random sequence, a weak historical binder, or a different target, the comparison does not support the claim.
Second, is the test set independent of the training data? If the antibodies used for evaluation could have appeared in the model's training data, a high score may reflect familiarity. The stronger claim requires targets and sequences the model has genuinely not seen.
Third, was the validation prospective? A prediction confirmed by making and assaying the antibody is real evidence; a prediction confirmed only by another model is not.
Fourth, were the properties that decide usability measured? Binding alone is not enough. Stability, expression, and immunogenicity determine whether a candidate is a usable molecule, and a claim that ignores them is incomplete.
Fifth, how large and how varied was the test? A result on one target or a handful of antibodies is a data point, not a general finding. The more targets and the more diverse the targets, the more the claim can bear.