An AI model for antibody design is not born knowing anything about antibodies. It learns from collections of antibodies that other people already determined. Two kinds of collection dominate.
Sequence databases hold the amino acid sequences of antibodies, mostly from natural human repertoires and from antibodies that have been studied or developed. They are large, and they are heavily weighted toward sequences that were easy to obtain or interesting enough to publish. Structure databases hold three-dimensional coordinates of antibody–antigen complexes, obtained by X-ray crystallography or cryo-electron microscopy. They are far smaller, because determining a structure is slow and expensive, but they carry the information that sequences alone do not: which residues actually sit at the interface and how the two surfaces fit together.
Both kinds of data are biased. Natural repertoires over-represent common germline families and common binding problems; solved structures over-represent antibodies that were stable enough to crystallize and interesting enough to fund. A model trained on this material inherits those biases, so it tends to be strong on antibody-like sequences and weak on genuinely unusual ones. Data quality therefore matters as much as data quantity: a model cannot learn a pattern that the data never contained, and it will confidently reproduce patterns that the data contained too often.