Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

AI Antibody Design vs Human Design: A Conceptual Overview

1Why Antibody Design Matters2How Humans Design Antibodies3How AI Learns to Design Antibodies4Comparing AI and Human Design5Evidence, Limits, and Open Questions
How AI Learns to Design Antibodies

Where the Model's Knowledge Comes From

1 / 4
Look at the two boxes on the left. The upper one is sequence data: millions of antibody amino acid sequences, mostly natural ones. The lower one is structure data: three-dimensional coordinates of antibody bound to antigen, and there are far fewer of these because each one took a costly experiment. Both flow into the model in the middle, and that is the whole point of the picture. The model has no independent knowledge of antibodies; it only has what these collections contain. So when you see a confident prediction later, remember it is a statement about patterns in this data, not about biology in general. The sequence box is large but shallow, and the structure box is small but deep.
0:00 / 0:00

An AI model for antibody design is not born knowing anything about antibodies. It learns from collections of antibodies that other people already determined. Two kinds of collection dominate.

Sequence databases hold the amino acid sequences of antibodies, mostly from natural human repertoires and from antibodies that have been studied or developed. They are large, and they are heavily weighted toward sequences that were easy to obtain or interesting enough to publish. Structure databases hold three-dimensional coordinates of antibody–antigen complexes, obtained by X-ray crystallography or cryo-electron microscopy. They are far smaller, because determining a structure is slow and expensive, but they carry the information that sequences alone do not: which residues actually sit at the interface and how the two surfaces fit together.

Both kinds of data are biased. Natural repertoires over-represent common germline families and common binding problems; solved structures over-represent antibodies that were stable enough to crystallize and interesting enough to fund. A model trained on this material inherits those biases, so it tends to be strong on antibody-like sequences and weak on genuinely unusual ones. Data quality therefore matters as much as data quantity: a model cannot learn a pattern that the data never contained, and it will confidently reproduce patterns that the data contained too often.

Previous1 / 4Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion