Generation is cheap, so a model can propose far more candidates than anyone could ever make and test. The pipeline therefore ends with a narrowing step: candidates are scored, filtered, and ranked until a small shortlist remains for real experiments.
Scoring assigns each candidate one or more numbers — a predicted binding strength, a predicted structural quality, a plausibility score from the generative model. Filtering removes candidates that fail hard constraints: sequences that look non-antibody-like, that carry obvious liabilities, or that are too similar to something already on the list. Ranking then orders what survives, so the most promising candidates are tested first.
The funnel shape is the point. Thousands of generated sequences become hundreds after filtering, then tens after ranking, and finally a handful that a lab can actually produce and measure. Each stage discards most of its input, and that is not a failure of the pipeline; it is how the pipeline converts cheap computation into an expensive experiment that is worth running. The scores are predictions, so the shortlist is a prioritized set of hypotheses, not a set of answers. The experiment at the end of the funnel is still what decides.