An Antibody Sequence Is Not a Research Plan: Antibody Foundation Models in 2026
“Humanise this antibody, keep its binding, improve its developability, and tell me which model to use.”
That sounds like one request. It is at least five. Is the input a heavy chain, a light chain, a native pair, a nanobody, or a structure with its antigen? Which residues are framework and which are complementarity-determining regions? Does “humanise” mean a higher human-repertoire likelihood or a measured immune response? Does “developability” mean expression, aggregation, polyspecificity, thermal stability, viscosity, or one named assay?
The model name cannot answer those questions. The experimental contract has to.
Part 2, A Protein Embedding Is Not an Explanation separated a pretrained representation from its pooling rule, adaptation method, downstream head and split. Part 3, A Generated Protein Is Not a Result separated a computational candidate from an experimentally tested result. Antibody modelling makes both distinctions unavoidable. The same sequence can be embedded, folded, infilled, scored for “nativeness”, or passed to a property predictor. Those outputs answer different questions.
The central decision is therefore not general model or antibody model? It is which biological unit, task, evidence standard and runnable artifact match the research question? This is a guide to making that decision for research use. It is not medical guidance or a claim about therapeutic efficacy.
One sequence, four missing contexts
An antibody variable region is not generic protein text with a special prefix. A conventional binding site is formed by paired heavy-chain and light-chain variable domains. Each chain contains relatively conserved framework regions interrupted by three complementarity-determining regions, or CDRs. V(D)J recombination assembles germline gene segments; somatic hypermutation adds further variation. CDR H3 is especially diverse and difficult to model, but it does not act alone. The light chain and residues outside a chosen CDR can alter binding properties (Olsen, Boyles and Deane; Leem et al.).
The first useful model diagram is therefore an input contract, not an architecture chart.
Figure 1. Original editorial synthesis of the antibody input contract. Region labels, pairing, provenance and antigen context are experimental inputs, not metadata to reconstruct after prediction. Biological and data concepts derive from the OAS update, AntiBERTa study, native-pairing study, and AntiFold input specification.
Numbering matters because “CDR” is not one universally indexed slice. IMGT, Kabat and Chothia place some boundaries differently (ANARCI). AntiFold, for example, expects IMGT-numbered variable-domain structures and can sample named IMGT regions (official repository). Passing Kabat positions under the same labels changes the design mask. Pin the numbering scheme and software version before creating labels or mutation constraints.
The available corpus also shapes the representation. The 2022 Observed Antibody Space update reported 1.5 billion unpaired sequences from 80 studies, but only 121,838 paired records from five studies at that time (OAS update). Scale and biological completeness pull in opposite directions. Separate heavy and light chains provide enormous statistical coverage. They do not reveal which chains met in the same B cell.
Topology changes the task too. A paired antibody, a nanobody and a T-cell receptor do not present the same chain arrangement to a predictor.

Figure 2. Immune-protein family and chain topology change the prediction problem. Source: Brennan Abanades et al., “ImmuneBuilder: Deep-Learning models for predicting the structures of immune proteins”, Fig. 1. Licensed under CC BY 4.0. Embedded image component reproduced without pixel edits.
A runnable input can still be the wrong experiment. Concatenating arbitrary heavy and light chains produces a tensor, but not a native pair. Feeding a VHH through a paired-Fv system may produce coordinates, but not evidence that the task contract was respected.
Representation: ask what the embedding was trained to make easy
Masked language models such as AntiBERTa, AntiBERTy and AbLang learn to recover hidden residues from sequence context. Their outputs may include one vector per residue, a pooled sequence vector, attention maps, and masked-residue probabilities. Those artifacts support different computations. AntiBERTa’s paratope experiment used supervised fine-tuning; it was not a zero-shot property of the encoder (Leem et al.). AbLang exposes residue encodings, mean-pooled sequence encodings and amino-acid likelihoods, while its headline evaluation concerned restoring missing residues, a task closely aligned with masked pretraining (Olsen, Moal and Deane).
These are useful capabilities. They are not evidence that an embedding explains specificity or affinity.
General protein models such as ESM-2 and ProtT5 remain essential controls. They bring broad protein statistics and can supply frozen features or an initialization for adaptation. In a controlled native-pairing study, the authors reported benefits from paired training and also found that ESM-2 could acquire antibody-pair features through fine-tuning (Burbach and Briney). Specialisation can enter through the corpus, objective, input serialization, positional scheme or downstream architecture. It need not mean training every parameter from scratch.
The system boundary is where many comparisons go wrong.
Figure 3. Original editorial synthesis of an antibody representation experiment. The output belongs to the complete extraction and adaptation recipe. Sources: AntiBERTa, AbLang, AntiBERTy, and the native-pairing study.
For a fair test, pin chain input, separators, numbering, truncation and special tokens. Prespecify layers and whether weights are frozen. Hold pooling, head, hyperparameter budget and split constant. Compare with one-hot or physicochemical features and a general protein model. If several components change together, call the result a system comparison.
Pairing deserves the same discipline. Embedding each chain separately and concatenating vectors is not the same intervention as allowing attention across a native pair. If native pairing is unavailable, say so and bound the claim to an unpaired-chain task.
Structure: score the loops, not only the scaffold
Structure prediction is a separate system branch. IgFold combines antibody-language-model representations with geometric networks to propose antibody coordinates. ImmuneBuilder provides family-specific predictors for antibodies, nanobodies and T-cell receptors. BALMFold couples an antibody language model to a structure head. In each case, the coordinate predictor, refinement procedure and confidence estimator matter alongside the upstream representation (IgFold; ImmuneBuilder; BALM).
An evaluation should separate framework regions, individual CDR loops and especially CDR H3. A low whole-Fv error can be dominated by conserved framework residues while the loop relevant to the question is wrong. The IgFold study evaluated structures deposited after its training cutoff and reported region-level error rather than implying that all regions were equally solved (Ruffolo et al.).

Figure 4. Framework agreement can coexist with local disagreement in difficult CDR geometries. Source: Jeffrey A. Ruffolo et al., “Fast, accurate antibody structure prediction from deep learning on massive set of natural antibodies”, Fig. 2 component. Licensed under CC BY 4.0. Extracted structural-overlay component from Fig. 2; no scientific content within the retained component was altered.
The output is narrower than it looks. A plausible unbound Fv does not establish a bound loop conformation, epitope, docking pose, affinity or specificity. A confidence score estimates model-defined uncertainty; it is not a calibrated probability that a proposed interaction is biologically correct.
Compare predictors on the same targets, chain treatment, templates, refinement protocol and region definitions. Apply temporal separation and remove close structural homologues. Keep missing predictions and failures in the denominator. If antigen context matters, evaluate the complex task instead of silently treating an antibody-only fold as a binding prediction.
CDR design: three transactions, one misleading verb
“Design CDR H3” can describe at least three transactions. Sequence infilling proposes masked residues from surrounding sequence. Backbone-conditioned inverse folding proposes residues compatible with supplied coordinates. Assay-specific prediction ranks variants for a measured endpoint, possibly with antigen context. These computations overlap, but none implies the others.
AntiFold begins with inverse folding. It was fine-tuned from general ESM-IF1 on solved and predicted antibody structures, can optionally include an antigen chain, and returns residue log-likelihoods or sampled sequences for selected regions (Høie et al.; repository). AbMPNN adapts ProteinMPNN to antibody structures; its released archive includes weights and split files (Dreyer et al.; Zenodo artifact).
Both make the proxy ladder visible. Native-sequence recovery asks whether observed residues receive high probability. Refolding asks whether a second predictor agrees with the requested backbone. Neither measures antigen binding. Predicted structures used for training may also carry the biases of the predictor that generated them.
Figure 5. Original editorial synthesis of the CDR-design evidence funnel. The computational-to-experimental boundary follows the task distinctions in AntiFold, AbMPNN, and the prospective design tasks in the AIntibody challenge.
S2ALM takes another route. Its peer-reviewed 2025 paper describes an ESM-2-derived model trained with sequence-structure matching and cross-level reconstruction. Its CDR task is computational sequence infilling, and its benchmark results are developer-reported (Yin et al.). The paper contributes evidence about a training method. It does not provide wet-lab validation of generated CDRs.
It also demonstrates why publication and reproducibility need separate columns. As retrieved on 28 August 2026, the public S2ALM repository declares Apache-2.0 for its contents. Its tree contains a README, licence and five figures, but no runnable S2ALM implementation, environment, evaluation scripts or identifiable S2ALM checkpoint. The README’s downloads are labelled for the separate ProtET project. Those links do not supply the missing S2ALM artifacts. The paper is peer-reviewed, but its results cannot be reproduced from this repository alone as retrieved.
Humanisation: resemblance is not immunogenicity
Humanisation systems try to move a non-human or less human-like variable region toward patterns observed in human repertoires while respecting chosen constraints. Sapiens, exposed through BioPhi, proposes residue changes using models trained on human repertoire sequences (Prihoda et al.). AbNatiV uses a learned nativeness score for hit-selection and humanisation workflows (Ramon et al.). The safe interpretation is distributional: how compatible is a sequence, region or mutation with the model’s selected repertoire distribution?
That can prioritize edits. It does not predict immunogenicity, retained binding, expression or stability. Germline usage, donor composition and study processing shape the training distribution. A higher humanness or nativeness score is one computed feature, not an immune-response assay.
For research comparison, freeze the source antibody and numbering convention. State which CDR definition is protected, which framework positions may change, whether both chains are modelled, and how many proposals each method may make. Count candidates before and after motif filters. Then test preservation of the intended properties with the relevant measurements. A small laboratory case study supports that case, not a universal success rate.
Developability: let each assay answer its own question
Developability compresses several properties into one convenient word: expression, conformational and colloidal stability, aggregation, viscosity, self-association, polyspecificity, chemical liabilities and more. These endpoints correlate imperfectly and depend on format, formulation and assay.
The 137-antibody study made that plurality visible by expressing variable domains in a common IgG1 context and running 12 biophysical assays (Jain et al.). TAP instead flags unusual CDR length, hydrophobic or charged surface patches, and heavy-light charge asymmetry relative to clinical-stage distributions (Raybould et al.). A flag marks an outlying proxy. It does not prove failure.
FLAb collects reported assay data across expression, thermostability, aggregation, polyreactivity, affinity and other properties, preserving datasets as separate records with metadata and units (Chungyoun, Ruffolo and Gray; official repository). Its authors reported that no evaluated model correlated consistently across all properties or across multiple datasets measuring similar properties. That is a developer-reported conclusion from the original benchmark, not a timeless ranking of later models or repository revisions.
Use one endpoint per primary model. Keep units, censoring, replicates, molecular format, campaign and assay conditions. Compare learned representations with simple motifs and physicochemical baselines. Calibrate within the intended campaign. Reserve “validated” for measured outcomes.
Availability is a vector, not an open badge
A paper, codebase, checkpoint, processed dataset, hosted service and licence are different objects. The rows below report observable release facts rather than awarding an “open” label.
| Representative family | Task and evidence boundary | Official artifacts and terms, checked 28 Aug 2026 |
|---|---|---|
| AntiBERTy | Unpaired-chain embeddings, attention and pseudo-log-likelihood; no native-pair guarantee | Code and packaged weights; MIT software licence; exact training corpus not republished |
| AbLang / AbLang2 | Single-chain representations in AbLang; paired representation in AbLang2; original AbLang evidence centres on sequence restoration | AbLang and AbLang2 code with public weights; BSD-3-Clause terms; processed pretraining corpora not packaged |
| IgBert / IgT5 | Paired and unpaired representations; transfer depends on task, head and split | Model cards and weights, plus a Zenodo archive; cards declare MIT; no processed training snapshot |
| Sapiens / BioPhi | Humanness scoring and mutation proposals, not immunogenicity | Sapiens and BioPhi code; VH weights and VL weights; repositories and cards declare MIT |
| AbNatiV2 | Nativeness scoring and humanisation-oriented research workflows | Maintained code and download instructions; CC BY-NC-SA 4.0 licence; public scoring server advertised |
| AntiBERTa2 | Sequence model and structure-aware variant; no complete pretraining release located | Model card and weights; custom terms restrict model use and generated antibody sequences to non-commercial use |
| IgFold | Antibody or nanobody coordinates, not affinity or docking | Runnable code and model loading; JHU Academic Software License requires separate commercial permission |
| ImmuneBuilder | Family-specific antibody, nanobody and TCR structure prediction | Code, weights and Colab links; BSD-3-Clause licence |
| AntiFold | Structure-conditioned probabilities and sequence samples; computational recovery is not binding proof | Code, checkpoint and webserver; BSD-3-Clause licence |
| S2ALM | Sequence-structure representation and developer-run computational tasks | Peer-reviewed paper and Apache-licensed repository, but no runnable S2ALM code or identifiable S2ALM checkpoint retrieved |
This is not a leaderboard or legal advice. A permissive code licence does not automatically cover externally hosted weights, training data, dependencies, services or outputs. Conversely, the absence of a special output clause does not guarantee freedom from other rights or obligations. AntiBERTa2 is unusually explicit about generated sequences. IgLM and IgFold are public but carry academic or non-commercial terms. IgGM’s repository explicitly says its code and model are MIT, while optional PyRosetta introduces separate terms (IgLM; IgGM). Audit the exact artifact you plan to run.
Hosted access is the weakest reproducibility category because a service can change without preserving a runnable version. Prefer versioned code and weights. Say “released for inference” when a pretraining rebuild is not supplied.
The benchmark can answer the wrong biological question
Suppose a dataset contains variants from one clonal lineage. A random row split can place near-identical relatives in training and test. The model may recognize a family rather than transfer to a new one. Antibody data offer more shortcut axes: shared germline genes, repeated donors, study-specific processing, common antigens, homologous structures and overlap with public pretraining corpora (OAS-explore; native-pairing study; IgFold).
Exact deduplication removes exact copies. It does not remove close relatives, a dominant donor or a shared experimental batch. An OAS-explore analysis reported that 13 individuals contributed more than 70% of human OAS sequences. Its authors trained 17 models on differently composed datasets and found limited transfer across chain types and between human and mouse repertoires, plus individual- and batch-specific effects (Glänzer, Reddy and Yermanos). More sequences did not erase provenance.
Figure 6. Original editorial synthesis of leakage routes and claim-shaped split policies. Sources: OAS documentation, OAS-explore, the native-pairing study, and IgFold.
Choose the split from the claim. For unseen variants within a known lineage, hold out variants and report each test variant’s sequence identity to its nearest training neighbours. For new lineages, cluster first and hold out clonotypes or lineage groups. For repertoire transfer, hold out donors and preferably studies or laboratories. For antigen generalisation, separate antigens or antigen families. For structure prediction, apply a temporal cutoff and homology filtering. For pretrained encoders, compare test sequences with the documented pretraining snapshot. If the snapshot is unavailable, report that contamination cannot be fully audited.
Random sequence splits can answer a narrow interpolation question. They are weak evidence for biological generalisation because they protect none of those boundaries by default.
Prospective experiments change the verdict
Retrospective metrics are useful for debugging and triage. A blinded prospective experiment asks a harder question: can a frozen procedure choose candidates before measurements are revealed?
The 2026 AIntibody challenge tested 511 designed or predicted antibodies from 29 organisations across affinity maturation, within-HCDR3-cluster ranking and out-of-library CDR design. Organisers reported that several groups produced strong candidates, but success was exceptional and did not transfer reliably across tasks. Except for one model, submissions ranking high-affinity clones within clustered HCDR3 sets performed worse than random clone picking; out-of-library design was highly variable (AIntibody challenge consortium). These are prospective, assay-anchored results for one challenge setup, not a universal family ranking.

Figure 7. Prospective evaluation commits methods to candidates before independent laboratories reveal affinity and developability measurements. Source: AIntibody challenge consortium, “A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability”, Fig. 1a-c. Licensed under CC BY 4.0. Cropped adaptation retaining panels a-c.
The lesson is not that computational antibody work is useless. Performance belongs to a task and funnel. Preserve the denominator at every stage: generated, valid, unique, synthesized, expressed, purified, monomeric, assayed, binding-positive and passing the named property threshold. Report negative controls and failures. A score is valuable when it enriches the measured endpoint relative to an honest baseline.
An executable research-use decision framework
Return to the opening request. Do not choose one model for all four verbs. Run the following gates.
Figure 8. Original editorial decision tree for research use. Its gates synthesize the representation controls in AntiBERTa and the native-pairing study, artifact requirements illustrated by S2ALM, split risks in OAS-explore, and assay boundaries in the AIntibody challenge.
- Define one output. Choose residue embeddings, paired-chain representations, mutation likelihoods, Fv coordinates, CDR proposals, humanisation edits, or one measured property. Stop if the output cannot be named precisely.
- Match the biological unit. Record unpaired VH or VL, native VH-VL pair, VHH, antibody alone, or antibody-antigen complex. Pin species, germline annotations, numbering and CDR definition.
- Shortlist runnable artifacts. Require the checkpoint, tokenizer, code revision, weight source and environment. Audit code, weights, data, service, dependencies and output terms separately. S2ALM does not pass a reproducible local-use gate as retrieved.
- Build three baselines. Use simple sequence or physicochemical features, a general protein model such as ESM-2 or ProtT5, and an antibody-specialised model. Hold the head, budget and split fixed when comparing representations.
- Shape the split around deployment. Cluster before splitting. Hold out lineage, donor, study, antigen family or structure date as required. Audit pretraining overlap and publish identity-to-training strata.
- Separate model stages. Record whether the encoder is frozen, probed, adapted or fully fine-tuned. For design, separate generator, inverse folder, structure oracle, filters and ranker. Do not credit the base representation with the whole stack’s result.
- Validate the endpoint. Use region-level error for a structure question, an antigen-specific assay for a binding question, and the named biophysical assay for a property question. Nativeness and likelihood scores are proxies. RMSD answers a structure-comparison question, and predicted-developability scores do not replace measured properties.
- Predeclare hard stops. Stop if the split cannot support the claimed transfer, artifacts cannot reproduce inference, terms block the intended research use, the model cannot beat simple baselines, or computational rank fails to enrich a bounded pilot. Diagnose the failed stage before scaling.
The answer to “humanise this antibody and keep its binding” is conditional but executable. Use a paired, numbering-aware workflow to propose edits. Keep a general-model and simple baseline. Protect CDR and structural constraints explicitly. Validate across lineage and study boundaries. Then measure binding and the named property assays. No single embedding or “foundation model” completes that chain.
Part 5 will move from rearranged immune repertoires to genomic sequence models, where strand, coordinate system and cell context redefine the biological unit. The useful habit carries over: specify what the model saw, what the downstream system predicts, and which held-out observation could prove the claim wrong.
References
- Olsen TH, Boyles F, Deane CM. “Observed Antibody Space: A diverse database of cleaned, annotated, and translated unpaired and paired antibody sequences.” Protein Science (2022).
- Leem J et al. “Deciphering the language of antibodies using self-supervised learning.” Patterns (2022).
- Ruffolo JA, Gray JJ, Sulam J. “Deciphering antibody affinity maturation with language models and weakly supervised learning.” arXiv:2112.07782 (2021).
- Olsen TH, Moal IH, Deane CM. “AbLang: an antibody language model for completing antibody sequences.” Bioinformatics Advances (2022).
- Burbach SM, Briney B. “Improving antibody language models with native pairing.” Patterns (2024).
- Jing H et al. “Accurate prediction of antibody function and structure using Bio-inspired Antibody Language Model.” Briefings in Bioinformatics (2024).
- Yin M et al. “S2ALM: Sequence-Structure Pre-trained Large Language Model for Comprehensive Antibody Representation Learning.” Research (2025).
- Ruffolo JA et al. “Fast, accurate antibody structure prediction from deep learning on massive set of natural antibodies.” Nature Communications (2023).
- Abanades B et al. “ImmuneBuilder: Deep-Learning models for predicting the structures of immune proteins.” Communications Biology (2023).
- Høie MH et al. “AntiFold: improved antibody structure-based design using inverse folding.” Bioinformatics Advances 5(1) (2025).
- Dreyer FA et al. “Inverse folding for antibody sequence design using deep learning.” ICML Computational Biology Workshop (2023).
- Prihoda D et al. “BioPhi: a platform for antibody design, humanization, and humanness evaluation based on natural antibody repertoires and deep learning.” mAbs (2022).
- Ramon V et al. “Assessing antibody and nanobody nativeness for hit selection and humanization with AbNatiV.” Nature Machine Intelligence (2024).
- Chungyoun M, Ruffolo JA, Gray JJ. “FLAb: Benchmarking deep learning methods for antibody fitness prediction.” bioRxiv (2024).
- Jain T et al. “Biophysical properties of the clinical-stage antibody landscape.” PNAS (2017).
- Raybould MIJ et al. “Five computational developability guidelines for therapeutic antibody profiling.” PNAS (2019).
- Glänzer WS, Reddy ST, Yermanos A. “Revealing bias in antibody language models through systematic training data processing with OAS-explore.” OpenReview (2025).
- AIntibody challenge consortium. “A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability.” Nature Biotechnology (2026).
- Lin Z et al. “Language models of protein sequences at the scale of evolution enable accurate structure prediction.” Science (2023).
- Elnaggar A et al. “ProtTrans: Toward cracking the language of life’s code through self-supervised deep learning and high performance computing.” IEEE TPAMI (2022).
- Dunbar J, Deane CM. “ANARCI: antigen receptor numbering and receptor classification.” Bioinformatics (2016).
Comments and feedback
Spotted an error or have a counterpoint? Comment below. No account needed, a name is enough. Corrections and pushback are welcome.