A DNA Likelihood Is Not a Functional Assay: Genomic Foundation Models in 2026
A genomic model result is credible only when its tokenisation, strand rule, biological target, evaluation split, and artifact terms are explicit.
RNA models share an alphabet, not an estimand: choose them by molecular object, output, information inputs, evidence class, and artifact contract.
A genomic model result is credible only when its tokenisation, strand rule, biological target, evaluation split, and artifact terms are explicit.
Protein generators propose different biological objects. A defensible design starts with the assay, records the complete computational stack, and preserves every experimental denominator.
A practical guide to choosing, extracting, adapting and validating protein language-model representations without mistaking a system score for biological generalisation.
A task-first guide to antibody representations, structure prediction, CDR design, humanisation and developability, with the evidence and reproducibility checks that model scores leave out.
Suppose somebody hands you a FASTA file and asks for “the structure.” First decide whether the target is one chain, a protein assembly or a mixed complex containing ligands or nucleic acids.
The pitch for genomic foundation models is that one pretrained network now beats task-specific tools across the board, from regulatory annotation to clinical variant interpretation.
Sequence models, variant-effect prediction, and the work to read non-coding DNA.
Single-sequence folding, protein language models, and structure without alignment.
Contrastive screening, protein–ligand interaction, and design that survives the wet lab.
What models encode, where they fail, and how evaluation and oversight hold up.