Aligning protein AI with experimental fitness stabilizes influenza vaccines
From Pepkio Team · 17 August 2026 · 2 min read
Scientists report today in Nature Methods a new way to make protein-generating AI models follow experimental fitness data. The work, led by Brian L. Hie at Stanford University, with first author Talal Widatalla, introduces ProteinDPO, a version of the structure-conditioned protein language model ESM-IF1 aligned using direct preference optimization (DPO) — an algorithm originally developed to align text-generating AI with human preferences. Instead of human feedback, the model was trained on about 660,000 experimentally measured stability changes across 403 protein domains.
The alignment closed a key 'alignment gap': unsupervised models often fail at specific tasks like stability prediction despite broad protein knowledge. ProteinDPO outperformed both vanilla ESM-IF1 and a supervised fine-tuned version, achieving stability prediction competitive with specialized supervised models such as ThermoMPNN. It also generalized beyond stability — improving binding-affinity scoring of protein complexes and ranking thermal stability of multichain antibodies, tasks outside its training distribution.
To show practical value, the team used ProteinDPO to stabilize the prefusion form of H5N1 influenza hemagglutinin (HA), a key vaccine target. Testing only 45 designs, they found about 80% had increased or similar stability; the best variants improved melting temperature by up to 17°C. Mutations chosen from a 2004 strain also stabilized 2024 strains by up to 32°C, while preserving binding to broadly neutralizing antibodies. The model recovered known stabilizing mutations and identified new ones without any prior training on HA.
Caveats remain: stability improvements were measured in vitro, not as vaccine efficacy, and some scoring relied on computational predictions. Still, the framework is general and open-source, pointing toward a future where biological foundation models can be aligned to almost any measurable fitness.
Reference: Widatalla, T., Borah, A.A., King, S.H. et al. Aligning protein-generative models to experimental fitness with ProteinDPO. Nature Methods (2026). https://doi.org/10.1038/s41592-026-03137-3
Explore Pepkio
- Bioinformatics CRO
Reproducible, publication-style analyses with full source code and methods for academic labs and biotech teams.
- The bioinformatics outsourcing playbook
Cost, timelines, vendor selection, and reproducibility for labs weighing whether to outsource bioinformatics.
- Free AI-assisted lab tools
Browser calculators for serial dilutions, molarity, PCR setup, plate readers, and more — no account required.