MIT framework targets more reliable design of novel proteins

Researchers in MIT’s Department of Biology have developed PottsMPNN, a machine-learning framework intended to improve computational protein design while reducing reliance on sequences already found in nature. The work, led by graduate student Foster Birnbaum and senior author Amy E. Keating, was published in PNAS under the title “Beyond native sequence recovery: Improved modeling of the sequence-energy landscape of protein structures.”
Protein design commonly begins with a target structure, followed by a model that proposes amino-acid sequences capable of folding into it. MIT’s researchers argue that recovering the particular sequence selected by evolution is not the most useful measure of success. Multiple amino-acid sequences can adopt the same protein structure, while a single sequence can take different structures depending on flexibility or a functional trigger.
Modeling feasible sequences rather than copying nature
PottsMPNN incorporates physical principles governing protein structure and stability. Its purpose is to improve both sequence generation and predictions of how mutations affect protein stability. The researchers describe this as stronger modeling of the sequence-energy landscape: the relationship between the amino acid at each position and a protein’s stability.
That distinction is important for proteins designed from scratch. A wholly new structure has no natural sequence against which a system can be compared. The relevant questions are instead whether a generated sequence is likely to fold into the intended structure, whether the model captures the energy landscape, and whether it can estimate the stability consequences of mutations.
Three changes to the training approach
The framework builds on research into the strategic use of training “noise,” meaning variations introduced into a protein structure during training. MIT says this reduces a model’s tendency to imitate native sequences too closely and broadens the range of structures for which it can generate sequences.
PottsMPNN also uses a pairwise distribution to represent interactions between amino acids. This allows the model to account for physical interactions among all 20 possible amino-acid choices at a pair of protein positions. The team identifies this capability as a key reason it models the sequence-energy landscape more accurately than other methods.
Finally, the researchers trained the framework with sets of evolutionarily related sequences, teaching it that distinct sequences may produce the same folded structure. Birnbaum noted the apparent tension: evolutionary information still draws on native sequences even as the method aims to move beyond them. Yet the results showed that reducing dependence on native sequences improved structural compatibility and energy prediction, including for novel proteins.
Implications for protein-design workflows
MIT positions PottsMPNN as an addition to protein-design pipelines where structural feasibility and mutation stability matter more than resemblance to a natural protein. Birnbaum suggested that further task-specific fine-tuning could improve predictions for particular mutation outcomes or consequences.
For organizations using computational protein design, the practical implication is to evaluate models against folding compatibility, sequence-energy understanding, and stability predictions, rather than treating native-sequence recovery as the primary benchmark for useful new-to-nature proteins.

