A new approach dubbed ProteinGuide aims to let researchers steer protein sequence generative models using auxiliary experimental or user-specified information—without the heavy burden of retraining the generative model itself. In a field where model updates typically require fresh computational learning, this “on-the-fly” strategy promises to make protein design faster and more adaptable to real-world constraints.
The core challenge is conditioning: most protein generators are pretrained to produce plausible sequences, but integrating additional signals—such as measured properties—usually means introducing extra training loops or specialized architectures. ProteinGuide instead offers a principled statistical framework that unifies how different generative paradigms can be guided at inference time.
Crucially, the method is compatible with a wide span of modern sequence generators. The authors demonstrate that ProteinGuide can work with masked language models such as ESM3, any-order autoregressive systems like ProteinMPNN, and diffusion or flow-matching models operating on discrete state spaces, including MultiFlow. This breadth suggests the approach is not tied to a single modeling philosophy, but rather to a common structure underlying conditioning.
As a proof of principle, the team uses pretrained generators to design proteins optimized for user-defined traits such as higher stability or activity. Rather than forcing the model to relearn protein–property relationships, ProteinGuide redirects sampling toward sequences expected to satisfy the desired objectives.
The work also tackles a familiar design dilemma: properties that conflict with each other. ProteinGuide can simultaneously optimize two target features, even when improving one tends to degrade the other—guiding the generator through a controlled balancing of objectives during sequence production.
To push beyond in silico success, the researchers pair ProteinGuide with wet-lab data generation. The target is an adenine base editor used in vivo, where editing performance is a practical bottleneck for genome engineering.
Rather than relying on many cycles of conventional optimization, ProteinGuide-supported design achieves a higher editing efficiency than had been reached previously after seven rounds of directed evolution. The result highlights the potential for guided generative sampling to reduce the experimental search space.
Overall, the study reframes protein engineering as a controllable sampling problem. By delivering inference-time conditioning across multiple model classes, ProteinGuide could become a versatile interface between pretrained generative intelligence and experimental reality—especially where retraining is costly or slow.
Subject of Research: Property guidance for protein sequence generative models
Article Title: Property guidance for protein sequence generative models with ProteinGuide
Article References: Xiong, J., Gaur, I., Lukarska, M. et al. Property guidance for protein sequence generative models with ProteinGuide. Nat Biotechnol (2026). https://doi.org/10.1038/s41587-026-03207-z
Image Credits: AI Generated
DOI: https://doi.org/10.1038/s41587-026-03207-z
Keywords: Protein engineering, generative models, on-the-fly conditioning, ESM3, ProteinMPNN, diffusion models, base editing, directed evolution

