Tianyu Lu, Po-Ssu Huang
A timely, well-informed perspective offering a coherent physics-grounding frame for protein generative modeling, but as an opinion piece with no original results its impact is bounded to citation/framing influence within one subfield.
Physical equations in protein modeling appear to have been replaced by generative models trained directly on structure data. By learning a mapping from noise to data, sampling de novo protein structures has become much more efficient. However, such models can also learn non-physical features and break down with out-of-distribution settings which are typical in protein design campaigns. In this review, we highlight where the physics persists in protein generative modeling pipelines and the issues that linger when physically relevant components of a macromolecular system are left unmodeled. We present a perspective that respecting the underlying physics of macromolecular systems, increasingly through learned representations that are physically grounded and updating generative models with experimental data, is foundational to generative modeling for functional protein design.
This is a review/perspective paper arguing that despite the apparent replacement of physics-based protein modeling (force fields, Rosetta energy functions) by data-driven generative models, physics remains foundational to protein generative modeling — and neglecting it leads to characteristic failure modes. The central conceptual framing is that generative models act as *physical priors* concentrating probability mass over physically plausible structures, and that protein design tasks inherently require sampling from out-of-distribution regions where physical grounding becomes essential. The authors organize the field around three mechanisms for "shaping the model's physical world view": energy-based losses (e.g., SLAE's energy-supervised latent space), architectural regularization (Potts-model parameterizations restricting to second-order terms), and experimental data feedback (design-build-test-learn / Bayesian optimization / lab-in-the-loop). The contribution is synthetic and framing-oriented rather than a new method or result.
As a perspective piece, there are no original experiments or proofs. The rigor lies in the marshalling of evidence for its thesis. The paper does this competently, citing concrete failure cases: adversarial charge-inversion mutations that fail to perturb co-folding predictions (Masters et al.), the meta-analysis of 3,766 binders showing ipSAE's low precision (Overath et al.), and length-generalization degradation as evidence that "length is a statistical property of the training set rather than a physical variable." This last argument is a genuinely sharp analytical insight. The paper is careful to distinguish deliberate vs. consequential omissions of physical components (hydrogens vs. activated water in serine hydrolase catalysis). However, the argument is largely assertive/illustrative rather than systematic — it does not, for example, quantify how much physics grounding improves generalization across a controlled set of models. The claim that "second-order terms are necessary and sufficient" is stated somewhat strongly given the cited support.
The paper is timely and addresses a real conceptual tension in a fast-moving, high-profile field (de novo protein design has produced Nobel-recognized work). Its framing — that OOD generalization failures in protein design trace to unmodeled physics — could productively guide method development, particularly the emphasis on physically grounded representations (SLAE, MLIPs) and experimental feedback loops. The synthesis of emerging high-throughput assays (mass-spec aggregation, HDX energy landscapes, automated NMR, bead display) as future training data sources is a useful forward-looking contribution that could orient experimentalists and modelers alike. That said, review/perspective papers in this niche typically influence framing and citations rather than directly enabling new capabilities. Much of the content promotes the authors' own recent work (SLAE, Caliby, SHAPES, Protpardelle, the metalloenzyme design), which somewhat narrows its neutral-synthesis value.
Highly timely. The paper engages with 2025-2026 developments (RFdiffusion3, Genie 3, Boltz-2, BindCraft, serine hydrolase design) and captures a live debate: whether co-folding models "learn physics" or merely memorize/interpolate. This is a genuine current bottleneck — the reliability of structure-prediction oracles as design filters is an actively contested problem, and the paper provides a coherent conceptual lens for it.
Strengths: (1) A clear, unifying conceptual frame (generative models as physical priors requiring OOD shift) that is intellectually coherent and pedagogically useful. (2) Several sharp analytical points, especially the "length is not a physical variable" argument and the interpretation of low-temperature sampling as averaging out unmodeled latent variables. (3) Excellent currency and breadth of citations spanning ML methods, structural biology, and emerging experimental assays. (4) The bridging of computational and experimental communities (new assays as future training signal) is valuable and non-obvious.
Limitations: (1) No original data, method, or theory — impact is bounded by what perspectives typically achieve. (2) Heavy reliance on the authors' own group's tools risks reading as advocacy rather than neutral survey. (3) Several key claims (necessary-and-sufficiency of second-order terms; the interpretation of ProteinMPNN's designability drop) are asserted with limited critical scrutiny. (4) The physics-vs-data framing, while clean, is somewhat philosophical and may not translate to concrete, actionable design prescriptions beyond "add energy supervision and experimental feedback." (5) The paper is a short opinion/review, so it lacks the systematic taxonomy or quantitative meta-analysis that would give it durable reference value.
The paper is well-written and accessible to the target audience, with clear figures conceptualizing the prior/plausible/objective distributions. Its two-figure structure is economical. Reproducibility, evidence-strength, and surprisingness dimensions are largely inapplicable given the paper type — there are no novel empirical results to reproduce or verify. The refutation angle (co-folding models don't learn physics) is inherited from cited works rather than originated here, so its refutation value is modest. Overall this is a solid, timely, well-informed perspective likely to be cited as a framing reference within the protein design subfield, but unlikely to independently change practice or serve as a foundational primitive.
Generated Sep 7, 2026
A timely, well-informed perspective offering a coherent physics-grounding frame for protein generative modeling, but as an opinion piece with no original results its impact is bounded to citation/framing influence within one subfield.