Pith. sign in

REVIEW 5 major objections 6 minor 38 references

The Dance of Atoms-De Novo Protein Design with Diffusion Model

T0 review · 5 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read In a head-to-head evaluation of eight diffusion-based protein backbone generators, RFdiffusion and Chroma show the most balanced performance across efficiency, structural plausibility, designability, novelty, naturalness, and diversity.

desk verdict A useful diffusion-model review whose headline ranking is compromised by an unvalidated ProteinMPNN/OmegaFold reconstruction for Cα-only models. read the letter →

arxiv 2504.16479 v1 pith:Q7Q2ENUF submitted 2025-04-23 q-bio.BM cs.AI

classification q-bio.BMcs.AI
keywords denovoproteindesigndiffusionmodelsbackbonegenerationgenerativeAIRFDiffusionChromastructureevaluationself-consistentTM-score
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review argues that diffusion models have become the most promising generative-AI route to de novo protein design, outpacing fragment-based and physics-guided methods in success rate and cost. To back that claim, it compares eight publicly available diffusion-based backbone generators under one protocol, measuring efficiency, structural plausibility, designability, novelty, naturalness, and diversity. Its central result is that RFdiffusion and Chroma are the most balanced choices overall, while FrameDiff leads on structural plausibility, RFdiffusion on designability, Chroma on diversity, and FoldingDiff on novelty and speed. The paper also collects experimentally validated successes for RFdiffusion, Chroma, and SCUBA-D, and it is candid that current models ignore conformational flexibility, protein–ligand interactions, and direct functional optimization.

What carries the argument

The load-bearing machinery is the diffusion process applied to a chosen representation of the protein backbone, plus the six-metric evaluation protocol that lets the models be compared. Backbone generators are grouped by representation: 2D feature maps (ProteinSGM, SCUBA-D), point clouds of $C_\alpha$ coordinates (ProtDiff, Genie), frame clouds in which each residue carries a translation and rotation in $\mathrm{SE}(3)$ (FrameDiff, RFdiffusion, Chroma, FoldingDiff), and latent-space vectors (PVQD). Frame-cloud diffusion is the representation behind the two models judged most balanced, because retaining inter-residue orientations lets the model capture chirality and residue interactions. The evaluation protocol defines designability through self-consistent TM-score: a generated backbone is reverse-folded into a sequence by ProteinMPNN or CarbonDesign, that sequence is refolded by OmegaFold, and the TM-score between generated and predicted structures measures whether the design is realizable; the same protocol supplies the dihedral-angle proxy that lets Cα-only models be scored for structural plausibility.

What would settle it

Re-running the sc-TM evaluation with AlphaFold2 in place of OmegaFold, or expressing the generated proteins and measuring which designs actually fold, would settle the ranking; if RFdiffusion and Chroma no longer lead, the balanced-performance claim is an artifact of the evaluation pipeline.

Watch

Extended reading notes

Core claim

The paper's central claim is that diffusion-based generative models now define the leading edge of de novo protein design, and that within this family no single model dominates: RFdiffusion and Chroma show the most balanced performance across all six evaluation axes, with each other model winning a specific niche. FrameDiff generates backbones whose dihedral-angle statistics most closely match native proteins; RFdiffusion achieves the highest self-consistent TM-scores across all tested lengths; Chroma produces the most structurally diverse populations while retaining natural secondary-structure composition; and FoldingDiff is the fastest and most novel generator for short chains. The authors further claim that Chroma completes all six conditional design tasks in their table—binder design, symmetric oligomers, secondary-structure-constrained design, category-based design, structure refinement, and partial-sequence-constrained design—whereas RFdiffusion excels specifically at spatially constrained tasks. The review supports these claims with a shared evaluation protocol and with experimental case studies in which designed binders, soluble proteins, and heme-binding proteins were verified in the lab.

Load-bearing premise

The load-bearing premise is that the computational round-trip—designing an amino acid sequence for a generated backbone, predicting its folded structure, and taking the best agreement score—reveals which generated backbones can genuinely fold into their intended shapes; if that round-trip is unfaithful, the rankings measure the pipeline rather than the generators.

Editorial extensions

If this is right

  • For a general-purpose de novo backbone design task, the default should be RFdiffusion or Chroma, since the benchmark treats them as the most balanced across all six criteria.
  • When a design goal prioritizes native-like backbone geometry, FrameDiff is the indicated generator; when it prioritizes designability, RFdiffusion; when diversity of conformations matters, Chroma.
  • For quick exploration of short, structurally novel proteins, FoldingDiff offers the best efficiency–novelty trade-off among the tested models.
  • Conditional design—especially binder and symmetric-oligomer construction—is currently best served by Chroma, with RFdiffusion as the strong alternative for spatially constrained tasks.
  • Because Cα-only generators (ProtDiff, Genie) must be evaluated through a reverse-folding/refolding proxy, their apparent efficiency comes with an added layer of uncertainty that full-atom or frame-based models do not carry.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The balanced-performance ranking may be specific to unconditional generation; in conditional tasks, the paper's own table shows Chroma dominating, so a practitioner optimizing binder design might reasonably weight RFdiffusion's validated binder successes more heavily than the six-metric average.
  • Because the six-metric protocol evaluates structure-generation models only, a direct comparison that adds sequence-diffusion models (EvoDiff, TaxDiff) and co-generation models under the same metrics could show whether the backbone-first advantage is real or an artifact of benchmark scope.
  • A cheap falsification check would be to re-run the benchmark with AlphaFold2 in place of OmegaFold in the sc-TM pipeline; if the relative rankings of FrameDiff, RFdiffusion, and Chroma shift materially, the paper's conclusions describe that particular pipeline.
  • The review's emphasis on experimental validation suggests a concrete next step: express a matched set of designs from several generators in the same host and compare soluble expression and binding success rates, which would test whether the in silico 'balance' translates to the lab.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This manuscript combines a review of diffusion-based de novo protein design with a new empirical benchmark. The authors compare eight backbone-generation models (Genie, Chroma, ProtDiff, ProteinSGM, FoldingDiff, FrameDiff, SCUBA-D, and RFdiffusion) across six metrics: efficiency, structural plausibility, designability, novelty, naturalness, and diversity, generated at six residue lengths on an A100 GPU. The central conclusion, stated in Sections 4 and 6, is that RFdiffusion and Chroma exhibit the most balanced performance, with FrameDiff best on structural plausibility, RFdiffusion best on designability, Chroma best on diversity, and FoldingDiff most novel and efficient. The paper also reviews conditional-generation capabilities and summarizes experimentally validated successes of RFdiffusion, Chroma, and SCUBA-D. The review portion is useful for orientation, but the benchmark portion, which underpins the headline ranking, has several methodological gaps that need to be addressed before the comparison can be considered reliable.

Significance. If the evaluation were rigorous, the paper would provide a valuable practical resource for choosing among generative models for protein backbone design, a question of current methodological interest. The taxonomy of diffusion formulations (feature map, point cloud, frame cloud, latent space) and the compilation of experimental success stories are informative. The authors make the effort to generate all structures on the same hardware and to use common in-silico proxies, which is commendable. However, the central ranking, and especially the 'most balanced performance' claim, currently rests on an undocumented reconstruction pipeline for Cα-only models, a post-hoc designability filter of unquantified effect, and comparisons without error bars or significance tests. These issues are load-bearing for the paper's main conclusion; the review content alone would not justify the current framing.

major comments (5)
  1. [Section 4.2] The evaluation of Cα-only models (FoldingDiff, Genie, ProtDiff) is not performed on the raw generated structures but on a ProteinMPNN/OmegaFold round-trip reconstruction with best-of-N selection. The text states that ProteinMPNN sequences are designed from the structure and OmegaFold predicts a structure for that sequence, and 'finally select the highest-scoring prediction as our evaluation targets.' No procedure is given for reconstructing the N, C, and O backbone atoms that ProteinMPNN requires from Cα-only coordinates, and no validation of this reconstruction is provided. Consequently, all metrics reported for these three models (dihedral angles, designability, naturalness, novelty, diversity) reflect properties of an undocumented reconstruction pipeline rather than of the generators themselves. This is internally inconsistent with the statement in Section 4.1 that 'All subsequent evaluations are based on the structures generated in this stage,' and it makes the cross-model comparison against full-atom models unfair. The ranking of RFdiffusion and Chroma as 'most balanced' directly depends on this comparison and therefore cannot be assessed from the current evidence.
  2. [Section 4.2] The 5th-percentile sc-TM filter applied to generated structures is a post-hoc selection step whose effect on the downstream novelty, naturalness, and diversity analyses is not characterized. Because the filter removes the least designable structures and the retained fraction likely differs across models (e.g., RFdiffusion with high sc-TM retains more structures than ProtDiff), the subsequent metrics are conditional on designability. A model with low raw designability could appear improved in naturalness or diversity simply because a more selected subset is analyzed. The paper should report the number and fraction of structures that pass the threshold for each model and provide results both before and after filtering, or use a statistically principled correction for selection effects.
  3. [Sections 4.1 and 4.4] No error bars, confidence intervals, or significance tests are reported for any of the quantitative comparisons. Efficiency times, dihedral-angle MSE values, KL divergences, and cluster counts are presented as point estimates, but diffusion sampling is stochastic and only 100 structures per length were generated. The differences that underlie the headline ranking (e.g., 'RFDiffusion maintains the highest structural consistency across all lengths' or 'Genie-generated structures exhibit the highest similarity to native proteins') may be within sampling variability. At minimum, bootstrap or replicated-run standard errors should be provided, and a statistical test should accompany claims of model superiority on each metric.
  4. [Section 4.3 and Table 3] The novelty metric is computed against each model's own training set, but the training sets differ substantially in both size and composition (PDB versus CATH 4.3.0 versus SCOPe). A model trained on CATH 4.3.0 (e.g., FoldingDiff) will appear more novel if measured against its smaller reference set than a PDB-trained model, not necessarily because it generates more original structures. This confound makes the stated ranking on novelty across models uninterpretable. The authors should either use a common reference database for all models, or explicitly analyze and discuss how the choice of reference set affects the novelty comparison. The same issue may affect structural-plausibility and naturalness comparisons if the native reference set is defined by sequence-length bins only.
  5. [Sections 4.1 and 4.4] The handling of residue-length limitations is not transparent. FoldingDiff and ProteinSGM can generate only up to 128 residues, and Genie up to 256 residues, yet the paper reports aggregate metrics across lengths 50-500. It is unclear whether models with length caps are evaluated only on the lengths they can generate, whether the pooled distributions are computed with unequal sample sizes per length, and how the summary statements such as 'for longer protein chains, Chroma, RFDiffusion, and SCUBA-D generate more diverse structures' incorporate the fact that some models have no data at those lengths. The paper should clearly state per-length sample counts and restrict cross-model comparisons to length ranges where all evaluated models can generate.
minor comments (6)
  1. [Section 2] The heading 'The Fundation of Diffusion Models' contains a typo; it should read 'The Foundation of Diffusion Models.'
  2. [Equations (1)-(4)] The notation in Equation (1) and the surrounding text is garbled (e.g., the definition of s(t) is not properly typeset), and the symbols in Equations (2)-(4) are not fully defined in the main text. Please provide clear definitions for all variables, including the Wiener process, the covariance matrix R, and the wrapped-normal parameters.
  3. [Table 2 and Section 4.2] There are several typos in Table 2, including 'Molde' instead of 'Model' and 'SCUB-D' instead of 'SCUBA-D.' Model names are also inconsistent between 'ProtDiff' and 'ProDiff' in Section 3.1.2; please standardize.
  4. [Figure 3 and Section 4.2] The text refers to 'Figure 3b' for both the dihedral-angle MSE and the designability across lengths, which appears to be a figure-numbering error. Please renumber the panels so that each reference is unambiguous.
  5. [Table 4] The conditional-generation table uses checkmarks without stating the evaluation procedure. It is unclear whether these were determined from documentation, from the authors' experiments, or from published results. Since conditional capabilities are part of the model comparison, please specify the source and criteria for each checkmark.
  6. [General] The effective sample sizes after the sc-TM filter are not reported. Please state, for each model and length, how many of the 100 generated structures survived the 5th-percentile threshold and were used for the novelty, naturalness, and diversity analyses.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: this is an empirical review with an openly stated evaluation proxy, not a derivation that reduces to its inputs.

full rationale

No circular step is present. The paper is a review and comparative benchmark of diffusion-based protein design models; its headline conclusion (RFdiffusion and Chroma are the most balanced) is an empirical ranking supported by Section 4's evaluation, not by an equation fitted to data and then relabeled as a prediction. Section 4.2 explicitly acknowledges that for ProtDiff, Genie, and FoldingDiff, 'directly obtaining dihedral angle distributions of these models’ outputs is not feasible,' so the authors 'utilize ProteinMPNN to predict the amino acid sequence of the structure, then employ Omegafold to predict the corresponding structure for that sequence and calculate the TM-score between the two structures, and finally select the highest-scoring prediction as our evaluation targets.' This is an openly stated proxy, and it is not definitionally circular because ProteinMPNN and OmegaFold are external, pretrained models rather than parameters fitted from the generated structures under evaluation. The manuscript omits the mechanical detail of how N, C, and O atoms are reconstructed from Cα-only inputs before ProteinMPNN, and the best-of-N TM-score selection can bias the proxy; these are validity concerns, not circularity. One self-citation exists: reference [33] (CarbonDesign) shares co-author Changyong Yu with the present paper, and CarbonDesign is used alongside ProteinMPNN in the designability assessment; however, CarbonDesign is an independently peer-reviewed external tool and the conclusion does not rest on its citation alone, so this does not constitute load-bearing circularity. The paper is self-contained against external benchmarks, and its rankings are falsifiable by re-running the stated protocol with different sequence-design or structure-prediction back ends.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new model or theory, so it carries no invented entities. Its quantitative conclusions depend on several hand-chosen thresholds and domain assumptions about evaluation proxies, especially the ProteinMPNN/OmegaFold pipeline for Cα-only models and the sc-TM filtering step.

free parameters (5)
  • sc-TM filtering threshold = 5th percentile of native sc-TM distribution
    The authors filter generated structures before novelty, naturalness, and diversity evaluation, so this cutoff shapes every downstream ranking and is not derived from theory.
  • Foldseek clustering coverage threshold = 0.9
    This threshold defines whether two structures are in the same Foldseek cluster in the diversity evaluation (Section 4.5).
  • Hierarchical clustering TM-score threshold = 0.5
    Used to define clusters in TM-score-based hierarchical clustering, so it directly controls the diversity counts (Section 4.5).
  • Hierarchical clustering RMSD threshold = 2
    Used to define clusters in RMSD-based hierarchical clustering, affecting diversity counts (Section 4.5).
  • Residue length bins = 50, 100, 150, 200, 300, 500
    These hand-chosen lengths, and the slightly different native reference bins (25-75 up to 475-525), determine comparability of the per-length analyses.
assumptions (4)
  • domain assumption Self-consistency TM-score computed with ProteinMPNN and OmegaFold is a valid proxy for designability.
    Section 4.2 uses sc-TM to rank models and to filter structures before other metrics, which assumes the inverse folding and structure prediction models are accurate enough for this purpose.
  • domain assumption Native PDB chains matched by length form an appropriate reference for structural plausibility and naturalness.
    Section 4.2 and 4.4 compare generated dihedral angle and secondary-structure distributions against native PDB chains, assuming the native set is the correct baseline.
  • domain assumption The training sets listed in Table 3 are complete and correctly identified for each model.
    Section 4.3 novelty scores depend on comparing generated structures to each model's actual training set; an incomplete or wrong training set would bias novelty rankings.
  • domain assumption Generated structures that pass the sc-TM filter are representative of each model's output.
    The designability filter changes the evaluated population, and the paper does not report the number or fraction of structures removed per model, so cross-model bias cannot be ruled out.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Dance of Atoms-De Novo Protein Design with Diffusion Model." pith.science (2026). https://pith.science/paper/Q7Q2ENUF

@misc{pith2026250416479,
  author       = {Pith},
  title        = {Pith review of: The Dance of Atoms-De Novo Protein Design with Diffusion Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q7Q2ENUF}},
  note         = {Machine review of arXiv:2504.16479}
}
read the original abstract

The de novo design of proteins refers to creating proteins with specific structures and functions that do not naturally exist. In recent years, the accumulation of high-quality protein structure and sequence data and technological advancements have paved the way for the successful application of generative artificial intelligence (AI) models in protein design. These models have surpassed traditional approaches that rely on fragments and bioinformatics. They have significantly enhanced the success rate of de novo protein design, and reduced experimental costs, leading to breakthroughs in the field. Among various generative AI models, diffusion models have yielded the most promising results in protein design. In the past two to three years, more than ten protein design models based on diffusion models have emerged. Among them, the representative model, RFDiffusion, has demonstrated success rates in 25 protein design tasks that far exceed those of traditional methods, and other AI-based approaches like RFjoint and hallucination. This review will systematically examine the application of diffusion models in generating protein backbones and sequences. We will explore the strengths and limitations of different models, summarize successful cases of protein design using diffusion models, and discuss future development directions.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

38 extracted references · 20 canonical work pages

  1. [1]

    Korendovych, I. V. & DeGrado, W. F. De novo protein design, a retrospective. Quart. Rev. Biophys. 53, e3 (2020)

  2. [2]

    De novo design of beta-sheet proteins

    Hecht, M. De novo design of beta-sheet proteins. Proceedings of the National Academy of Sciences of the United States of America 91, 8729–30 (1994)

  3. [3]

    UniProt: the Universal Protein knowledgebase

    Apweiler, R. UniProt: the Universal Protein knowledgebase. Nucleic Acids Research 32, 115D – 119 (2004)

  4. [4]

    Berman, H. M. The Protein Data Bank. Nucleic Acids Research 28, 235–242 (2000)

  5. [5]

    K., Brenner, S

    Fox, N. K., Brenner, S. E. & Chandonia, J.-M. SCOPe: Structural Classification of Proteins—extended, integrating SCOP and ASTRAL data and classification of new structures. Nucl. Acids Res. 42, D304–D309 (2014)

  6. [6]

    Repecka, D. et al. Expanding functional protein sequence spaces using generative adversarial networks. Nat Mach Intell 3, 324–333 (2021)

  7. [7]

    R., Choe, C

    Eguchi, R. R., Choe, C. A. & Huang, P.-S. Ig-VAE: Generative modeling of protein structure by direct 3D coordinate generation. PLoS Comput Biol 18, e1010271 (2022)

  8. [8]

    Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100 (2023)

Show all 38 references
  1. [9]

    Guo, Z. et al. Diffusion Models in Bioinformatics: A New Wave of Deep Learning Revolution in Action

  2. [10]

    & Ermon, S

    Song, Y., Garg, S., Shi, J. & Ermon, S. Sliced Score Matching: A Scalable Approach to Density and Score Estimation. Preprint at http://arxiv.org/abs/1905.07088 (2019)

  3. [11]

    & Abbeel, P

    Ho, J., Jain, A. & Abbeel, P. Denoising Diffusion Probabilistic Models. Preprint at http://arxiv.org/abs/2006.11239 (2020)

  4. [12]

    & Ermon, S

    Song, J., Meng, C. & Ermon, S. Denoising Diffusion Implicit Models. Preprint at http://arxiv.org/abs/2010.02502 (2022)

  5. [13]

    Song, Y. et al. Score-Based Generative Modeling through Stochastic Differential Equations. Preprint at http://arxiv.org/abs/2011.13456 (2021)

  6. [14]

    & Schneider, G

    Atz, K., Grisoni, F. & Schneider, G. Geometric deep learning on molecular representations. Nat Mach Intell 3, 1023–1032 (2021)

  7. [15]

    S., Kim, J

    Lee, J. S., Kim, J. & Kim, P. M. ProteinSGM: Score-Based Generative Modeling for de Novo Protein Design. http://biorxiv.org/lookup/doi/10.1101/2022.07.13.499967 (2022) doi:10.1101/2022.07.13.499967

  8. [16]

    Liu, Y. et al. De novo protein design with a denoising diffusion network independent of pretrained structure prediction models. Nat Methods (2024) doi:10.1038/s41592-024-02437-w

  9. [17]

    Wu, K. E. et al. Protein structure generation via folding diffusion. Nat Commun 15, 1059 (2024)

  10. [18]

    Trippe, B. L. et al. Diffusion probabilistic modeling of protein backbones in 3D for the motif-scaffolding problem. Preprint at http://arxiv.org/abs/2206.04119 (2023)

  11. [19]

    & AlQuraishi, M

    Lin, Y., Lee, M., Zhang, Z. & AlQuraishi, M. Out of Many, One: Designing and Scaffolding Proteins at the Scale of the Structural Universe with Genie 2. Preprint at http://arxiv.org/abs/2405.15489 (2024)

  12. [20]

    & McIlvin, M

    Saito, M. & McIlvin, M. R. Detection of Iron Protein Supercomplexes in Pseudomonas aeruginosa by Native Metalloproteomics. Preprint at https://doi.org/10.1101/2025.01.15.633287 (2025)

  13. [21]

    Yim, J. et al. SE(3) diffusion model with application to protein backbone generation. Preprint at http://arxiv.org/abs/2302.02277 (2023)

  14. [22]

    Ingraham, J. B. et al. Illuminating protein space with a programmable generative model. Nature 623, 1070–1078 (2023)

  15. [23]

    P., Salimans, T., Poole, B

    Kingma, D. P., Salimans, T., Poole, B. & Ho, J. Variational Diffusion Models. Preprint at http://arxiv.org/abs/2107.00630 (2023)

  16. [24]

    & Dhariwal, P

    Nichol, A. & Dhariwal, P. Improved Denoising Diffusion Probabilistic Models. Preprint at http://arxiv.org/abs/2102.09672 (2021)

  17. [25]

    & Liu, H

    Liu, Y., Chen, L. & Liu, H. Diffusion in a Quantized Vector Space Generates Non-Idealized Protein Structures and Predicts Conformational Distributions. http://biorxiv.org/lookup/doi/10.1101/2023.11.18.567666 (2023) doi:10.1101/2023.11.18.567666

  18. [26]

    Ni, B., Kaplan, D. L. & Buehler, M. J. Generative design of de novo proteins based on secondary-structure constraints using an attention-based diffusion model. Chem 9, 1828–1849 (2023)

  19. [27]

    J., Yim, J., Barzilay, R

    Yang, J. J., Yim, J., Barzilay, R. & Jaakkola, T. Fast non-autoregressive inverse folding with discrete diffusion. Preprint at http://arxiv.org/abs/2312.02447 (2023)

  20. [28]

    Alamdari, S. et al. Protein Generation with Evolutionary Diffusion: Sequence Is All You Need. http://biorxiv.org/lookup/doi/10.1101/2023.09.11.556673 (2023) doi:10.1101/2023.09.11.556673

  21. [29]

    Zongying, L. et al. TaxDiff: Taxonomic-Guided Diffusion Model for Protein Sequence Generation. Preprint at http://arxiv.org/abs/2402.17156 (2024)

  22. [30]

    Meshchaninov, V. et al. Diffusion on language model encodings for protein sequence generation. Preprint at https://doi.org/10.48550/arXiv.2403.03726 (2025)

  23. [31]

    & Achim, T

    Anand, N. & Achim, T. Protein Structure and Sequence Generation with Equivariant Denoising Diffusion Probabilistic Models. Preprint at http://arxiv.org/abs/2205.15019 (2022)

  24. [32]

    Dauparas, J. et al. Robust deep learning–based protein sequence design using ProteinMPNN. Science 378, 49–56 (2022)

  25. [33]

    & Zhang, H

    Ren, M., Yu, C., Bu, D. & Zhang, H. Accurate and robust protein sequence design with CarbonDesign. Nat Mach Intell 6, 536–547 (2024)

  26. [34]

    Wu, R. et al. High-resolution de novo structure prediction from primary sequence. Preprint at https://doi.org/10.1101/2022.07.21.500999 (2022)

  27. [35]

    Barrio-Hernandez, I. et al. Clustering predicted structures at the scale of the known protein universe. Nature 622, 637–645 (2023)

  28. [36]

    & Samudrala, R

    Hung, L.-H. & Samudrala, R. Accelerated protein structure comparison using TM-score-GPU. Bioinformatics 28, 2191–2192 (2012)

  29. [37]

    Bennett, N. R. et al. Atomically accurate de novo design of single-domain antibodies. Preprint at https://doi.org/10.1101/2024.03.14.585103 (2024)

  30. [38]

    Vázquez Torres, S. et al. De novo design of high-affinity binders of bioactive helical peptides. Nature 626, 435–442 (2024). Supplement for The Dance of Atoms: De Novo Protein Design with Diffusion Model Yujie Qin1,2,Ming He1,3, Changyong Yu3, Ming Ni1,*, Xian Liu1,*, Xiaochen...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.