{"id":"23f078ed-7100-4ba6-84b9-178ec1425ac0","arxiv_id":"2504.16479","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A review and comparative benchmark of eight diffusion-based protein design models reports RFdiffusion and Chroma as the most balanced across six evaluation criteria.","lead":"This paper reviews AI diffusion models for designing new proteins and compares eight of them on speed, structure quality, designability, novelty, naturalness, and diversity. It reports that RFdiffusion and Chroma are the most balanced tools overall, while other models lead in specific metrics.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Rankings of Cα-only models rest on an undocumented ProteinMPNN/OmegaFold reconstruction; the 'balanced' conclusion may be an artifact of that proxy.","rationale":"The reader's weakest_assumption is the same as mine: the comparison of Cα-only generators through a ProteinMPNN/OmegaFold reconstruction pipeline is the load-bearing point. The paper's strongest claim is explicitly comparative ('RFdiffusion and Chroma exhibit the most balanced performance'), so the evaluation must be measurement-neutral across models. For full-atom models, metrics like dihedral plausibility are computed on generated structures; for Cα-only models they are computed on reconstructed, best-of-N predicted structures. That asymmetry is not validated anywhere in Section 4 or the supplement. The sc-TM 5th-percentile filter further couples the designability proxy to model ranking, and the paper does not report how many structures are discarded per model. The absence of code, seeds, and error bars compounds the issue, but the unvalidated reconstruction is the sharpest concrete defect. The proposed experiment—degrading full-atom outputs to Cα and running the same proxy—directly tests whether the proxy is measurement-neutral. If it is not, the central 'most balanced' claim is not supported. The reader's conditional verdict therefore remains appropriate, with the condition being validation of the Cα reconstruction pipeline or a redesign of the comparison.","tokens_in":12866,"tokens_out":4121,"duration_ms":41391,"concrete_test":"Re-run the benchmark with a controlled pipeline check: take full-atom designs from RFdiffusion and Chroma, discard all atoms except Cα, rebuild missing backbone atoms with the same protocol used for FoldingDiff/Genie/ProtDiff, then run ProteinMPNN, OmegaFold, and select the highest sc-TM prediction. Recompute sc-TM and dihedral MSE on these proxy structures and compare with scores computed directly on the original full-atom structures. If the proxy causes a systematic shift or a rank change among models, the Cα-only model rankings are not trustworthy. Also report the fraction of structures removed by the 5th-percentile sc-TM filter for each model, since that filter can mask un-designable generations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 states that for FoldingDiff, Genie, and ProtDiff, which output only Cα coordinates, the authors use ProteinMPNN to design sequences, OmegaFold to predict structures, and select the highest-TM prediction as evaluation targets. Three problems follow. First, no procedure is given for reconstructing the N/C/O backbone atoms that ProteinMPNN requires from a Cα-only input. Second, structural plausibility (dihedral angles), naturalness, novelty, and diversity metrics are then computed on these best-of-N round-trip reconstructions rather than on the generated backbones, so Cα-only models are compared through a proxy that full-atom models are not subject to. Third, the highest-TM selection injects an optimization step that can make a Cα-only generator look more native-like than its raw output. Because the headline claim—that RFdiffusion and Chroma are the most balanced—depends on a fair cross-model comparison, an unvalidated reconstruction pipeline is the weakest load-bearing assumption. If the proxy favors or penalizes Cα-only models, the ranking and the balanced conclusion change.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript combines a review of diffusion-based de novo protein design with a new empirical benchmark. The authors compare eight backbone-generation models (Genie, Chroma, ProtDiff, ProteinSGM, FoldingDiff, FrameDiff, SCUBA-D, and RFdiffusion) across six metrics: efficiency, structural plausibility, designability, novelty, naturalness, and diversity, generated at six residue lengths on an A100 GPU. The central conclusion, stated in Sections 4 and 6, is that RFdiffusion and Chroma exhibit the most balanced performance, with FrameDiff best on structural plausibility, RFdiffusion best on designability, Chroma best on diversity, and FoldingDiff most novel and efficient. The paper also reviews conditional-generation capabilities and summarizes experimentally validated successes of RFdiffusion, Chroma, and SCUBA-D. The review portion is useful for orientation, but the benchmark portion, which underpins the headline ranking, has several methodological gaps that need to be addressed before the comparison can be considered reliable.","tokens_in":13037,"tokens_out":6006,"duration_ms":59148,"significance":"If the evaluation were rigorous, the paper would provide a valuable practical resource for choosing among generative models for protein backbone design, a question of current methodological interest. The taxonomy of diffusion formulations (feature map, point cloud, frame cloud, latent space) and the compilation of experimental success stories are informative. The authors make the effort to generate all structures on the same hardware and to use common in-silico proxies, which is commendable. However, the central ranking, and especially the 'most balanced performance' claim, currently rests on an undocumented reconstruction pipeline for Cα-only models, a post-hoc designability filter of unquantified effect, and comparisons without error bars or significance tests. These issues are load-bearing for the paper's main conclusion; the review content alone would not justify the current framing.","major_comments":[{"comment":"The evaluation of Cα-only models (FoldingDiff, Genie, ProtDiff) is not performed on the raw generated structures but on a ProteinMPNN/OmegaFold round-trip reconstruction with best-of-N selection. The text states that ProteinMPNN sequences are designed from the structure and OmegaFold predicts a structure for that sequence, and 'finally select the highest-scoring prediction as our evaluation targets.' No procedure is given for reconstructing the N, C, and O backbone atoms that ProteinMPNN requires from Cα-only coordinates, and no validation of this reconstruction is provided. Consequently, all metrics reported for these three models (dihedral angles, designability, naturalness, novelty, diversity) reflect properties of an undocumented reconstruction pipeline rather than of the generators themselves. This is internally inconsistent with the statement in Section 4.1 that 'All subsequent evaluations are based on the structures generated in this stage,' and it makes the cross-model comparison against full-atom models unfair. The ranking of RFdiffusion and Chroma as 'most balanced' directly depends on this comparison and therefore cannot be assessed from the current evidence.","section":"Section 4.2"},{"comment":"The 5th-percentile sc-TM filter applied to generated structures is a post-hoc selection step whose effect on the downstream novelty, naturalness, and diversity analyses is not characterized. Because the filter removes the least designable structures and the retained fraction likely differs across models (e.g., RFdiffusion with high sc-TM retains more structures than ProtDiff), the subsequent metrics are conditional on designability. A model with low raw designability could appear improved in naturalness or diversity simply because a more selected subset is analyzed. The paper should report the number and fraction of structures that pass the threshold for each model and provide results both before and after filtering, or use a statistically principled correction for selection effects.","section":"Section 4.2"},{"comment":"No error bars, confidence intervals, or significance tests are reported for any of the quantitative comparisons. Efficiency times, dihedral-angle MSE values, KL divergences, and cluster counts are presented as point estimates, but diffusion sampling is stochastic and only 100 structures per length were generated. The differences that underlie the headline ranking (e.g., 'RFDiffusion maintains the highest structural consistency across all lengths' or 'Genie-generated structures exhibit the highest similarity to native proteins') may be within sampling variability. At minimum, bootstrap or replicated-run standard errors should be provided, and a statistical test should accompany claims of model superiority on each metric.","section":"Sections 4.1 and 4.4"},{"comment":"The novelty metric is computed against each model's own training set, but the training sets differ substantially in both size and composition (PDB versus CATH 4.3.0 versus SCOPe). A model trained on CATH 4.3.0 (e.g., FoldingDiff) will appear more novel if measured against its smaller reference set than a PDB-trained model, not necessarily because it generates more original structures. This confound makes the stated ranking on novelty across models uninterpretable. The authors should either use a common reference database for all models, or explicitly analyze and discuss how the choice of reference set affects the novelty comparison. The same issue may affect structural-plausibility and naturalness comparisons if the native reference set is defined by sequence-length bins only.","section":"Section 4.3 and Table 3"},{"comment":"The handling of residue-length limitations is not transparent. FoldingDiff and ProteinSGM can generate only up to 128 residues, and Genie up to 256 residues, yet the paper reports aggregate metrics across lengths 50-500. It is unclear whether models with length caps are evaluated only on the lengths they can generate, whether the pooled distributions are computed with unequal sample sizes per length, and how the summary statements such as 'for longer protein chains, Chroma, RFDiffusion, and SCUBA-D generate more diverse structures' incorporate the fact that some models have no data at those lengths. The paper should clearly state per-length sample counts and restrict cross-model comparisons to length ranges where all evaluated models can generate.","section":"Sections 4.1 and 4.4"}],"minor_comments":[{"comment":"The heading 'The Fundation of Diffusion Models' contains a typo; it should read 'The Foundation of Diffusion Models.'","section":"Section 2"},{"comment":"The notation in Equation (1) and the surrounding text is garbled (e.g., the definition of s(t) is not properly typeset), and the symbols in Equations (2)-(4) are not fully defined in the main text. Please provide clear definitions for all variables, including the Wiener process, the covariance matrix R, and the wrapped-normal parameters.","section":"Equations (1)-(4)"},{"comment":"There are several typos in Table 2, including 'Molde' instead of 'Model' and 'SCUB-D' instead of 'SCUBA-D.' Model names are also inconsistent between 'ProtDiff' and 'ProDiff' in Section 3.1.2; please standardize.","section":"Table 2 and Section 4.2"},{"comment":"The text refers to 'Figure 3b' for both the dihedral-angle MSE and the designability across lengths, which appears to be a figure-numbering error. Please renumber the panels so that each reference is unambiguous.","section":"Figure 3 and Section 4.2"},{"comment":"The conditional-generation table uses checkmarks without stating the evaluation procedure. It is unclear whether these were determined from documentation, from the authors' experiments, or from published results. Since conditional capabilities are part of the model comparison, please specify the source and criteria for each checkmark.","section":"Table 4"},{"comment":"The effective sample sizes after the sc-TM filter are not reported. Please state, for each model and length, how many of the 100 generated structures survived the 5th-percentile threshold and were used for the novelty, naturalness, and diversity analyses.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is positioned as a review but contains original benchmark analyses that are central to its conclusions. The major risk is that the ranking of Cα-only models is built on an undocumented ProteinMPNN/OmegaFold reconstruction, and the 'balanced performance' claim may not survive a rigorous re-analysis. I would encourage the editor to require that the authors make the evaluation scripts and the reconstruction protocol available, or at least provide a detailed step-by-step description, before further consideration. The novelty comparison is also confounded by differing training sets; this should be addressed before publication. The paper's topic is timely and the review portions are useful, so a major revision is warranted rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this before citing the benchmark numbers: the review half is solid, but the head-to-head ranking of Cα-only models is built on an undocumented reconstruction pipeline. For FoldingDiff, Genie, and ProtDiff, the authors use ProteinMPNN to design sequences from Cα-only outputs, OmegaFold to fold those sequences, and then keep the highest-TM prediction. They never describe how the missing N, C, O backbone atoms were reconstructed, how many sequences were sampled per structure, or how often the round-trip failed. Because RFdiffusion and Chroma output full backbones and escape this proxy, the \"balanced performance\" claim is not a like-for-like comparison.\n\nWhat is actually new: the benchmark of eight models on shared tasks—efficiency, structural plausibility, designability, novelty, naturalness, diversity—and the conditional-generation capability table. I have not seen that specific head-to-head in earlier reviews. The figures are legible, the metric choices are reasonable, and the survey of point-cloud, frame-cloud, feature-map, and latent-space families is competent. Readers new to the area will get a useful map.\n\nThe soft spots are real. There are no error bars or significance tests anywhere in Section 4. The designability filter uses the 5th percentile of native sc-TM as a hard cutoff, but the paper never says how many generated structures were discarded, so the downstream novelty, naturalness, and diversity numbers describe a filtered subset whose size is unknown. Novelty is computed against each model's training set, but the sets are only listed, not deduplicated in any described way. And with no code, data, or seeds, the benchmark is not reproducible. The conditional-generation table is filled from the authors' judgment, not new experiments, and should be labeled as such.\n\nIs the qualitative conclusion wrong? Probably not—RFdiffusion and Chroma are genuinely strong, widely used tools. But the evidence presented here does not support a precise quantitative ranking, and the stress-test concern about Cα proxies is exactly on target.\n\nThe paper is for someone who wants a compact survey and a rough sense of model strengths, not for someone who needs reliable performance numbers. I would send it to peer review with the expectation of major revision: document or fix the Cα reconstruction, report variance, disclose filter counts, and release code and data. If the authors cannot do that, the benchmark should be published as a qualitative comparison, not a ranked evaluation.","headline":"A useful diffusion-model review whose headline ranking is compromised by an unvalidated ProteinMPNN/OmegaFold reconstruction for Cα-only models.","tokens_in":13612,"tokens_out":3058,"would_cite":false,"duration_ms":26644,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In a head-to-head evaluation of eight diffusion-based protein backbone generators, RFdiffusion and Chroma show the most balanced performance across efficiency, structural plausibility, designability, novelty, naturalness, and diversity.","keywords":["de novo protein design","diffusion models","protein backbone generation","generative AI","RFDiffusion","Chroma","protein structure evaluation","self-consistent TM-score"],"falsifier":"Re-running the sc-TM evaluation with AlphaFold2 in place of OmegaFold, or expressing the generated proteins and measuring which designs actually fold, would settle the ranking; if RFdiffusion and Chroma no longer lead, the balanced-performance claim is an artifact of the evaluation pipeline.","tokens_in":12622,"feed_emoji":"🧬","tokens_out":9028,"duration_ms":79549,"temperature":0.7,"pith_summary":"This review argues that diffusion models have become the most promising generative-AI route to de novo protein design, outpacing fragment-based and physics-guided methods in success rate and cost. To back that claim, it compares eight publicly available diffusion-based backbone generators under one protocol, measuring efficiency, structural plausibility, designability, novelty, naturalness, and diversity. Its central result is that RFdiffusion and Chroma are the most balanced choices overall, while FrameDiff leads on structural plausibility, RFdiffusion on designability, Chroma on diversity, and FoldingDiff on novelty and speed. The paper also collects experimentally validated successes for RFdiffusion, Chroma, and SCUBA-D, and it is candid that current models ignore conformational flexibility, protein–ligand interactions, and direct functional optimization.","feed_headline":"Benchmark names RFdiffusion and Chroma most balanced protein designers","feed_subtitle":"A side-by-side comparison across six design criteria tells practitioners which generator to try first.","key_machinery":"The load-bearing machinery is the diffusion process applied to a chosen representation of the protein backbone, plus the six-metric evaluation protocol that lets the models be compared. Backbone generators are grouped by representation: 2D feature maps (ProteinSGM, SCUBA-D), point clouds of $C_\\alpha$ coordinates (ProtDiff, Genie), frame clouds in which each residue carries a translation and rotation in $\\mathrm{SE}(3)$ (FrameDiff, RFdiffusion, Chroma, FoldingDiff), and latent-space vectors (PVQD). Frame-cloud diffusion is the representation behind the two models judged most balanced, because retaining inter-residue orientations lets the model capture chirality and residue interactions. The evaluation protocol defines designability through self-consistent TM-score: a generated backbone is reverse-folded into a sequence by ProteinMPNN or CarbonDesign, that sequence is refolded by OmegaFold, and the TM-score between generated and predicted structures measures whether the design is realizable; the same protocol supplies the dihedral-angle proxy that lets Cα-only models be scored for structural plausibility.","core_discovery":"The paper's central claim is that diffusion-based generative models now define the leading edge of de novo protein design, and that within this family no single model dominates: RFdiffusion and Chroma show the most balanced performance across all six evaluation axes, with each other model winning a specific niche. FrameDiff generates backbones whose dihedral-angle statistics most closely match native proteins; RFdiffusion achieves the highest self-consistent TM-scores across all tested lengths; Chroma produces the most structurally diverse populations while retaining natural secondary-structure composition; and FoldingDiff is the fastest and most novel generator for short chains. The authors further claim that Chroma completes all six conditional design tasks in their table—binder design, symmetric oligomers, secondary-structure-constrained design, category-based design, structure refinement, and partial-sequence-constrained design—whereas RFdiffusion excels specifically at spatially constrained tasks. The review supports these claims with a shared evaluation protocol and with experimental case studies in which designed binders, soluble proteins, and heme-binding proteins were verified in the lab.","pith_inferences":["The balanced-performance ranking may be specific to unconditional generation; in conditional tasks, the paper's own table shows Chroma dominating, so a practitioner optimizing binder design might reasonably weight RFdiffusion's validated binder successes more heavily than the six-metric average.","Because the six-metric protocol evaluates structure-generation models only, a direct comparison that adds sequence-diffusion models (EvoDiff, TaxDiff) and co-generation models under the same metrics could show whether the backbone-first advantage is real or an artifact of benchmark scope.","A cheap falsification check would be to re-run the benchmark with AlphaFold2 in place of OmegaFold in the sc-TM pipeline; if the relative rankings of FrameDiff, RFdiffusion, and Chroma shift materially, the paper's conclusions describe that particular pipeline.","The review's emphasis on experimental validation suggests a concrete next step: express a matched set of designs from several generators in the same host and compare soluble expression and binding success rates, which would test whether the in silico 'balance' translates to the lab."],"forward_implications":["For a general-purpose de novo backbone design task, the default should be RFdiffusion or Chroma, since the benchmark treats them as the most balanced across all six criteria.","When a design goal prioritizes native-like backbone geometry, FrameDiff is the indicated generator; when it prioritizes designability, RFdiffusion; when diversity of conformations matters, Chroma.","For quick exploration of short, structurally novel proteins, FoldingDiff offers the best efficiency–novelty trade-off among the tested models.","Conditional design—especially binder and symmetric-oligomer construction—is currently best served by Chroma, with RFdiffusion as the strong alternative for spatially constrained tasks.","Because Cα-only generators (ProtDiff, Genie) must be evaluated through a reverse-folding/refolding proxy, their apparent efficiency comes with an added layer of uncertainty that full-atom or frame-based models do not carry."],"supporting_citations":[{"why":"Introduces RFdiffusion and the experimentally validated binder-design results the review treats as its flagship success case.","marker":"[8]"},{"why":"Introduces Chroma, the other model judged most balanced, and the solubility-validation experiments cited.","marker":"[22]"},{"why":"Introduces FoldingDiff, the efficient and high-novelty baseline evaluated in the comparison.","marker":"[17]"},{"why":"Introduces FrameDiff, the structural-plausibility leader in the benchmark.","marker":"[21]"},{"why":"Introduces ProtDiff and the self-consistent TM-score method used to measure designability.","marker":"[18]"},{"why":"Supplies ProteinMPNN, one of the two inverse-folding networks used to backfold generated backbones in the evaluation.","marker":"[32]"},{"why":"Supplies OmegaFold, the structure predictor used to refold designed sequences and compute sc-TM.","marker":"[34]"},{"why":"Supplies CarbonDesign, the second inverse-folding network used in the evaluation.","marker":"[33]"},{"why":"Introduces SCUBA-D, a feature-map-based generator and the source of the heme-binding and Ras design case studies.","marker":"[16]"},{"why":"Supplies Foldseek, the structural clustering method used for the diversity assessment.","marker":"[35]"}],"fun_headline_variants":["Diffusion models choreograph de novo protein design","RFdiffusion and Chroma top balanced protein design","Diffusion models each excel at a protein design task","No one diffusion model wins all protein design challenges","Chroma and RFdiffusion: balanced protein design leaders"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the computational round-trip—designing an amino acid sequence for a generated backbone, predicting its folded structure, and taking the best agreement score—reveals which generated backbones can genuinely fold into their intended shapes; if that round-trip is unfaithful, the rankings measure the pipeline rather than the generators.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion models choreograph de novo protein design","RFdiffusion and Chroma top balanced protein design","Diffusion models each excel at a protein design task","No one diffusion model wins all protein design challenges","Chroma and RFdiffusion: balanced protein design leaders"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000873,"raw_usage":{"total_tokens":3778,"prompt_tokens":942,"completion_tokens":2836,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":2761}},"tokens_in":558,"tokens_out":2836,"duration_ms":19784,"temperature":1.0,"reasoning_tokens":2761,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:02:08.533441+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-running the sc-TM evaluation with AlphaFold2 in place of OmegaFold, or expressing the generated proteins and measuring which designs actually fold, would settle the ranking; if RFdiffusion and Chroma no longer lead, the balanced-performance claim is an artifact of the evaluation pipeline.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces RFdiffusion and the experimentally validated binder-design results the review treats as its flagship success case."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces FoldingDiff, the efficient and high-novelty baseline evaluated in the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies ProteinMPNN, one of the two inverse-folding networks used to backfold generated backbones in the evaluation."},{"cited_title":"& Zhang, H","cited_arxiv_id":null,"evidence_quote":"Supplies CarbonDesign, the second inverse-folding network used in the evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Foldseek, the structural clustering method used for the diversity assessment."}],"review_version":1}