{"id":"951a3097-03a4-41d2-aeb3-2061c7a01e01","arxiv_id":"2608.10598","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Full-chain structural novelty is not evidence of fold invention, because generated backbones mostly contain known domains and a zero-training retrieval baseline reproduces the same novelty profile.","lead":"Protein design papers often claim to invent new folds when their generated structures look unlike any known protein. This study shows the measure is misleading: most 'novel' backbones contain known structural pieces, and a simple retrieval-based method can match the same statistics without learning anything.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that tTM@0.7 is the only calibrated novelty score rests on a domain-level negative control; a chain-level negative control, directly constructible from RetFold's manifest, is missing and could change the calibration conclusion.","rationale":"The reader's weakest_assumption correctly identifies the leave-one-topology-out negative control as the hinge for the calibration conclusion, and notes the domain/chain mismatch. I agree that this is the most load-bearing concern. I would add that the paper's own caveat in Appendix G explicitly disclaims transferring FPRs to chains, which creates a tension with the unqualified abstract claim that only tTM@0.7 supports novelty reading. A concrete chain-level negative control is feasible using RetFold's manifest, and it would directly settle the transferability question. The central claim that full-chain novelty is insufficient evidence of fold-level innovation is robust: RetFold's 20.0% qTM retrieval on known-by-construction recombination already proves that the conventional protocol misclassifies recombination as novelty. Therefore the overall verdict should remain CONDITIONAL, pending either the missing chain-level control or a more carefully scoped conclusion. The reader's verdict is unchanged; my concern reinforces the condition rather than overturning it.","tokens_in":22495,"tokens_out":7950,"duration_ms":77124,"concrete_test":"For each of the 500 RetFold chains, use the construction manifest to remove from the CATH S40 search set (i) the three exact source fragments and (ii) all CATH domains sharing their topology; recompute the full-chain tTM>=0.7 retrieval rate. If this rate remains near 5.38%, the domain-level FPR transfers to chains and the calibration claim holds. If it is substantially higher, the paper's assertion that tTM@0.7 is the only novelty-supporting score at chain level is unsupported and would require either a new chain-level negative control or a qualified conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central prescriptive claim is that only reference-normalized TM at tau=0.7 supports reading a retrieval rate as a novelty estimate. This is supported by the leave-one-topology-out negative control in Appendix G, where tTM reaches a 5.38% false positive rate at tau=0.7 on domain-sized queries. However, the actual novelty decisions in Table 3 are made on full-chain queries searched against the same domain reference. The paper explicitly disclaims transferring these FPRs to chains ('we therefore do not use them to judge whether a given chain-sized retrieval rate exceeds chance'), yet the abstract and conclusion generalize the calibration: 'reference-normalized TM reaches a 5.38% false positive rate only at 0.7' and 'only one of the three supports reading a retrieval rate as a novelty estimate' are stated without the domain/chain qualifier. A full chain has more surface area and more opportunities to partially cover a reference domain, so its tTM@0.7 FPR could be substantially higher than the domain-level 5.38%. The positive control (RetFold, 100% saturated at tTM>=0.7) demonstrates sensitivity on chains but says nothing about specificity on chains. Without a chain-level negative control, the conclusion that tTM@0.7 separates recombined from invented folds at the chain level is not established. This is not a hidden flaw—the authors flag the limitation—but it is load-bearing because the calibrated criterion is the paper's main positive contribution beyond the RetFold demonstration.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that low full-chain similarity to known proteins is insufficient evidence of fold-level novelty in protein backbone generation. It introduces the Domain Retrieval Rate (DRR), which decomposes generated backbones into structural domains and measures retrieval against the CATH S40 domain database, and proposes RetFold, a zero-training baseline that assembles known CATH domains into recombined chains via geometry-based helix-linker optimization. The authors calibrate three structural-similarity scores (alnTM, qTM, tTM) against RetFold as a positive control with known composition and against a leave-one-topology-out negative control, concluding that only reference-normalized tTM at a threshold of 0.7 supports reading a retrieval rate as a novelty estimate. Across eight backbone generators, they report that most outputs contain retrievable domains even when full-chain retrieval is low, and that RetFold lands inside the retrieval envelope of learned generators at roughly two orders of magnitude lower cost.","tokens_in":22804,"tokens_out":9947,"duration_ms":82965,"significance":"If the chain-level calibration is established, the paper makes a significant methodological contribution. It demonstrates a granularity mismatch in existing novelty evaluation, provides a strong positive control with known composition, and quantifies score-specific false-positive rates. The paper is unusually thorough: it includes multiple segmentations (Merizo and Chainsaw), quality filtering, per-length analysis, sensitivity scans over thresholds and coverage rules, and explicit self-flagged limitations. The RetFold baseline is a creative and useful diagnostic, and the paper's core message—that full-chain novelty is not evidence of fold invention—is well supported even independently of the specific calibration threshold.","major_comments":[{"comment":"The central prescriptive claim—that reference-normalized tTM at τ=0.7 is the only score that supports reading a retrieval rate as a novelty estimate—rests on a negative control measured on domain-sized queries. Table 2 reports a 5.38% FPR for tTM at τ=0.7 on CATH S40 domains after leave-one-topology-out exclusion, and Appendix G explicitly states: 'we therefore do not use them to judge whether a given chain-sized retrieval rate exceeds chance.' Yet Table 3 applies the criterion to full-chain queries, and the abstract and conclusion state the 5.38% figure and the 'only one of the three' conclusion without the domain/chain qualifier. A full chain has more surface area and more opportunities to partially cover a reference domain, so its tTM@0.7 FPR could be substantially higher than the domain-level 5.38%. The positive control (RetFold, 100% saturated at tTM≥0.7) demonstrates sensitivity on chains but says nothing about specificity on chains. Without a chain-level negative control—for example, assembling CATH domains into chimeric chains with known non-matching topologies—the calibrated separation of RetFold from the unconditional generators in Table 3 is not established. The authors should either add such a control or restrict the conclusion to domain-level queries.","section":"Appendix G, Table 3, Abstract"}],"minor_comments":[{"comment":"The 5.38% figure is described as a false positive rate, but Appendix G defines it as an upper bound because the leave-one-topology-out exclusion removes a discrete label rather than a region of structure space; the text should consistently say 'at most 5.38%' or 'upper bound of 5.38%'.","section":"Abstract and Section 'A zero-training baseline exposes the resolution limit'"},{"comment":"The sentence 'the true floor is correspondingly higher' is ambiguous; it should clarify that the reference Δ value is correspondingly higher because the qTM false-positive rate on chains is even lower than the domain-level 21.06%.","section":"Appendix G, 'Noise Floor Details'"},{"comment":"AUC is reported but not defined in the caption or text; please define it as the area under the ROC curve for distinguishing same-topology from cross-topology queries.","section":"Table 2"},{"comment":"The label 'RFD3' is used for RFDiffusion without prior definition; please introduce the abbreviation at first use in the text or figure caption.","section":"Figures 3 and 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid methodological contribution with an unusually honest limitations section. The main scientific gap is the chain-level calibration of tTM@0.7; this is fixable with a chain-level negative control or by qualifying the conclusions. If the authors address this, I would support publication. No concerns about novelty or citation practices."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe short version: this paper makes a needed point—low full-chain similarity is a weak proxy for fold invention—and backs it with a genuinely useful diagnostic (DRR) and a clever zero-training baseline (RetFold). The headline number, RetFold retrieving at only 20% full-chain while recombined by construction, should make anyone in the field pause. It deserves a serious referee.\n\nWhat is actually new: DRR decomposes generated backbones into domains and measures how many contain at least one CATH S40-matching unit. RetFold assembles unmodified CATH domains with helix-linker refinement, giving a positive control with known composition. The calibration of the three TM-score conventions against positive and negative controls is the reality check the novelty literature has been missing. The paper is also unusually honest: it flags its own limitations, including the domain/chain gap, the non-exhaustive CATH S40, and the PXDesign ablation mode.\n\nThe main soft spot is the one in the stress-test note. The 5.38% false positive rate for tTM@0.7 comes from a leave-one-topology-out control on domain-sized queries, but the application in Table 3 is to full chains. The paper explicitly says the FPRs do not transfer, yet the abstract and conclusion state the calibration without that qualifier. That is an over-generalization, and it is load-bearing for the specific claim that tTM@0.7 is the only calibrated score. A chain-level negative control—directly constructible from RetFold's manifest—could move that number substantially. It would likely not erase the large separation between RetFold and the unconditional generators, but it could change the threshold recommendation.\n\nThe bigger reproducibility issue is the absence of code and data. The methods are described in enough detail to reimplement, but not to verify the eight model runs or the exact retrievals. Given the paper's role is to set a new evaluation standard, that is a real gap.\n\nMy verdict: the qualitative claim—full-chain novelty coexists with high domain retrievability, and simple recombination reproduces the same retrieval profile—is solid. The quantitative calibration claim needs refinement. I'd send it to peer review with a request for chain-level calibration and code/data release.","headline":"A solid evaluation paper that makes a real point about full-chain novelty, but its calibrated-threshold claim is over-generalized from domain-level to chain-level queries.","tokens_in":23358,"tokens_out":2573,"would_cite":true,"duration_ms":23782,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Low full-chain novelty does not prove a protein generator invented a new fold.","keywords":["protein backbone generation","domain retrieval rate","novelty evaluation","RetFold","TM-score calibration","CATH S40","fold-space exploration"],"falsifier":"A direct test would be to build a chain-sized negative control: take RetFold's 500 chains, remove the specific CATH domains used to construct each chain from the reference set by structural exclusion (not by topology label), and measure how often each score retrieves the remaining structure above threshold. If reference-normalized tTM at 0.7 shows a false positive rate well above 5.38 percent on such chains, the paper's claim that only tTM at 0.7 supports a novelty reading would be weakened.","tokens_in":22178,"feed_emoji":"🧬","tokens_out":3541,"duration_ms":26332,"temperature":0.7,"pith_summary":"The paper argues that reporting low full-chain similarity to known proteins is not evidence that a generative model has invented a new fold. It introduces the Domain Retrieval Rate (DRR), which decomposes each generated backbone into domains and checks whether any domain matches a known domain in the CATH S40 database. Across eight backbone generation models, most outputs contain locally alignable known structure, even when full-chain retrieval rates are low. To calibrate what retrieval alone can achieve, the paper builds RetFold, a zero-training baseline that assembles backbones from real CATH domains with geometry-based linker refinement. RetFold is recombined by construction, yet the conventional full-chain protocol retrieves it only 20 percent of the time, showing that low full-chain similarity cannot certify fold-level invention. The paper concludes that novelty rates should be reported alongside calibration of the score used.","feed_headline":"Low full-chain novelty is not proof of a new fold","feed_subtitle":"A retrieval-only baseline reproduces the novelty profiles of learned generators, so novelty scores need calibration.","key_machinery":"The central object is RetFold, a zero-training retrieval-and-assembly baseline that constructs backbones by retrieving three fragments from CATH S40 (one core domain at 40-60 percent of target length and two extensions at 15-35 percent each), using an ensemble of ESM-2, Foldseek 3Di, and ProstT5 embeddings, then connecting them with idealized alpha-helix linkers optimized for geometry and diversity. Because RetFold's composition is known by construction, it serves as a positive control for any retrieval criterion. The paper also uses DRR as the diagnostic metric: it decomposes generated backbones with Merizo and asks whether any domain matches a CATH S40 domain at aligned-length TM above a threshold. The argument hinges on calibrating three TM-score normalizations: query-normalized qTM (which carries a length ceiling), aligned-length alnTM (which has high false positive rates under leave-one-topology-out), and target-normalized tTM (the only one that passes both positive and negative controls at threshold 0.7).","core_discovery":"The paper's central claim is that apparent full-chain novelty coexists with high domain-level retrievability, and that the conventional protocol of reporting low full-chain TM-score against a structural database does not distinguish a genuinely new fold from a novel assembly of known domains. It demonstrates this with the Domain Retrieval Rate (DRR), which finds that a majority of generated backbones from eight models contain domains aligned to known CATH S40 entries, even when full-chain retrieval rates are near zero. The paper then introduces RetFold, a training-free pipeline that retrieves CATH domains and assembles them into chains with helix linkers; because its composition is known by construction, any criterion applied to it has a known answer. The conventional full-chain query-normalized protocol retrieves RetFold at only 20.0 percent, while the reference-normalized, coverage-controlled protocol retrieves it at 100.0 percent even at a threshold of 0.9. Against a leave-one-topology-out negative control, only the reference-normalized TM-score reaches a false positive rate below 10 percent (5.38 percent at threshold 0.7), while aligned-length TM retrieves 90.04 percent of queries whose fold class has been removed from the searched set. Therefore, the paper concludes, the retrieval envelope reported for learned generators does not by itself certify fold invention, and novelty claims should be calibrated against positive and negative controls.","pith_inferences":["A testable extension is to run a chain-sized negative control that excludes near matches by structural exclusion rather than by discrete CATH topology labels, which would directly calibrate the false positive rates for full-chain queries.","The paper implicitly suggests a 'novelty calibration report' standard: every novelty rate should be accompanied by the rate the same protocol assigns to a known-recombined positive control and a known-absent negative control.","Because RetFold's 20 percent full-chain retrieval comes from its large verbatim core domain, the result may extend to other fragment-assembly methods: any generator that uses long unmodified reference fragments will be misclassified as novel by qTM for length reasons alone.","A natural next experiment is to apply the calibrated tTM criterion to conditional binder generation in documented inference mode, where the paper's current task-matched cohorts were run under target-free ablations."],"forward_implications":["Novelty rates in protein structure generation should be reported with the calibration of the score used, including false positive rates on a negative control whose fold class is absent from the reference set.","A generated backbone that fails full-chain retrieval is not evidence of a new fold; it may be a recombination of known domains or a distorted version of a known fold.","The domain–full-chain gap, DRR minus FC, should not be read as a measure of structural content unless the reference value under a true-negative control is subtracted.","Reference-normalized TM-score with coverage control, at a threshold of 0.7, separates a recombined baseline from learned unconditional generators, suggesting a protocol that can distinguish recombination from generated novelty.","Designability screens can pass near-identical retrieved copies if the reference database contains close homologs, so reported pass rates should be accompanied by a matched trivial-baseline comparison."],"supporting_citations":[{"why":"Establishes CATH as the domain hierarchy used to map known fold space and to define the reference database CATH S40.","marker":"(Orengo et al. 1997)"},{"why":"Supplies the conventional reading of TM-score 0.5 as a same-fold threshold used to set the retrieval thresholds.","marker":"(Xu and Zhang 2010)"},{"why":"Provides Merizo, the domain segmentation tool that defines the units in the DRR decomposition.","marker":"(Lau, Kandathil, and Jones 2023)"},{"why":"Provides Foldseek's 3Di structural search, used as one of RetFold's retrieval embeddings and for the TM-score rescoring protocol.","marker":"(Van Kempen et al. 2024)"},{"why":"Provides ESM-2 sequence embeddings used in RetFold's ensemble retrieval of CATH domains.","marker":"(Lin et al. 2023)"},{"why":"Provides ProstT5 sequence-structure embeddings used in RetFold's ensemble retrieval.","marker":"(Heinzinger et al. 2024)"},{"why":"One of the key learned generators (RFDiffusion) whose reported novelty rates are recalibrated against the controls.","marker":"(Watson et al. 2023)"},{"why":"Supports the modularity premise that natural proteins are composed of recombining domains.","marker":"(Chothia 1992)"}],"fun_headline_variants":["Full-chain novelty masks known-domain reuse","Most generated proteins contain known domains","Retrieval baseline matches learned generators","Novelty scores need positive and negative controls","Generated proteins often assemble known folds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The key assumption is that removing a query's CATH topology label from the searched reference set removes all true matches, so that any hit above threshold in that control is a false positive; but because the exclusion is a discrete label rather than a structural region, cross-topology hits at TM above 0.5 include genuine similarity, and the authors state their false positive rates are upper bounds.","fun_headline_variants_meta":{"raw":{"variants":["Full-chain novelty masks known-domain reuse","Most generated proteins contain known domains","Retrieval baseline matches learned generators","Novelty scores need positive and negative controls","Generated proteins often assemble known folds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1429,"prompt_tokens":981,"completion_tokens":448,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":387}},"tokens_in":597,"tokens_out":448,"duration_ms":5338,"temperature":1.0,"reasoning_tokens":387,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:12:47.242414+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be to build a chain-sized negative control: take RetFold's 500 chains, remove the specific CATH domains used to construct each chain from the reference set by structural exclusion (not by topology label), and measure how often each score retrieves the remaining structure above threshold. If reference-normalized tTM at 0.7 shows a false positive rate well above 5.38 percent on such chains, the paper's claim that only tTM at 0.7 supports a novelty reading would be weakened.","supporting_citations":[{"cited_title":"2024 , publisher=","cited_arxiv_id":null,"evidence_quote":"Provides ProstT5 sequence-structure embeddings used in RetFold's ensemble retrieval."},{"cited_title":"De novo design of protein structure and function with","cited_arxiv_id":null,"evidence_quote":"One of the key learned generators (RFDiffusion) whose reported novelty rates are recalibrated against the controls."}],"review_version":1}