Pith. sign in

REVIEW 1 major objections 4 minor 34 references

Is Retrieval All You Need? Assessment and Emergence of Novelty in Protein Structure Generation

T0 review · 1 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Low full-chain novelty does not prove a protein generator invented a new fold.

desk verdict A solid evaluation paper that makes a real point about full-chain novelty, but its calibrated-threshold claim is over-generalized from domain-level to chain-level queries. read the letter →

arxiv 2608.10598 v1 pith:S3KEU3NR submitted 2026-08-11 q-bio.BM

classification q-bio.BM
keywords proteinbackbonegenerationdomainretrievalratenoveltyevaluationRetFoldTM-scorecalibrationCATHS40fold-spaceexploration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that reporting low full-chain similarity to known proteins is not evidence that a generative model has invented a new fold. It introduces the Domain Retrieval Rate (DRR), which decomposes each generated backbone into domains and checks whether any domain matches a known domain in the CATH S40 database. Across eight backbone generation models, most outputs contain locally alignable known structure, even when full-chain retrieval rates are low. To calibrate what retrieval alone can achieve, the paper builds RetFold, a zero-training baseline that assembles backbones from real CATH domains with geometry-based linker refinement. RetFold is recombined by construction, yet the conventional full-chain protocol retrieves it only 20 percent of the time, showing that low full-chain similarity cannot certify fold-level invention. The paper concludes that novelty rates should be reported alongside calibration of the score used.

What carries the argument

The central object is RetFold, a zero-training retrieval-and-assembly baseline that constructs backbones by retrieving three fragments from CATH S40 (one core domain at 40-60 percent of target length and two extensions at 15-35 percent each), using an ensemble of ESM-2, Foldseek 3Di, and ProstT5 embeddings, then connecting them with idealized alpha-helix linkers optimized for geometry and diversity. Because RetFold's composition is known by construction, it serves as a positive control for any retrieval criterion. The paper also uses DRR as the diagnostic metric: it decomposes generated backbones with Merizo and asks whether any domain matches a CATH S40 domain at aligned-length TM above a threshold. The argument hinges on calibrating three TM-score normalizations: query-normalized qTM (which carries a length ceiling), aligned-length alnTM (which has high false positive rates under leave-one-topology-out), and target-normalized tTM (the only one that passes both positive and negative controls at threshold 0.7).

What would settle it

A direct test would be to build a chain-sized negative control: take RetFold's 500 chains, remove the specific CATH domains used to construct each chain from the reference set by structural exclusion (not by topology label), and measure how often each score retrieves the remaining structure above threshold. If reference-normalized tTM at 0.7 shows a false positive rate well above 5.38 percent on such chains, the paper's claim that only tTM at 0.7 supports a novelty reading would be weakened.

Watch

Extended reading notes

Core claim

The paper's central claim is that apparent full-chain novelty coexists with high domain-level retrievability, and that the conventional protocol of reporting low full-chain TM-score against a structural database does not distinguish a genuinely new fold from a novel assembly of known domains. It demonstrates this with the Domain Retrieval Rate (DRR), which finds that a majority of generated backbones from eight models contain domains aligned to known CATH S40 entries, even when full-chain retrieval rates are near zero. The paper then introduces RetFold, a training-free pipeline that retrieves CATH domains and assembles them into chains with helix linkers; because its composition is known by construction, any criterion applied to it has a known answer. The conventional full-chain query-normalized protocol retrieves RetFold at only 20.0 percent, while the reference-normalized, coverage-controlled protocol retrieves it at 100.0 percent even at a threshold of 0.9. Against a leave-one-topology-out negative control, only the reference-normalized TM-score reaches a false positive rate below 10 percent (5.38 percent at threshold 0.7), while aligned-length TM retrieves 90.04 percent of queries whose fold class has been removed from the searched set. Therefore, the paper concludes, the retrieval envelope reported for learned generators does not by itself certify fold invention, and novelty claims should be calibrated against positive and negative controls.

Load-bearing premise

The key assumption is that removing a query's CATH topology label from the searched reference set removes all true matches, so that any hit above threshold in that control is a false positive; but because the exclusion is a discrete label rather than a structural region, cross-topology hits at TM above 0.5 include genuine similarity, and the authors state their false positive rates are upper bounds.

Editorial extensions

If this is right

  • Novelty rates in protein structure generation should be reported with the calibration of the score used, including false positive rates on a negative control whose fold class is absent from the reference set.
  • A generated backbone that fails full-chain retrieval is not evidence of a new fold; it may be a recombination of known domains or a distorted version of a known fold.
  • The domain–full-chain gap, DRR minus FC, should not be read as a measure of structural content unless the reference value under a true-negative control is subtracted.
  • Reference-normalized TM-score with coverage control, at a threshold of 0.7, separates a recombined baseline from learned unconditional generators, suggesting a protocol that can distinguish recombination from generated novelty.
  • Designability screens can pass near-identical retrieved copies if the reference database contains close homologs, so reported pass rates should be accompanied by a matched trivial-baseline comparison.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to run a chain-sized negative control that excludes near matches by structural exclusion rather than by discrete CATH topology labels, which would directly calibrate the false positive rates for full-chain queries.
  • The paper implicitly suggests a 'novelty calibration report' standard: every novelty rate should be accompanied by the rate the same protocol assigns to a known-recombined positive control and a known-absent negative control.
  • Because RetFold's 20 percent full-chain retrieval comes from its large verbatim core domain, the result may extend to other fragment-assembly methods: any generator that uses long unmodified reference fragments will be misclassified as novel by qTM for length reasons alone.
  • A natural next experiment is to apply the calibrated tTM criterion to conditional binder generation in documented inference mode, where the paper's current task-matched cohorts were run under target-free ablations.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. The paper argues that low full-chain similarity to known proteins is insufficient evidence of fold-level novelty in protein backbone generation. It introduces the Domain Retrieval Rate (DRR), which decomposes generated backbones into structural domains and measures retrieval against the CATH S40 domain database, and proposes RetFold, a zero-training baseline that assembles known CATH domains into recombined chains via geometry-based helix-linker optimization. The authors calibrate three structural-similarity scores (alnTM, qTM, tTM) against RetFold as a positive control with known composition and against a leave-one-topology-out negative control, concluding that only reference-normalized tTM at a threshold of 0.7 supports reading a retrieval rate as a novelty estimate. Across eight backbone generators, they report that most outputs contain retrievable domains even when full-chain retrieval is low, and that RetFold lands inside the retrieval envelope of learned generators at roughly two orders of magnitude lower cost.

Significance. If the chain-level calibration is established, the paper makes a significant methodological contribution. It demonstrates a granularity mismatch in existing novelty evaluation, provides a strong positive control with known composition, and quantifies score-specific false-positive rates. The paper is unusually thorough: it includes multiple segmentations (Merizo and Chainsaw), quality filtering, per-length analysis, sensitivity scans over thresholds and coverage rules, and explicit self-flagged limitations. The RetFold baseline is a creative and useful diagnostic, and the paper's core message—that full-chain novelty is not evidence of fold invention—is well supported even independently of the specific calibration threshold.

major comments (1)
  1. [Appendix G, Table 3, Abstract] The central prescriptive claim—that reference-normalized tTM at τ=0.7 is the only score that supports reading a retrieval rate as a novelty estimate—rests on a negative control measured on domain-sized queries. Table 2 reports a 5.38% FPR for tTM at τ=0.7 on CATH S40 domains after leave-one-topology-out exclusion, and Appendix G explicitly states: 'we therefore do not use them to judge whether a given chain-sized retrieval rate exceeds chance.' Yet Table 3 applies the criterion to full-chain queries, and the abstract and conclusion state the 5.38% figure and the 'only one of the three' conclusion without the domain/chain qualifier. A full chain has more surface area and more opportunities to partially cover a reference domain, so its tTM@0.7 FPR could be substantially higher than the domain-level 5.38%. The positive control (RetFold, 100% saturated at tTM≥0.7) demonstrates sensitivity on chains but says nothing about specificity on chains. Without a chain-level negative control—for example, assembling CATH domains into chimeric chains with known non-matching topologies—the calibrated separation of RetFold from the unconditional generators in Table 3 is not established. The authors should either add such a control or restrict the conclusion to domain-level queries.
minor comments (4)
  1. [Abstract and Section 'A zero-training baseline exposes the resolution limit'] The 5.38% figure is described as a false positive rate, but Appendix G defines it as an upper bound because the leave-one-topology-out exclusion removes a discrete label rather than a region of structure space; the text should consistently say 'at most 5.38%' or 'upper bound of 5.38%'.
  2. [Appendix G, 'Noise Floor Details'] The sentence 'the true floor is correspondingly higher' is ambiguous; it should clarify that the reference Δ value is correspondingly higher because the qTM false-positive rate on chains is even lower than the domain-level 21.06%.
  3. [Table 2] AUC is reported but not defined in the caption or text; please define it as the area under the ROC curve for distinguishing same-topology from cross-topology queries.
  4. [Figures 3 and 7] The label 'RFD3' is used for RFDiffusion without prior definition; please introduce the abbreviation at first use in the text or figure caption.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's controls, calibrations, and negative-control FPRs are externally grounded and explicitly qualified.

full rationale

The paper's load-bearing steps are empirical calibrations against external resources rather than reductions to their own inputs. DRR uses CATH S40 with Merizo segmentation, and the high aligned-length domain retrieval is then bounded by the leave-one-topology-out negative control, whose ground-truth labels come from CATH topology rather than from the scores under test. RetFold is a constructed positive control: its composition is known by construction, but it is used as a sensitivity probe and counterexample, not as a fitted predictor of learned generators. The central 'envelope does not certify invention' argument is an existence proof, a deliberately recombined chain shown to land inside the reported retrieval envelope, and it does not derive the generators' behavior from RetFold. The choice of tTM at tau=0.7 is transparently reported with its 5.38% FPR and 100% sensitivity on the positive control; no fitted parameter is renamed as a prediction. The paper explicitly flags the one place where circular reasoning would arise and avoids it: in Appendix G it states that using the class of a generated backbone's best hit for model-specific FPR corrections 'is circular when the hit itself may be the false positive.' The acknowledged limitation that the negative-control FPRs are measured on domain-sized queries and do not transfer directly to chain-sized retrieval rates is a scope caveat about the evidence, not a definitional equivalence. There are no self-citations used as load-bearing support, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in by citation; all external tools (CATH, Merizo, Foldseek, ProteinMPNN, AlphaFold3) are independent. The paper is self-contained against external benchmarks, so no significant circularity is present.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

DRR is a new metric and RetFold a new pipeline, but both are procedures, not postulated physical entities. No new forces, particles, dimensions, or energy terms are introduced.

free parameters (5)
  • RetFold ensemble embedding weights = 23% ESM-2, 38.5% Foldseek 3Di, 38.5% ProstT5
    Chosen by hand for Stage 1 retrieval; affects which fragments are assembled and thus the retrieval profile.
  • RetFold fragment length windows = core 40-60% L, extensions 15-35% L
    Hand-set assembly constraints; central to RetFold's construction and to its full-chain qTM behavior.
  • RetFold helix-linker parameters = 6-12 residue linkers, rise 1.5 A/res, radius 2.25 A, turn 100 deg/res, axis within +/-65 deg
    Hand-set geometric constants in Stage 2; define the junctions and affect the AF3 confidence results.
  • RetFold filtering thresholds = Calpha distance >3 A, Rg <2.1 x expected, 40 variants per linker length
    Hand-set filters that determine which 500 backbones are retained for evaluation.
  • Evaluation thresholds and coverage rule = tau 0.5/0.7, tcov>=0.7, pLDDT>=70, iPTM>=0.6
    Convention choices on which the entire calibration argument depends; the paper explicitly varies thresholds in sensitivity analyses.
assumptions (5)
  • domain assumption CATH S40 (34,653 non-redundant domains) represents known fold space for this evaluation.
    All retrieval rates are relative to this non-exhaustive reference, as the conclusion explicitly states.
  • domain assumption Merizo domain segmentation recovers the fold units of generated backbones.
    Primary DRR uses Merizo; the Chainsaw ablation in Appendix F shows rates shift by up to 44 percentage points, so the segmentation choice is material.
  • domain assumption Leave-one-topology-out exclusion removes true matches for a query.
    The negative control deletes a discrete CATH label rather than a structural neighborhood; the authors note cross-topology hits are genuine similarity, making false-positive rates upper bounds.
  • standard math TM-score conventions from Xu and Zhang 2010, TM>0.5 as same-fold, and the query-length bound hold.
    Eq. 2 and the 0.5 threshold are imported from cited literature; the paper does not re-derive them.
  • domain assumption AlphaFold3 pLDDT/iPTM and ProteinMPNN constitute a valid designability screen.
    The screen is the shared quality filter; Appendix A shows the screen passes near-copies, so its validity is load-bearing for quality-conditioned claims.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Is Retrieval All You Need? Assessment and Emergence of Novelty in Protein Structure Generation." pith.science (2026). https://pith.science/paper/S3KEU3NR

@misc{pith2026260810598,
  author       = {Pith},
  title        = {Pith review of: Is Retrieval All You Need? Assessment and Emergence of Novelty in Protein Structure Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S3KEU3NR}},
  note         = {Machine review of arXiv:2608.10598}
}
read the original abstract

Protein backbone generation models are often credited with exploring novel fold space based solely on low full-chain similarity to known proteins, yet this cannot distinguish a genuinely new fold from a novel assembly of known structural units. We first ask whether this granularity mismatch alone explains the reported rates, and introduce the Domain Retrieval Rate (DRR), the fraction of generated backbones for which any constituent domain matches a known domain in CATH S40. Applied to eight backbone generation models spanning diffusion and flow-matching paradigms, DRR finds locally alignable known structure in most outputs, while the fraction containing a substantially covered complete domain is considerably smaller and depends on the scoring convention. To calibrate what retrieval alone can achieve, we propose RetFold, a zero-training baseline that constructs backbones by retrieving CATH domains and refining inter-domain connections through geometry-based helix-linker optimization, at two orders of magnitude lower cost on CPU alone.

Figures

Figures reproduced from arXiv: 2608.10598 by the authors.

Figure 1
Figure 1. Contrasting full-chain and domain-level novelty. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of RetFold. CATH S40 domains are embedded with complementary representations. RetFold retrieves one [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Domain-level vs. full-chain retrievability under the unified Foldseek protocol and a fixed denominator of 500 per [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Retrieval-based binder design for all five targets. Target chains are shown in gray; ground-truth binders from the [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Reference-normalized, coverage-controlled domain retrieval over the all-generated cohort using a fixed denominator of [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: Sensitivity of any-domain retrieval to the operational protocol over the all-generated cohort with a fixed denominator [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Per-length DRR vs. FC for all eight learned generators under the primary top-100 protocol (all generated structures, no [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: RetFold length-wise quality-screen pass rate and domain-level retrieval. Any-domain retrieval at alnTM [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: Junction-resolved AF3 confidence for 466 RetFold backbones passing the whole-chain pLDDT threshold. (a) Mean [PITH_FULL_IMAGE:figures/full_fig_p019_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

34 extracted references · 19 canonical work pages

  1. [1]

    Journal of molecular biology , volume=

    Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions , author=. Journal of molecular biology , volume=. 1997 , publisher=

  2. [2]

    Methods in enzymology , volume=

    Protein structure prediction using Rosetta , author=. Methods in enzymology , volume=. 2004 , publisher=

  3. [3]

    Nature structural biology , volume=

    Protein building blocks preserved by recombination , author=. Nature structural biology , volume=. 2002 , publisher=

  4. [4]

    Proceedings of the National Academy of Sciences , volume=

    An all-atom protein generative model , author=. Proceedings of the National Academy of Sciences , volume=. 2024 , publisher=

  5. [5]

    bioRxiv , pages=

    Boltzgen: Toward universal binder design , author=. bioRxiv , pages=. 2025 , publisher=

  6. [6]

    Nature , volume=

    De novo design of protein structure and function with RFdiffusion , author=. Nature , volume=. 2023 , publisher=

  7. [8]

    arXiv preprint arXiv:2302.02277 , year=

    SE(3) Diffusion Model with Application to Protein Backbone Generation , author=. arXiv preprint arXiv:2302.02277 , year=

  8. [9]

    arXiv preprint arXiv:2310.05297 , year=

    Fast protein backbone generation with se (3) flow matching , author=. arXiv preprint arXiv:2310.05297 , year=

Show all 34 references
  1. [10]

    Nature biotechnology , volume=

    Fast and accurate protein structure search with Foldseek , author=. Nature biotechnology , volume=. 2024 , publisher=

  2. [11]

    Nature Communications , volume=

    Merizo: a rapid and accurate protein domain segmentation method using invariant point attention , author=. Nature Communications , volume=. 2023 , publisher=

  3. [12]

    Nature , volume=

    Accurate structure prediction of biomolecular interactions with AlphaFold 3 , author=. Nature , volume=. 2024 , publisher=

  4. [13]

    bioRxiv , pages=

    PXDesign: Fast, modular, and accurate de novo design of protein binders , author=. bioRxiv , pages=. 2025 , publisher=

  5. [14]

    De novo design of protein structure and function with

    Watson, Joseph L and Juergens, David and Bennett, Nathaniel R and Trippe, Brian L and Yim, Jason and Eisenach, Helen E and Ahern, Woody and Borber, Andrew J and Ragotte, Robert J and Milles, Lukas F and others , journal=. De novo design of protein structure and function with. ...

  6. [15]

    Yim, Jason and Trippe, Brian L and De Bortoli, Valentin and Mathieu, Emile and Doucet, Arnaud and Barzilay, Regina and Jaakkola, Tommi , journal=

  7. [16]

    Nature , volume=

    Illuminating protein space with a programmable generative model , author=. Nature , volume=. 2023 , publisher=

  8. [17]

    arXiv preprint arXiv:2310.16624 , year=

    Yim, Jason and Campbell, Andrew and Foong, Andrew Y K and Gastegger, Michael and Jim. arXiv preprint arXiv:2310.16624 , year=

  9. [18]

    arXiv preprint arXiv:2310.02391 , year=

    SE(3)-Stochastic Flow Matching for Protein Backbone Generation , author=. arXiv preprint arXiv:2310.02391 , year=

  10. [19]

    Proceedings of the National Academy of Sciences , volume=

    An all-atom protein generative model , author=. Proceedings of the National Academy of Sciences , volume=

  11. [20]

    Nature Machine Intelligence , volume=

    Predicting equilibrium distributions for molecular systems with deep learning , author=. Nature Machine Intelligence , volume=

  12. [21]

    Science , volume=

    Scaffolding protein functional sites using deep learning , author=. Science , volume=. 2022 , publisher=

  13. [22]

    Robust deep learning--based protein sequence design using

    Dauparas, Justas and Anishchenko, Ivan and Bennett, Nathaniel and Baek, Minkyung and Juergens, David and Ragotte, Robert J and Milles, Lukas F and Wicky, Basile I M and Galber, Alexis and Baker, David , journal=. Robust deep learning--based protein sequence design using. 2022 ...

  14. [23]

    Highly accurate protein structure prediction with

    Jumper, John and Evans, Richard and Pritzel, Alexander and Green, Tim and Figurnov, Michael and Ronneberger, Olaf and Tunyasuvunakool, Kathryn and Bates, Russ and. Highly accurate protein structure prediction with. Nature , volume=. 2021 , publisher=

  15. [24]

    Accurate structure prediction of biomolecular interactions with

    Abramson, Josh and Adler, Jonas and Dunger, Jack and Evans, Richard and Green, Tim and Pritzel, Alexander and Ronneberger, Olaf and Willmore, Lindsay and Ballard, Andrew J and others , journal=. Accurate structure prediction of biomolecular interactions with. 2024 , publisher=

  16. [25]

    Fast and accurate protein structure search with

    van Kempen, Michel and Kim, Stephanie S and Tumescheit, Charlotte and Mirdita, Milot and Lee, Jeongjae and Gilchrist, Cameron L M and S. Fast and accurate protein structure search with. Nature Biotechnology , volume=. 2024 , publisher=

  17. [26]

    Science , volume=

    Evolutionary-scale prediction of atomic-level protein structure with a language model , author=. Science , volume=. 2023 , publisher=

  18. [27]

    2024 , publisher=

    Heinzinger, Michael and Weissenow, Konstantin and Littmann, Maria and Steinegger, Martin and Rost, Burkhard , journal=. 2024 , publisher=

  19. [28]

    arXiv preprint arXiv:2407.04967 , year=

    Generative Augmentation Flows , author=. arXiv preprint arXiv:2407.04967 , year=

  20. [29]

    Nature , volume=

    One thousand families for the molecular biologist , author=. Nature , volume=. 1992 , publisher=

  21. [30]

    1997 , publisher=

    Orengo, Christine A and Michie, Alex D and Jones, Susan and Jones, David T and Swindells, Mark B and Thornton, Janet M , journal=. 1997 , publisher=

  22. [31]

    Journal of Molecular Biology , volume=

    Estimating the number of protein folds and families from complete genome data , author=. Journal of Molecular Biology , volume=. 2000 , publisher=

  23. [32]

    Proceedings of the National Academy of Sciences , volume=

    Nature of the protein universe , author=. Proceedings of the National Academy of Sciences , volume=. 2009 , publisher=

  24. [33]

    2021 , publisher=

    Sillitoe, Ian and Bordin, Nicola and Dawson, Natalie and Waman, Vaishali P and Ashford, Paul and Scholes, Harry M and Pang, Chan and Woodridge, Laura and Rauer, Clemens and Sen, Neera and Abbasian, Maryam and LeCornu, Steven and Lam, Su Datt and Berka, Karel and Varekova, Radk...

  25. [34]

    Bioinformatics , volume=

    Chainsaw: protein domain segmentation with fully convolutional neural networks , author=. Bioinformatics , volume=. 2024 , publisher=

  26. [35]

    Bioinformatics , volume=

    How significant is a protein structure similarity with TM-score= 0.5? , author=. Bioinformatics , volume=. 2010 , publisher=

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.