{"id":"7c134104-c411-4734-ab66-1578df58a511","arxiv_id":"2501.01477","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative survey of deep learning in protein bioinformatics that frames protein design as the inverse of structure and function prediction.","lead":"This paper surveys deep learning methods for proteins, grouping prior work into structure prediction, functional prediction, and protein design. It argues that progress in prediction is now the main engine for design, the hardest and most useful protein task.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's map is only as reliable as its second-hand method descriptions, and the already-documented citation misattributions in §3.1 and §4.2.2–4.2.3 make the central design-as-inverse synthesis unverifiable without a full reference audit.","rationale":"I read the paper as an expository survey, not a research contribution with original experiments. The central claim, that design can be viewed as the inverse of prediction and that prediction advances have contributed to design, is plausible and consistent with the broader field. The reader's verdict and evidence are accurate. The most load-bearing assumption is that the survey's second-hand descriptions map cleanly to the primary literature; the two documented citation errors confirm this assumption is not satisfied. My stress-test does not find a separate flaw in the inverse-design logic or the three-category taxonomy. It does find that the current text cannot be fully trusted as a reference until a citation-claim audit is done. Since the reader already assigned CONDITIONAL, I recommend UNCHANGED without altering the verdict. No ad hominem intended.","tokens_in":25904,"tokens_out":4660,"duration_ms":46192,"concrete_test":"Perform a complete citation-claim audit: for every method described in §§3–5, open the cited paper and verify (a) the cited work exists and (b) the described mechanism appears there. Start with the four design-as-inverse examples ([8] trRosetta hallucination, [81] trRosetta energy-landscape optimization, [110] trRosetta motif scaffolding, [50] Wasserstein GAN with structure oracle), then check the structural-prediction methods [117], [104], [119], [49], [92], and the structure-function claims in §4.2.2–4.2.3. If the misattribution rate exceeds the two already found, the survey map cannot support the central claim; if the audit is clean elsewhere, targeted corrections suffice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is not a quantitative theorem; it is an organizational claim that a reader can trust the survey's map and that design has benefited from predictors. That claim depends on the accuracy of dozens of second-hand method descriptions. This dependence is load-bearing, and it is already observably fragile: §4.2.2 and §4.2.3 attribute two different models, dMASIF and ScanNet, to the same reference [112]; the actual dMASIF paper (Sverrisson et al., 2021) is absent. §3.1 says 'Altschul in 1997 [4] explored using neural networks through SLPs for contact prediction', but [4] is the PSI-BLAST paper and contains no neural-network contact predictor; early contact prediction by neural networks is [68] (Lund et al., 1997). These are not stylistic choices: they prevent a reader from tracing the synthesis to sources. Since the paper contains no experiments, its only evidence is the accuracy of these descriptions. If other descriptions in §5 carry similar misattributions, the central assertion that prediction advances 'directly contributed' to design (supported mostly by hallucination papers [8], [81], [110], [50]) must be re-examined. The survey itself notes in §5.3 that AlphaFold2 'has yet to make its impact' on design, so the inverse-design claim is really about trRosetta-era oracles; the textual errors make it unclear which methods are actually being credited.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys deep learning methods in protein bioinformatics, organizing the field into three categories—structural prediction, functional prediction, and protein design—and argues that protein design can be viewed as the inverse of structural and functional prediction. It reviews historical and modern structure prediction methods (e.g., RaptorX, AlphaFold1/2, RoseTTAFold), functional prediction from sequence and structure (including GNNs and geometric deep learning), and design methods such as latent-space generation, GANs with structure oracles, directed evolution, and deep hallucination. The paper contains no new experiments; its contribution is a high-level synthesis plus a discussion of challenges (e.g., in silico versus in vitro validation) and future directions. The manuscript text indicates a submission date of October 2021, though it is posted on arXiv in January 2025.","tokens_in":26141,"tokens_out":4054,"duration_ms":40403,"significance":"If the survey's descriptions are accurate, it offers a useful introductory map of deep learning in protein bioinformatics and a clear articulation of the inverse-design viewpoint. The paper explicitly and repeatedly flags the gap between in silico and in vitro validation, it gives a reasonable high-level account of the trajectory from contact prediction to AlphaFold2, and it covers a broad range of methods including geometric deep learning and unsupervised language models. However, because the paper is a survey, its evidentiary value rests entirely on the correctness and completeness of its second-hand method descriptions. The manuscript's own text contains at least two concrete citation errors that affect the traceability of the synthesis, and it omits major post-2021 developments; these issues currently make the map less reliable than a survey should be.","major_comments":[{"comment":"The sentence \"Altschul in 1997 [4] explored using neural networks through SLPs for contact prediction\" is a factual misattribution. Reference [4] is the PSI-BLAST paper, which describes a sequence-search algorithm and contains no neural-network contact predictor. The correct early neural-network contact prediction work appears to be Lund et al. 1997, which the survey itself cites as [68] in §3. This error is load-bearing for the historical narrative in the structure-prediction section, because a reader cannot trace the claimed development of contact prediction from the cited sources. The passage should be corrected and the surrounding historical claims re-verified against their primary sources.","section":"§4.2.2–4.2.3"},{"comment":"Two different methods, dMASIF (spelled \"dMASIV\" in the text) and ScanNet, are both attributed to the same reference [112], which is the ScanNet preprint by Tubiana et al. The actual dMASIF paper (Sverrisson et al., 2021) is not cited anywhere. This makes it impossible for a reader to verify the descriptions of either method and casts doubt on the reliability of the \"function from structure\" review as a whole. The authors must add the correct reference for dMASIF and audit the surrounding subsections for similar citation errors.","section":"§4.2.2 and §4.2.3"},{"comment":"The central claim that advances in structural prediction have \"directly contributed\" to protein design is supported in the text almost exclusively by trRosetta-based hallucination works ([8], [81], [110]) and oracle-based generative methods ([38], [50]). Yet §5.3 itself states that AlphaFold2 \"has yet to make its impact in protein design.\" As written, the inverse-design claim is broader than the evidence presented. The authors should either restrict the claim to the specific prediction models actually used in the design methods they review (e.g., trRosetta-era oracles) or provide a more careful account of which structural-prediction advances have and have not been exploited in design, and why.","section":"§5.3"},{"comment":"The manuscript is dated October 2021 but posted in January 2025, and its coverage appears to end around 2021. It does not discuss major post-2021 developments such as RFdiffusion, ProteinMPNN/InverseFolding, ESMFold, AlphaFold3, or the widespread use of diffusion models in design. For a survey whose stated purpose is to map the current state of deep learning in protein bioinformatics, the omission of these widely used methods is a load-bearing incompleteness. The authors should either update the survey to cover the 2022–2024 literature or clearly restrict the claimed scope and title to a historical snapshot.","section":"Overall"}],"minor_comments":[{"comment":"There are numerous typographical errors, including \"outperfrom,\" \"convoluational,\" \"dMASIV,\" \"disearable,\" \"interporlate,\" \"baysian,\" \"hyrdophobic,\" \"millisconds,\" \"peer reviewewd,\" and inconsistent capitalization such as \"Alphafold2\" versus \"AlphaFold2.\" A careful proofreading pass is needed.","section":"Throughout"},{"comment":"Several figures are reproduced from external sources ([5], [16], [31], [86], [50]) without explicit permission statements, and the provenance is given only in captions. The authors should confirm that permission or license terms are in order for a published survey.","section":"Figures"},{"comment":"The text ends with a list of bare blog URLs that are not integrated into the reference list or cited in the body. These should either be incorporated as formal references with access dates or removed.","section":"End matter"},{"comment":"The claim that DeepGO \"surpassed GOLabeler and NetGO\" is stated without a benchmark or time qualification, while the later discussion of TALE says it achieved state-of-the-art results for two of three subclasses. The authors should make the comparison consistent and specify which CAFA evaluation is being referenced.","section":"§4.1.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be a 2021 submission that was only posted to arXiv in January 2025. The editor may wish to clarify the submission timeline, since the text has not been updated to reflect the past three years of literature. The citation errors in §3.1 and §4.2 are individually fixable, but they indicate that a full citation audit of the entire survey is necessary before the paper can be relied upon as a map of the field. If the authors complete that audit and update the scope, the organizational narrative could be a useful contribution to the q-bio.BM literature; in its current form it is not yet dependable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a competent introductory survey and nothing more. It introduces no new method, dataset, or result; the three-category taxonomy is explicitly inherited from the standard paradigm, and the design-as-inverse-prediction framing comes from the trRosetta hallucination papers it cites. That is not a flaw for a review, but the paper's value is entirely as a map for newcomers.\n\nWhat it does well: the structure-prediction narrative (RaptorX, AlphaFold1/2, RoseTTAFold, MSA Transformer) is accurate at a high level; the repeated caution about in silico versus in vitro validation in design is honest and appropriately stressed; and the design section covers a representative slice of the 2021 landscape—hallucination, GANs with structure oracles, autoregressive generation, directed evolution. The prose is readable and the figures are well chosen.\n\nSoft spots, in proportion: the citation errors are real and matter for a survey. Section 3.1 attributes neural-network contact prediction to Altschul 1997 [4], which is the PSI-BLAST paper; the correct early neural work (Lund 1997 [68]) is cited two paragraphs earlier, so this is careless but not conceptually deep. More annoying: Sections 4.2.2 and 4.2.3 assign two different methods, dMASIF and ScanNet, to the same reference [112]; dMASIF is actually Sverrisson et al. 2021, missing from the bibliography. That is the kind of error that makes a survey a bad pointer for people who need to follow the thread. The stress-test note worries that these errors make the central design-as-inverse synthesis unverifiable. I disagree. The design sections cite Anishchenko, Norn, Tischer, and Karimi correctly, and those are the papers that actually carry the inverse-design argument. The misattributions sit in the functional-prediction background, so they erode trust without toppling the thesis.\n\nAlso worth noting: the text is a 2021 submission posted without visible update in 2025. The claim that AlphaFold2 has yet to impact design was fair in late 2021 but is now stale. A revision should either update the design section or explicitly date the survey's coverage.\n\nFor whom: a graduate student who wants the lay of the land in deep-learning protein bioinformatics, not a researcher looking for new insight. It deserves a serious referee, but only with the condition that the reference list is audited and the scope is dated. I would not cite it until then.","headline":"A useful but dated survey whose citation errors are fixable and don't sink the design-as-inverse thesis; send it to review with an order to audit the references.","tokens_in":26720,"tokens_out":3491,"would_cite":false,"duration_ms":32970,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that protein design is the inverse of structural and functional prediction, and that recent deep-learning advances in prediction have become the main engine for designing new proteins.","keywords":["deep learning","protein bioinformatics","protein structure prediction","protein function prediction","protein design","deep hallucination","geometric deep learning","generative models"],"falsifier":"Design 100 de novo proteins by hallucinating a state-of-the-art structure predictor and an equal number by physics-based energy minimization, synthesize all 200, and compare in vitro folding rates; if the hallucinated set does not fold more often, or at least as often, the survey's claim that stronger predictors directly drive design loses its load-bearing evidence.","tokens_in":25636,"feed_emoji":"🧬","tokens_out":5680,"duration_ms":52714,"temperature":0.7,"pith_summary":"The paper is a survey, and its contribution is an organizing claim: protein design is best understood as the inverse of structural and functional prediction, and the recent deep-learning successes in prediction are what have made design tractable. It reviews structure prediction from sequence, functional prediction from sequence and structure, and design as sequence generation from function or structure, with the three tasks linked in a single cycle. A sympathetic reader should care because the claim reframes design as a problem that improves automatically as predictors improve, and it explains why methods that work for AlphaFold2-style models—especially back-propagating through a trained network—can be recycled as design engines. The survey also stresses that experimental validation remains the bottleneck that separates plausible designs from real proteins.","feed_headline":"Protein design is prediction run backwards","feed_subtitle":"A field survey shows how AlphaFold-era predictors double as design engines for new protein sequences.","key_machinery":"The mechanism that carries the argument is reverse-mode gradient descent on trained prediction networks—the 'deep hallucination' procedure used by several papers it surveys. A structure predictor such as trRosetta maps a sequence to a distribution over inter-residue distances and orientations; by freezing the network's weights and back-propagating the gradient of a structural objective into a randomly initialized sequence, the sequence itself becomes the optimizable variable. This inversion is what allows design to be described as prediction run backwards, and it is the concrete object the survey points to when it says structural and functional prediction advances have directly contributed to design tasks.","core_discovery":"The paper's central assertion is that the three canonical problems of protein bioinformatics form a directed cycle: sequence determines structure, structure and sequence determine function, and design is the inverse map from function or structure back to sequence. Within that frame, the paper claims that the strongest recent design results did not come from better energy functions or fragment libraries but from reusing deep-learning predictors—most notably trRosetta and AlphaFold2—either as scoring oracles in generative models or as differentiable networks that can be run backward by gradient ascent to 'hallucinate' sequences for a desired structure or motif. It treats this inversion as the main explanation for why design methods have improved since CASP13, while acknowledging that in silico performance has repeatedly failed to survive in vitro synthesis. The survey also claims the structure-first, function-second ordering is useful because functional labels are sparser than structural data, making function the weakest link in the cycle.","pith_inferences":["Editorial inference: the same back-propagation trick should in principle work on any differentiable function predictor, not just structure predictors, so protein design could be targeted at activity or binding directly if such predictors reach sufficient accuracy.","Editorial inference: a direct controlled comparison of hallucination-based design against physics-based Rosetta design on the same test structures would isolate how much of the recent design progress is due to prediction accuracy rather than to the search procedure.","Editorial inference: the survey's three-way map suggests a testable prediction: as AlphaFold2-class models are adopted as oracles, reported in vitro success rates of de novo designs should rise; if they stay flat, the causal link from prediction to design would be weaker than claimed."],"forward_implications":["If design is the inverse of prediction, then any sustained improvement in structure or function prediction should translate into better design methods without new design-specific ideas.","Structure predictors can serve as oracle feedback inside generative models such as GANs, reinforcement learning, and directed evolution, steering generated sequences toward stable folds.","Inverse back-propagation, or hallucination, turns a trained predictor into a sequence generator, so design no longer requires an explicit energy function or fragment library.","Because in silico validation alone has repeatedly failed to predict in vitro folding, the field's rate of progress will be capped by experimental synthesis throughput, not just model accuracy.","Benchmarks for design should move toward in vitro fold-success rates rather than sequence-recovery or reconstruction scores."],"supporting_citations":[{"why":"Supplies the survey's central example: AlphaFold2 achieved near-experimental structure prediction accuracy.","marker":"[49]"},{"why":"Provides the trRosetta predictor whose distance and orientation outputs are inverted by hallucination methods.","marker":"[119]"},{"why":"Demonstrates de novo design by back-propagating through trRosetta to hallucinate new sequences, the paper's key example of design as inverse prediction.","marker":"[8]"},{"why":"Extends hallucination to proteins with specified binding motifs, showing inverse prediction can target function.","marker":"[110]"},{"why":"Applies reverse-mode optimization to trRosetta with a Bayesian prior for higher-probability sequence design.","marker":"[81]"},{"why":"Establishes structure-based functional prediction with geometric deep learning, underpinning the function-from-structure category.","marker":"[31]"}],"fun_headline_variants":["Running prediction backward to design proteins","Design is prediction in reverse: a survey","Inverting deep learning for protein design","Predictors become designers in protein survey"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's central map is only as reliable as its second-hand descriptions of dozens of primary papers; those descriptions already contain documented errors, including attributing two distinct methods (dMASIF and ScanNet) to the same reference and crediting an early neural-network contact predictor to the PSI-BLAST paper.","fun_headline_variants_meta":{"raw":{"variants":["Running prediction backward to design proteins","Design is prediction in reverse: a survey","Inverting deep learning for protein design","Predictors become designers in protein survey"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1312,"prompt_tokens":937,"completion_tokens":375,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":324}},"tokens_in":553,"tokens_out":375,"duration_ms":4523,"temperature":1.0,"reasoning_tokens":324,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:35:29.511504+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Design 100 de novo proteins by hallucinating a state-of-the-art structure predictor and an equal number by physics-based energy minimization, synthesize all 200, and compare in vitro folding rates; if the hallucinated set does not fold more often, or at least as often, the survey's claim that stronger predictors directly drive design loses its load-bearing evidence.","supporting_citations":[{"cited_title":"Design of proteins presenting discontinuous functional sites using deep learning","cited_arxiv_id":null,"evidence_quote":"Extends hallucination to proteins with specified binding motifs, showing inverse prediction can target function."},{"cited_title":"Protein sequence design by explicit energy landscape optimization","cited_arxiv_id":null,"evidence_quote":"Applies reverse-mode optimization to trRosetta with a Bayesian prior for higher-probability sequence design."}],"review_version":1}