{"id":"3531b413-5e33-42b2-bbc5-6756b6912dd7","arxiv_id":"2502.10173","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"VibeGen generates de novo protein sequences conditioned on a target normal-mode vibration profile, using a protein designer plus a screening predictor, and validates the resulting dynamics with all-atom normal mode analysis.","lead":"VibeGen uses two AI models, one that writes amino acid sequences and one that checks their predicted vibrations, to invent new proteins that match a chosen pattern of low-frequency motion. The authors validate designs with molecular simulations and report that many generated sequences are unlike any natural protein.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on CHARMM19/BNM normal modes being a valid proxy for real protein dynamics; without independent validation (explicit-solvent MD or experiment), the design-and-validate loop may be self-referential.","rationale":"I identify the validity of the CHARMM19/BNM proxy as the single load-bearing assumption. The paper's design, curation, and validation all use the same force field, minimizer, and BNM method, so the reported accuracy could be an artifact of the evaluation protocol rather than a property of the designed proteins. This is exactly the reader's weakest_assumption, so I agree. The missing random-sequence baseline is a real but secondary issue: it affects the interpretation of effect size, whereas the proxy question determines whether the headline claim is about physical dynamics or only about a computational phenotype. The proposed MD check is a decisive test because it uses an independent, physically more realistic model; if the correlation survives, the proxy concern is largely resolved, and if not, the paper's conclusions must be weakened. I recommend no change to the reader's CONDITIONAL verdict, since the reader already flagged this issue and the appropriate outcome is the same.","tokens_in":18426,"tokens_out":6068,"duration_ms":68261,"concrete_test":"Take 25–50 generated designs and the same number of source PDB proteins; for each, run explicit-solvent all-atom MD (CHARMM36m or AMBER ff14SB, 300 K, three 200 ns replicates) starting from the OmegaFold structure. Compute the per-residue Cα root-mean-square fluctuation (RMSF) and the first principal-component amplitude profile from the trajectories, and compare each to the prescribed normal mode shape smoothed as in Fig. S2. If the median Pearson correlation between MD-derived profiles and targets is far below the reported 0.72 smoothed BNM median, the validation loop is not transferable and the central claim should be narrowed to 'matches a CHARMM19/BNM proxy,' not real tailored dynamics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that VibeGen 'accurately reproduces prescribed normal mode amplitudes' and establishes a direct sequence-to-vibrational-behavior link. The target and the validation are both produced by the same pipeline: CHARMM19 all-atom energy minimization with implicit Gaussian solvent, Block Normal Mode analysis, with the first non-trivial mode's per-residue amplitude normalized by Eq. 2. Generated sequences are folded with OmegaFold, energy-minimized with the same CHARMM19 setup, and the resulting BNM shape is compared to the target. This loop is coherent, but it only demonstrates matching within this specific coarse NMA protocol. If CHARMM19/BNM low-frequency mode amplitude profiles are not representative of true conformational dynamics in aqueous solution, the designed proteins may not exhibit the tailored motions claimed. The paper itself notes experimental validation is needed (Section 3), and no independent force field, explicit-solvent MD, NMR, or B-factor comparison is provided. A second fragile link: structures are OmegaFold predictions with no reported confidence filtering; normal modes computed on poorly predicted structures would further detach the claims from physical reality.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces VibeGen, a generative framework for de novo protein design conditioned on low-frequency normal mode shapes. It curates a dataset of ~12,924 PDB protein chains (≤126 residues) by computing the lowest non-trivial normal mode amplitude profile (the \"normal mode shape vector\") using CHARMM19 energy minimization with an implicit Gaussian solvent and Block Normal Mode analysis. Two protein language diffusion models are trained: a protein designer (PD) that generates sequences from a target mode shape, and a protein predictor (PP) that predicts mode shapes from sequences. At inference, the PP screens candidate sequences from the PD to select the most accurate designs. The authors report median Pearson correlations of 0.53 (raw) and 0.72 (after low-pass filtering retaining the lowest 10% of FFT frequencies) between measured and target mode shapes across 1,293 test cases, plus evidence of diverse, often BLAST-novel sequences. The paper claims a direct, bidirectional link between sequence and vibrational behavior and positions VibeGen as a step toward dynamics-informed protein engineering.","tokens_in":18623,"tokens_out":3625,"duration_ms":40677,"significance":"If the central claim survives scrutiny, the paper makes a genuine contribution: it is, to my knowledge, one of the first demonstrations of end-to-end sequence generation conditioned directly on a protein dynamics signature, and the two-agent formulation (PD plus PP screening) is a sensible architectural choice. Strengths include public release of code and model weights, a newly curated normal-mode dataset, and a large-scale held-out evaluation with 1,293 designs. The reported diversity and de novo novelty of the generated sequences are notable. However, the significance is substantially tempered by the fact that the training labels, the design targets, and the validation metric are all produced by the same CHARMM19/BNM protocol, making the accuracy assessment self-referential with respect to that specific simulation model.","major_comments":[{"comment":"The training set and the validation protocol are both based on the same CHARMM19 energy function with implicit Gaussian solvent, energy minimization, and Block Normal Mode analysis. The design objective (input mode shape) and the measured output (mode shape of the generated sequence) therefore come from one and the same computational pipeline. The reported accuracy (median ρ = 0.53 raw, 0.72 after low-pass filtering) is a measure of self-consistency of that pipeline, not of transferability to physical protein dynamics in solution or in a test tube. The manuscript acknowledges the need for experimental validation only in the Conclusion, while the Abstract and Section 2 make the stronger claim of establishing a \"direct, bidirectional link between sequence and vibrational behavior.\" I recommend either adding an independent check on a subset (e.g., explicit-solvent MD, comparison with experimental B-factors or NMR S² order parameters) or softening the central claim to explicitly state that the mapping is within the CHARMM19/BNM representation.","section":"Fig. 5A-B and low-pass filter"},{"comment":"The headline accuracy numbers rely on a post-hoc low-pass filter that retains only the lowest 10% of FFT frequencies of the mode shape vectors. This cutoff is a free parameter, and the raw median Pearson correlation is only 0.53 (median relative L2 error 0.57). Because the paper claims to \"accurately reproduce the prescribed normal mode amplitudes across the backbone\" (Abstract), the unfiltered metric is the more direct test of that claim. The smoothed values are informative for large-scale shape matching, but the 10% cutoff must be justified independently of the observed improvement. Please report sensitivity to the cutoff value and present the raw and filtered distributions side by side for the same test cases.","section":"Fig. 7 and PP screening"},{"comment":"The claim that the protein predictor (PP) improves design accuracy is supported by comparing the PP-predicted-best and PP-predicted-worst designs among the PD's 40 candidates, showing median Pearson correlations of 0.53 versus 0.31. However, a random-selection baseline is missing. Even a weak ranking model would be expected to separate the extremes of a candidate pool. To quantify the actual benefit of the PP, the authors should report the median accuracy of all 40 candidates (or of a random subset), and ideally the expected median of the best/worst of 40 draws under random selection. Without such a baseline, the improvement cannot be attributed to the PP's predictive skill.","section":"§4 (Protein folding) and §2 (validation)"},{"comment":"All validation of generated sequences is performed on OmegaFold-predicted structures, with no confidence filtering or quality threshold reported. If OmegaFold produces inaccurate structures for some designs, the subsequent normal mode analysis on those structures may be meaningless, and the Pearson correlations would be degraded for reasons unrelated to the generative model. Please report the distribution of OmegaFold confidence scores (e.g., pLDDT) for the 1,293 generated sequences, and either restrict the accuracy analysis to high-confidence predictions or show that the results are insensitive to structure-prediction quality.","section":"§4 (Protein folding)"}],"minor_comments":[{"comment":"The Abstract states that validation is performed \"via full-atom molecular simulations,\" but the protocol is energy minimization followed by normal mode analysis; no MD trajectories are generated. Please rephrase to avoid overstatement.","section":"Abstract"},{"comment":"The caption labels panel (D) as the BLAST novelty distribution, but the text in Section 2 refers to \"Fig. 5D\" while the caption lists \"(F) shows the distribution\" — the panel label is inconsistent.","section":"Fig. 5 caption"},{"comment":"There are numerous typos and inconsistent abbreviations: \"animo acids,\" \"BLSAT,\" \"frequences,\" \"confirmational,\" \"dynamical shits,\" and the model name alternates between \"pLMD\" and \"pLDM.\" A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The text says cases B–E in Fig. 4 are all \"no significant similarity found,\" but Table 1 lists C as NSSF and B and D as NSSF; case E in the table has a match (68.25% identical with 3FZ9_A). Please reconcile the text with the table.","section":"Fig. 4 and Table 1"},{"comment":"De novo status is defined solely by \"no significant similarity found\" in a BLAST search against the nr database (or against PDB). This is sequence-level novelty; a generated sequence could be highly similar in structure to a natural protein despite low sequence identity. Consider reporting structural similarity (e.g., TM-score against the closest natural structure) to support the claim of \"de novo\" design.","section":"§2 (Novelty analysis)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a natural continuation of the Buehler group's prior work on ForceGen and protein normal-mode datasets, and the code and data release are commendable. My main reservation is the self-referential validation loop: the target shapes and the measured outputs are both produced by the identical CHARMM19/BNM protocol. This is a correctable weakness within the manuscript's scope if the authors either add an independent validation component (even a small subset with explicit-solvent MD or a comparison to experimental B-factors) or explicitly restrict the claims to the in-silico representation. The missing random-selection baseline for the PP screening and the unreported OmegaFold confidence filtering are additional load-bearing gaps. I would not reject the paper, but I cannot recommend acceptance without addressing these points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real, incremental step in dynamics-informed protein design, and the authors deserve credit for shipping code, model weights, and data. But the strongest claims need tempering: the reported 'accurate reproduction' of normal mode shapes is measured with the same coarse-grained NMA protocol used to generate the training labels, and the median raw Pearson correlation is 0.53, which is modest. The low-pass filtered 0.72 is presented fairly as capturing overall shape, but it is a post-hoc smoothing.\n\nWhat is new: conditioning a protein language diffusion model on the amplitude profile of the lowest non-trivial normal mode, with a second predictor agent to screen candidates. That extends the group's ForceGen work to a new design axis. The two-agent setup demonstrably improves selection: the predicted-best group has median Pearson 0.53 vs 0.31 for predicted-worst. The diversity results for the same target shape are genuinely interesting, and the BLAST analysis shows many sequences have no significant similarity to natural proteins.\n\nThe main soft spot is the evaluation loop. The target and the validation are both the CHARMM19/BNM normal mode amplitude shape. So the design-and-validate loop proves self-consistency within one simulation protocol, not that the designed proteins will show those motions in a test tube. The paper does acknowledge experimental validation is needed, but the abstract's 'direct validation' language oversells it. A second issue: the OmegaFold structures are used without confidence filtering; if a structure prediction is poor, the subsequent normal mode analysis is suspect.\n\nMinor: the PP screening lacks a random-selection baseline. Comparing best vs worst does not tell us how much better the PP is than picking a random candidate. The gap is suggestive, but a random baseline would make the benefit quantitative.\n\nOverall, this is a coherent, honest paper with reproducible artifacts. The central idea is plausible; the limitations are acknowledged and fixable. It is not a breakthrough, but it is a useful contribution to the growing area of dynamics-aware design. I would send it to review with requests for a random baseline, confidence filtering on OmegaFold, and better framing of the validation as protocol-specific. Who is this for? Researchers working on generative protein design and normal-mode-based dynamics. I would probably cite it.","headline":"Solid, reproducible step in dynamics-conditioned protein design, but the accuracy claims are weaker than the abstract suggests and the validation is self-referential to the simulation protocol.","tokens_in":19182,"tokens_out":2807,"would_cite":true,"duration_ms":29826,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-agent protein language diffusion model, VibeGen, generates de novo amino acid sequences whose low-frequency vibration profiles match prescribed targets.","keywords":["protein design","de novo proteins","normal mode analysis","protein dynamics","language diffusion model","vibrational mode shape","generative AI","sequence-structure-dynamics"],"falsifier":"Run explicit-solvent all-atom molecular dynamics on a set of VibeGen-designed proteins and compare the lowest non-trivial mode's per-residue amplitude profile to the prescribed target; if the median Pearson correlation falls near zero rather than around 0.53 to 0.72, the claimed sequence-to-vibration mapping would be an artifact of the validation force field.","tokens_in":18191,"feed_emoji":"🧬","tokens_out":5558,"duration_ms":55686,"temperature":0.7,"pith_summary":"VibeGen is a generative system that tries to solve the inverse problem of protein dynamics: given a desired pattern of low-frequency vibration along a protein's backbone, it produces amino acid sequences that, when folded, actually vibrate in that pattern. The paper claims this works end to end, without first designing a backbone structure, by training two complementary models—a designer that turns a vibration shape into sequences and a predictor that scores sequences by their predicted vibration shape. On 1,293 held-out targets, the measured vibration shapes of designed proteins correlate with the requested shapes with median Pearson $\\rho=0.53$, rising to $0.72$ when only the smooth, large-scale shape is compared. Many generated sequences have no significant BLAST match to natural proteins, so the designs reach sequence space beyond evolution. If correct, this establishes a direct, bidirectional map between sequence and vibrational behaviour that could be used to engineer flexible enzymes, dynamic scaffolds, and responsive biomaterials.","feed_headline":"VibeGen designs de novo proteins with prescribed vibrational profiles","feed_subtitle":"A two-agent diffusion model maps sequence to low-frequency motion; median correlation hits 0.72 on smoothed shapes.","key_machinery":"The load-bearing object is the normal mode shape vector, defined as the $\\mathbb{R}^N$ vector of C$\\alpha$ displacement amplitudes of the protein's lowest non-trivial normal mode, normalized so that $\\|\\vec{V}\\|=\\sqrt{N}$; it is a coordinate-invariant descriptor of the vibrational displacement distribution along the backbone. The generative machinery is a two-agent protein language diffusion model: a frozen 150M-parameter pretrained protein language model (ESM-2) embeds sequences, and a trainable 1D U-Net diffusion model performs denoising conditioned on the vibration target (designer) or on sequence representations (predictor). The designer proposes candidate sequences from a target shape; the predictor evaluates them on the fly, and the best-scoring candidates are validated by full-atom CHARMM19 energy minimization and Block Normal Mode analysis, the same protocol used to build the training dataset of 12,924 PDB chains.","core_discovery":"The central claim is that the lowest non-trivial normal mode shape of a protein—the per-residue amplitude profile $\\vec{V}=(d_1,\\dots,d_N)$ of the first non-rigid vibrational mode, normalized so that its L2 norm equals the sequence length $N$—can serve as a design condition for de novo protein generation. The paper reports that VibeGen, built from two protein language diffusion models, generates sequences whose measured normal mode shapes \"closely follow\" the prescribed targets, with a median Pearson correlation of 0.53 across 1,293 test cases (0.72 after low-pass filtering), and that many of the sequences are de novo by BLAST. It further claims that the predictor agent reliably ranks designs, so that selecting the predicted-best candidate from a batch of 40 significantly improves accuracy over the predicted-worst, and that the generated proteins fold into stable structures with secondary-structure motifs that plausibly explain the vibration pattern (helices and sheets suppress amplitude; loops and termini amplify it).","pith_inferences":["Beyond the paper: because the model uses only amplitude and drops directional information, the degeneracy it exploits may be even larger than reported; conditioning on full displacement vectors or on multiple modes could produce a richer family of designs per target shape.","Beyond the paper: a direct experimental check is available—pick a set of VibeGen designs, measure backbone dynamics by NMR spin relaxation or single-molecule FRET, and compare the measured flexibility profile to the prescribed normal-mode amplitude; the paper itself lists such validation as future work.","Beyond the paper: the same two-agent diffusion scheme could be transferred to other collective coordinates, such as mechanical unfolding force profiles or domain-interface motions, where a fast forward predictor can screen a generative inverse model."],"forward_implications":["Designing for dynamics becomes a direct sequence-level task: given any smooth target amplitude profile, the model can generate candidate sequences whose predicted vibration shapes match it, so dynamics can be combined with other sequence-level design objectives in one pipeline.","The two-agent screening scheme separates good from bad designs without running expensive physics for every candidate, since the predictor's ranking correlates with measured normal-mode accuracy.","De novo sequences with prescribed vibration profiles expand the searchable protein space beyond natural homologs, giving access to folds and motions that evolution may not have explored.","If the sequence-to-vibration map is real, it implies dynamics-conditioned design can be applied to functional properties known to depend on low-frequency motion, such as enzyme loop flexibility, allosteric coupling, and mechanosensitive response."],"supporting_citations":[{"why":"Supplies the full-atom MD and normal mode analysis protocol used to curate the training dataset and to validate the designed proteins.","marker":"66"},{"why":"Provides the protein language diffusion model architecture on which both the designer and predictor are built.","marker":"55"},{"why":"Supplies the frozen protein language model that carries sequence knowledge into the diffusion process.","marker":"30"},{"why":"Predicts 3D structures of the generated sequences, which are then used for normal mode analysis and secondary-structure assessment.","marker":"71"},{"why":"Used in the protein BLAST analysis to establish that many designed sequences have no significant similarity to known proteins.","marker":"72"},{"why":"Supplies the CHARMM all-atom force field used for energy minimization and Hessian-based normal mode calculations.","marker":"68"},{"why":"Provides the Block Normal Mode method that makes efficient normal mode analysis possible for large numbers of protein chains.","marker":"81"}],"fun_headline_variants":["VibeGen turns vibrational specs into new proteins","Agentic AI designs proteins to match vibrational targets","VibeGen: de novo proteins with prescribed motion profiles","VibeGen: AI for proteins that vibrate on demand","Two-agent diffusion model crafts de novo proteins with prescribed dynamics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole loop assumes that the lowest-frequency vibration computed in silico—using CHARMM19 with implicit Gaussian solvent on structures predicted by OmegaFold—is a faithful stand-in for how the protein would actually move in a test tube; if that proxy is wrong, the training and validation are self-referential and the reported accuracy would not transfer to experimental dynamics.","fun_headline_variants_meta":{"raw":{"variants":["VibeGen turns vibrational specs into new proteins","Agentic AI designs proteins to match vibrational targets","VibeGen: de novo proteins with prescribed motion profiles","VibeGen: AI for proteins that vibrate on demand","Two-agent diffusion model crafts de novo proteins with prescribed dynamics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001599,"raw_usage":{"total_tokens":6394,"prompt_tokens":989,"completion_tokens":5405,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":605,"completion_tokens_details":{"reasoning_tokens":5327}},"tokens_in":605,"tokens_out":5405,"duration_ms":37029,"temperature":1.0,"reasoning_tokens":5327,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T19:06:39.157023+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run explicit-solvent all-atom molecular dynamics on a set of VibeGen-designed proteins and compare the lowest non-trivial mode's per-residue amplitude profile to the prescribed target; if the median Pearson correlation falls near zero rather than around 0.53 to 0.72, the claimed sequence-to-vibration mapping would be an artifact of the validation force field.","supporting_citations":[],"review_version":1}