{"id":"e66bd9fc-f0a9-4681-b9ad-3ace664300ce","arxiv_id":"2504.13075","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"APM generates multi-chain protein complexes at all-atom resolution via a three-module flow-matching design, achieving strong computed binding affinities in antibody and peptide design benchmarks.","lead":"APM is a new generative model that creates multi-chain protein complexes with full atomic detail, including amino acid sequence, backbone, and sidechains. It reports strong in-silico binding energies for antibodies and peptides, and the authors have released the code.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 4's central comparison may be confounded: fixed chain-length regimes rather than per-complex native lengths, and Rosetta ΔG after relaxation is the sole readout, so the 'binding capability' claim rests on a proxy the paper itself concedes needs wet-lab validation.","rationale":"The reader's weakest assumption is exactly the Rosetta ΔG proxy concern, and I agree with it. My added specificity is that Table 4's comparison is also confounded by length regime (fixed two-chain lengths, no per-complex native length matching), and that the abstract's 'binding capabilities' step is stronger than the paper's own cautious phrasing in Section F. Given the paper's otherwise strong engineering contributions—an all-atom representation, an integrated three-module architecture, released code, and solid in-silico baseline performance on folding and inverse-folding—the right verdict is CONDITIONAL: the central claim is plausible but not yet demonstrated as binding capability in an experimentally meaningful sense. A relaxation-independent readout and per-sample error bars would substantially de-risk the claim.","tokens_in":28569,"tokens_out":1740,"duration_ms":16075,"concrete_test":"Re-run the Table 4 generation and evaluation for a single length combination (100-100) with all other settings fixed, replacing the Rosetta ΔG readout with a relaxation-independent binding quality metric, and report per-sample error bars across at least 10 independent seeds. Concretely: (a) compute interface buried surface area, interface shape complementarity, or AlphaFold-multimer ipTM on the designed sequences; (b) compare APM versus Chroma* on that metric; (c) check whether the Rosetta ΔG ranking of individual designs correlates with the relaxation-independent metric. If the ΔG gap persists with non-overlapping distributions and a relaxation-independent metric also favors APM, the proxy concern is materially reduced.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that APM 'designs protein complexes with binding capabilities from scratch' rests primarily on Table 4, where APM reports a mean Rosetta all-atom binding energy ΔGRAA of -130.31 for 100-100 complexes versus -62.15 for Chroma redesigned with ProteinMPNN. Two load-bearing issues weaken this comparison. First, the evaluation is not controlled for design difficulty: APM generates complexes at fixed length combinations (50-100, 100-100, 100-200), whereas the difficulty of achieving a favorable Rosetta energy depends on chain lengths and composition; a length-matched control across a wider range of native complex geometries would be needed to attribute the energy gap to the generative model rather than to the sampled length regime. Second, and more fundamentally, every readout in Table 4 and in the binder-design experiments of Section F relies on pyRosetta relaxation and ΔG as a proxy for binding. The authors explicitly state in Section F: 'the actual effectiveness still requires validation through wet lab experiments.' As stated, the abstract's claim of 'binding capabilities from scratch' is therefore only as strong as this proxy. The paper contains no wet-lab or affinity measurement (e.g., ITC, SPR, or co-purification), so the central claim reduces to: APM generates complexes that score favorably under one physics-based energy function after relaxation. This proxy assumption is load-bearing for the headline claim, and no independent evidence is provided that would bridge proxy energetics to actual binding capability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents APM, an all-atom generative model for multi-chain protein complexes. The model is composed of three modules: a sequence/backbone flow-matching module, a sidechain prediction module, and a refinement module, trained in two phases with a mixture of single-chain and multi-chain protein data and ESM2 embeddings. The authors report benchmarks on single-chain folding and inverse-folding, multi-chain folding and inverse-folding, unconditional multi-chain complex generation, antibody CDR-H3 co-design on the RAbD benchmark, peptide design on PepBench/LNR, and zero-shot binder design against several targets. The headline claim is that APM designs protein complexes with binding capabilities from scratch, supported mainly by pyRosetta binding-energy calculations and by Boltz-1 confidence and DockQ metrics.","tokens_in":28889,"tokens_out":6028,"duration_ms":52240,"significance":"APM addresses a timely gap in multi-chain all-atom protein generation. If the reported results are reliable, the model is a useful contribution: it natively supports multi-chain generation without pseudo-linkers, integrates sidechain torsion angles for inter-chain modeling, and achieves competitive or superior scores on standard antibody and peptide benchmarks. The authors release code, and the two-phase training scheme with a consistency loss and a refine module is clearly described. However, the paper's strongest claim, that APM designs complexes 'with binding capabilities from scratch,' is only supported by computational proxies (Rosetta ΔG, pLDDT/ipTM), which the authors themselves note require wet-lab validation; the headline should be scaled back accordingly.","major_comments":[{"comment":"The central claim of the abstract ('designing protein complexes with binding capabilities from scratch') is supported only by pyRosetta relaxation and ΔG calculations in Table 4 and by pLDDT/ipTM confidence scores in Section F. The authors explicitly concede in Section F that 'the actual effectiveness still requires validation through wet lab experiments.' No ITC, SPR, co-purification, or other affinity measurement is reported. As written, the headline claim is load-bearing and stronger than the evidence: the results establish that APM generates complexes with favorable Rosetta energies under the chosen relaxation protocol, not that it designs binders. I recommend either adding experimental validation or qualifying the abstract and Section 4.3.2 claims as computational binding-affinity predictions.","section":"§4.3.2, Table 4; §F"},{"comment":"The claim that APM's peptide designs show 'significantly outperforming other methods' on functionality is not supported by Table 5. In that table, RFDiffusion achieves a better mean ΔG (-23.27 vs -19.90 for APMSFT) and a higher fraction of favorable complexes (%<0 = 78.58 vs 69.34). RFDiffusion also has higher pLDDT (69.65 vs 60.36), ipTM (0.73 vs 0.66), and Success (46.28% vs 29.22%). APM's advantage is limited to DockQ and the high-quality DockQ fraction. The text should be revised to report these performance gaps accurately and to describe APM as competitive on functionality and foldability rather than superior.","section":"§4.4.2, Table 5"},{"comment":"The comparison for binding energies uses only three fixed length combinations (50-100, 100-100, 100-200), and the reported averages and medians are presented without variance, confidence intervals, or statistical tests. Because the difficulty of achieving favorable Rosetta energies depends strongly on chain lengths and composition, a near-constant length regime cannot establish a general claim of superior inter-chain modeling. I ask for seed-level variability, analogous to Table 7 in Appendix D.2, and, if feasible, a benchmark on a broader set of native complex geometries with matched chain-length distributions.","section":"§4.3.2, Table 4"},{"comment":"The sentence 'this proves the importance of the all-atom information in the inter-chain interactions modeling' is too strong for the reported ablation. APMBB differs from full APM not only in the absence of all-atom information in the backbone generation path, but also in the absence of the SidechainModule and RefineModule entirely, including their sequence-level corrections. The observed energy gap could be due to the refinement module or to differences in the sampling schedule rather than to the sidechain torsions per se. A cleaner ablation would feed the same Seq&BBModule predictions through the Sidechain/Refine modules with and without the torsion features, or otherwise isolate the information channel.","section":"§4.3.2, APMBB ablation"}],"minor_comments":[{"comment":"The APM entry for 100-100 ΔGRAA reads '-130.31-134.57' and should be '-130.31/-134.57'; this formatting issue appears in at least one other table row.","section":"Table 4"},{"comment":"The text should state more explicitly that the multi-chain folding comparison favors APM only against Boltz-1 without MSA; the gap to Boltz-1 with MSA (RMSD 12.6 vs 5.40, TM 0.64 vs 0.87) is large and deserves a clearer caveat in the main discussion.","section":"Table 3, §4.3.1"},{"comment":"The word 'resdiue' should be 'residue'.","section":"§A.3"},{"comment":"The consistency loss in Eq. (19) uses tS and tT before the notation for the decoupled noising times is formally defined; a brief definition would improve readability.","section":"§3.3.1, Eq. (19)"},{"comment":"The sequence sampling temperature schedule in Eq. (24) uses hyperparameters Tmax and λ whose numerical values appear only in Appendix D.1; stating them in the main text would aid reproducibility.","section":"§3.4, Eq. (24)"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the ICML audience, and the code release is valuable. My main concern is that the abstract's 'binding capabilities from scratch' claim outruns the evidence; please encourage the authors to soften the claim or add experimental support. The misstatement about RFDiffusion in Table 5 should be fixed before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real contribution—a three-module all-atom generative model that natively handles multi-chain complexes, with code out and a lot of careful engineering detail. The reader's conditional verdict and the stress-test note are about right, and I'd send it to review.\n\nWhat's actually new: APM is, as far as I can tell, one of the first systems to generate multi-chain complexes end-to-end with explicit sidechain torsions, using flow matching for sequence+backbone, a separate sidechain module, and an all-atom refine module. The decoupled noising and two-phase training are sensible original choices, and the ablation APM vs APMBB in Table 4 does show the all-atom modules matter for interface quality. The multi-chain inverse-folding results (AAR 61% vs ProteinMPNN 46%) are strong, and the antibody CDR-H3 numbers look competitive. The paper is honest about its own limitations—folding is weak, RefineModule is limited, multi-chain generation degrades beyond two chains—and the appendix is thorough.\n\nSoft spots, in proportion:\n\n1. The central claim in the abstract, 'designing protein complexes with binding capabilities from scratch,' rests entirely on Rosetta ΔG values after relaxation. The authors say in Section F that wet-lab validation is still needed. That is exactly right, and it means the headline claim should be read as 'favorable Rosetta binding energies,' not demonstrated binding. The reader's stress-test note on this is fair. It's a standard proxy in the field, but it is load-bearing and unverified.\n\n2. The text overstates one table. In peptide design (Table 5), the paper says APM 'significantly outperforms other methods' on functionality, but RFDiffusion has better mean ΔG (-23.27 vs -19.90) and a higher fraction below zero (78.6% vs 69.3%). That sentence should be corrected or qualified; it's a small thing but it undercuts trust.\n\n3. The multi-chain generation comparison is thin: only Chroma (+ProteinMPNN) as baseline, no error bars or significance tests, and fixed length combinations that may not reflect native complex difficulty. This is addressable with more baselines (MultiFlow, RFDiffusion) and per-complex statistics.\n\nNone of these are fatal. The architecture is described precisely enough to reproduce, code is released, and the honest tone gives me confidence the authors are not hiding the load-bearing assumption.\n\nWho this is for: computational protein designers and anyone tracking generative models for complexes. It deserves a serious referee; I'd accept it and ask for claim softening, error bars, and one more baseline. I'd probably cite it as a point of comparison in my own work.","headline":"Genuinely new multi-chain all-atom generative model with honest limitations, but the 'binding capability' claim rests on Rosetta energies and one SOTA sentence is undercut by its own peptide table.","tokens_in":29436,"tokens_out":4468,"would_cite":true,"duration_ms":37562,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"APM generates protein complexes from scratch, with all-atom sidechain modeling producing roughly twice the computed binding energy of a Chroma-based baseline.","keywords":["multi-chain protein generation","all-atom protein design","flow matching","sidechain torsion angles","protein complexes","antibody design","peptide design","binding energy"],"falsifier":"Synthesize genes for a set of APM-generated two-chain complexes spanning the reported ΔG range, express and purify the chains, and measure dissociation constants by surface plasmon resonance or isothermal titration calorimetry; if the measured affinities do not correlate with the reported Rosetta ΔG values, the claim that APM designs binding-capable complexes is unsupported.","tokens_in":28376,"feed_emoji":"🧬","tokens_out":9127,"duration_ms":83043,"temperature":0.7,"pith_summary":"APM is a generative model built to do what most protein foundation models do not: design multi-chain complexes, not just single chains, and output all-atom structures rather than backbones that need separate sidechain packing. The authors' central claim is that adding sidechain torsion angles to a residue-level flow-matching model is what makes inter-chain interactions learnable, and that this lets APM generate two-chain complexes from scratch with markedly more favorable computed binding energies than the Chroma baseline. In the headline comparison, 100-residue-plus-100-residue complexes generated by APM have a mean all-atom relaxed binding energy of -130.31 Rosetta energy units, versus -62.15 for Chroma redesigned with ProteinMPNN. The model also performs multi-chain folding and inverse folding and, after supervised fine-tuning or even zero-shot, improves antibody CDR-H3 and peptide design metrics. The authors are explicit that the actual binding effectiveness still requires wet-lab validation, so the claim is about computed binding strength until experimental confirmation.","feed_headline":"One model designs protein complexes atom-by-atom from scratch","feed_subtitle":"APM co-generates sequence, backbone, and sidechains, reporting binding energies about twice as favorable as Chroma","key_machinery":"The load-bearing object is the all-atom residue representation — amino acid type, backbone frame in $\\mathrm{SE}(3)$, and sidechain torsion angles $\\chi \\in [0,2\\pi)^4$ — plus the three-module architecture that uses it. Module one (Seq&BB) is a flow-matching generator over discrete sequence tokens and $\\mathrm{SE}(3)$ backbone frames, trained with decoupled noising of the two modalities plus a consistency loss; module two (Sidechain) is a one-step packer that predicts $\\chi$ from the generated sequence and backbone; module three (Refine) takes the predicted all-atom structure and corrects sequence and backbone before the next denoising step. The integrated design matters because sidechain prediction needs clean sequences and structures, so the sidechain cannot simply be another flow-matching head without amino-acid-type leakage; instead it is activated only late in sampling ($t \\ge 0.8$), and the Refine module lets all-atom information feed back into the backbone and sequence. A protein language model supplies sequence understanding to all three modules.","core_discovery":"On its own terms, APM's discovery is that a protein complex can be generated natively as a joint object — sequence, backbone, and sidechain conformations together — rather than as separate chains stitched together or as a backbone that is later packed. The model represents each residue by amino-acid type, an $\\mathrm{SE}(3)$ backbone frame, and up to four sidechain torsion angles $\\chi$, and it generates the sequence and backbone with flow matching while a dedicated sidechain module predicts $\\chi$ and a refinement module re-optimizes the whole all-atom structure during the last part of sampling. The authors argue that the all-atom loop is doing real work: ablating it (APMBB, residue-level only) lowers computed binding strength and raises interface RMSD, and the sidechain torsion angles carry amino-acid-type information that would leak if the sidechain were noised jointly. In multi-chain inverse folding, APM's amino acid recovery is 61.26% versus 46.17% for ProteinMPNN, and in the downstream tasks APM's fine-tuned antibody designs have lower total and binding energies than the compared methods while its peptides are the only ones producing a meaningful share of high-quality DockQ designs.","pith_inferences":["If the Rosetta proxy survives wet-lab testing, the natural next step is to measure the correlation between ΔG rankings and experimental affinities; a positive correlation would make the model's energies usable as a screening prior for binder leads.","The three-module split suggests a general recipe for discrete-continuous co-generation: keep the sidechain as a one-step conditional predictor rather than a noised channel, and let refinement feed all-atom information back into the backbone. The same pattern could transfer to other joint design problems, though the paper does not claim this.","The chain-by-chain sampling mode produces weaker interfaces and visible clashes, which suggests that simultaneous generation is important for interface complementarity; the paper reports this behavior but does not elevate it to a design principle.","Zero-shot antibody designs with low binding energy but unnatural CDR-H3 shapes imply that the model's binding prior is generic rather than antibody-specific, and that targeted fine-tuning changes the binding mode rather than merely improving affinity."],"forward_implications":["Designing a complex no longer has to be staged as separate backbone generation, sequence design, and sidechain packing: APM outputs all three together, and its ablation suggests the sidechain loop is essential to interface quality.","Complex generation with two chains of length 100-100 yields computed binding energies roughly double those of the Chroma-plus-ProteinMPNN pipeline, indicating the all-atom co-generation path is competitive at directly producing tightly bound interfaces.","Multi-chain inverse folding works: APM's 61.26% amino acid recovery on the multi-chain test set exceeds the single-chain-oriented ProteinMPNN baseline, so sequences of existing complexes can be redesigned while preserving structure.","Supervised fine-tuning turns the general model into a specialist: antibody CDR-H3 co-design and peptide design both improve, while zero-shot sampling still produces low computed binding energies, making the same checkpoint usable in both modes.","Longer binder design against targets such as IL-7RA, PD-1, and TNF-α is reachable zero-shot, with computed binding energies comparable to the RFdiffusion baseline across six of seven targets."],"supporting_citations":[{"why":"Supplies the AlphaFold2 structure-module trunk, the sidechain-torsion representation, and the FAPE-style losses that all three APM modules build on.","marker":"Jumper et al., 2021"},{"why":"Provides the flow-matching framework used by the Seq&BB module for continuous generation of backbone structures.","marker":"Lipman et al., 2023"},{"why":"Provides the discrete flow-matching formulation for sequence tokens and the MultiFlow baseline and data split used in single-chain comparisons.","marker":"Campbell et al., 2024"},{"why":"Supplies the ESM2 protein language model integrated into all APM modules for sequence understanding, and the ESMFold baseline for folding comparisons.","marker":"Lin et al., 2023"},{"why":"Provides ProteinMPNN, the inverse-folding baseline and the sequence-redesign baseline used in complex generation, peptide design, and binder design.","marker":"Dauparas et al., 2022"},{"why":"Provides Chroma, the unconditional complex-generation baseline against which APM's computed binding energies in Table 4 are compared.","marker":"Ingraham et al., 2023"},{"why":"Defines the Rosetta all-atom energy function used through pyRosetta for relaxation and ΔG scoring, which is the central evaluation proxy for binding.","marker":"Alford et al., 2017"},{"why":"Provides Boltz-1, the multi-chain folding baseline and the folding model used to assess foldability and interface confidence of designed peptides and binders.","marker":"Wohlwend et al., 2024"},{"why":"Provides RFdiffusion, the diffusion baseline for peptide and long-binder design, typically paired with ProteinMPNN for sequence design.","marker":"Watson et al., 2023"}],"fun_headline_variants":["APM co-designs protein complexes atom-by-atom","All-atom generative model builds protein complexes from scratch","Protein complex design: all-atom generation in one model","One model, all atoms: protein complexes from scratch"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that Rosetta's computed binding energy after relaxation is a trustworthy stand-in for real binding: if those energies do not predict what happens in a wet-lab binding assay, the paper's central claim about 'binding capabilities' reduces to a claim about favorable simulation scores.","fun_headline_variants_meta":{"raw":{"variants":["APM co-designs protein complexes atom-by-atom","All-atom generative model builds protein complexes from scratch","Protein complex design: all-atom generation in one model","One model, all atoms: protein complexes from scratch"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000927,"raw_usage":{"total_tokens":3985,"prompt_tokens":972,"completion_tokens":3013,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":2948}},"tokens_in":588,"tokens_out":3013,"duration_ms":20405,"temperature":1.0,"reasoning_tokens":2948,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:15:40.701582+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Synthesize genes for a set of APM-generated two-chain complexes spanning the reported ΔG range, express and purify the chains, and measure dissociation constants by surface plasmon resonance or isothermal titration calorimetry; if the measured affinities do not correlate with the reported Rosetta ΔG values, the claim that APM designs binding-capable complexes is unsupported.","supporting_citations":[{"cited_title":"Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design","cited_arxiv_id":null,"evidence_quote":"Provides the discrete flow-matching formulation for sequence tokens and the MultiFlow baseline and data split used in single-chain comparisons."},{"cited_title":"Evolutionary-scale prediction of atomic-level protein structure with a language model","cited_arxiv_id":null,"evidence_quote":"Supplies the ESM2 protein language model integrated into all APM modules for sequence understanding, and the ESMFold baseline for folding comparisons."},{"cited_title":"Boltz-1: Democratizing biomolecular interaction modeling","cited_arxiv_id":null,"evidence_quote":"Provides Boltz-1, the multi-chain folding baseline and the folding model used to assess foldability and interface confidence of designed peptides and binders."}],"review_version":1}