{"id":"2694664c-03c0-4ad9-9450-d348946e401a","arxiv_id":"2506.01995","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A STAR-based computational pipeline on bulk and single-cell mouse antibody data identified 67 SPR-validated anti-GFRAL antibodies, with 19 of 40 STAR hits confirmed as binders.","lead":"This paper combines computational screening of antibody repertoires with lab binding tests to find 67 new antibodies against GFRAL, a receptor linked to appetite and weight control. It reports that the STAR clustering method, when followed by single-cell pairing, yields a 50 percent success rate in identifying binding antibodies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 50% STAR success rate is confounded by single-cell presence; the paper's own control from the same STAR clusters absent from single-cell data yields only 19% binding, so the central efficiency claim lacks a matched baseline.","rationale":"The paper's strongest evidence for its central methodological claim is the 50% success rate of STAR-derived candidates relative to a 20% random baseline. But that comparison is not matched: STAR candidates were required to appear in the single-cell data, while the random baseline was drawn from single-cell sequences without any bulk-cluster requirement. The paper itself supplies a nearly perfect control: heavy chains from the same STAR clusters that were absent from single-cell data bound at only 19%, essentially the random rate. That control strongly suggests that the STAR cluster signal alone has little predictive power for binding, and that the higher success rate of the 40 hits is explained by their single-cell presence or frequency. If the 'rapid discovery' claim rests on STAR enriching candidate selection, this evidence undercuts it. The catalog of 67 SPR-validated antibodies remains a useful resource, but the claimed pipeline advantage is not supported by the presented comparisons. This is therefore a condition for acceptance rather than a reason to reject outright. The reader named light-chain pairing as the weakest assumption; I regard the single-cell-presence confound as more load-bearing, although the reader's rationale also mentions the confound, so our assessments partially align.","tokens_in":14362,"tokens_out":5976,"duration_ms":64134,"concrete_test":"Recompute the binding success rate of the 40 STAR hits versus the 20 random single-cell controls after matching for single-cell abundance (e.g., restrict both sets to sequences with single-cell count 1–5, then 6–20, then >20). If the STAR advantage over random controls disappears in these matched strata, the 50% claim is an artifact of single-cell presence or frequency rather than a property of the STAR signal.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing weakness is in attributing the headline success rate to STAR. In Section II.B, all 40 STAR hits were required to be present in the single-cell dataset; 19 of these 40 (50%) bound by SPR. The only control that isolates the STAR clustering signal—reconstructed antibodies from the same STAR clusters whose heavy chains were absent from single-cell data—yielded 19% binders (Section II.B, criterion 3), indistinguishable from the 20% success rate of randomly selected single-cell antibodies. This indicates that the apparent enrichment of STAR hits may be entirely due to their presence or frequency in single-cell data, not to the STAR cluster signal. The paper's statement that finding 19 binders among 40 has 'very low' probability is therefore not established without a frequency-matched or presence-matched control. Additionally, Section II.D reports 70 binders and 67 non-binders among the 137 modeled antibodies, conflicting with the 67 validated binders stated in the abstract and with the 19+26+22 binders enumerated in Section II.B; this inconsistency must be resolved before the catalog can be treated as reliable. The light-chain pairing limitation is acknowledged by the authors and is less damaging, since SPR validates each expressed pair as a binder.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an integrated antibody discovery pipeline for GFRAL using Trianni mice, combining longitudinal bulk B-cell receptor repertoire sequencing with single-cell paired-chain data. Bulk heavy-chain CDR3 repertoires are analyzed with the STAR method to identify clusters of closely related, putatively antigen-specific sequences; hits present in the single-cell dataset are paired with light chains, expressed, and tested by surface plasmon resonance (SPR). The authors report a 50% success rate among 40 STAR-derived candidates, a catalog of 67 experimentally validated anti-GFRAL antibodies, evidence of convergent selection across mice, AlphaFold3 predictions of antibody-antigen structure, and logistic-regression features predictive of binding.","tokens_in":14753,"tokens_out":7709,"duration_ms":75413,"significance":"If the claims hold, the paper provides a practical resource and a useful benchmark: a catalog of experimentally validated anti-GFRAL antibodies, plus a pipeline that leverages bulk repertoire depth for hit discovery and single-cell data for chain pairing. The external SPR readout is a genuine strength and avoids the circularity that can plague purely computational antibody-discovery studies. The paper also ships the STAR code on GitHub, reports multiple controls (including random single-cell antibodies), and is transparent about acknowledged limitations such as the absence of UMIs and the need for single-cell pairing. However, the strength of the central efficiency claim is currently limited by an incomplete control structure and by unresolved numerical inconsistencies in the counts of validated binders.","major_comments":[{"comment":"The headline 50% success rate is computed for 40 STAR hits that were additionally required to be present in the single-cell dataset, and the paper's controls do not isolate the STAR signal from single-cell presence or frequency. The single-cell frequency criterion gives an 80% success rate, random single-cell antibodies give 20%, and sequences from the same STAR clusters but absent from single-cell give 19%. To support the claim that STAR itself drives enrichment, the authors should compare STAR hits against random bulk sequences present in single-cell, or stratify STAR hits by single-cell frequency. The statement that finding 19 binders among 40 drawn from 1,530,511 bulk sequences has 'very low' probability is not a valid null model, because the 40 sequences were not randomly drawn; the appropriate baseline is the measured 20% random rate, under which 19/40 is indeed significant (binomial p ≈ 0.0004). Please add the missing frequency- or presence-matched control, or explicitly reframe the 50% figure as the success rate of the integrated STAR-plus-single-cell pipeline rather than of STAR in isolation.","section":"II.B / Fig. 3C"},{"comment":"The counts of validated binders are internally inconsistent. The abstract and Section II.B report a catalog of 67 validated binders (19 STAR + 26 single-cell-frequency + 22 bulk-cluster). Section II.D states that of 137 antibodies modeled with AlphaFold3, 70 were binders and 67 non-binders. Since 70+67 = 137, the sentence as written reverses the binder/non-binder split if the catalog contains 67 binders. This is a central deliverable, so the correct totals must be stated and reconciled with the SPR-tested sets enumerated in Section II.B and Figure 3C.","section":"II.D vs Abstract / II.B"},{"comment":"The control group used to assess the role of single-cell presence is under-reported. For each of the 40 STAR clusters the authors state they took '3 or 4 sequences' with high/medium/low bulk frequency and mutation levels, which should yield roughly 120–160 tested antibodies, but the manuscript reports only the 19% success rate without the exact number tested or the number of binders. Without the denominator, the comparison with the 20% random rate cannot be evaluated. The sentence 'this 19% success rate has far greater significance compared to the 20% rate' is also confusing, since 19% is not greater than 20%; presumably the intended meaning is that the implications differ because these sequences were absent from single-cell data. Please report exact counts, clarify how light chains were assigned to heavy chains not present in the single-cell dataset, and rephrase the comparison.","section":"II.B, criterion 3"}],"minor_comments":[{"comment":"The bulk blood sampling day is given as day 38 in Section II.A but as day 39 in the Methods (IV.B) and in the Figure 3A caption; please make the time points consistent.","section":"II.A / IV.B / Fig. 3A"},{"comment":"The sentence 'To the generalizability of the model more rigorously' appears to be missing a verb; please rephrase.","section":"II.E"},{"comment":"The text says there is 'no correlation' between KD and single-cell abundance while reporting Spearman rho = 0.26 with p = 0.03; this should be described as a weak or modest correlation rather than no correlation.","section":"Fig. 4B"},{"comment":"The sentence 'A given CDR3 heavy chain may be paired with multiple heavy and light chains' should refer to multiple light chains (and possibly multiple heavy chains in a broader sense); as written it is unclear.","section":"Discussion"},{"comment":"The data availability statement refers to 'the attached Antibody_SPR.xlsx file' but does not give a persistent link or accession; please provide a stable location for the binder catalog.","section":"IV.D"},{"comment":"The word 'humaninized' is a typo for 'humanized'.","section":"Significance Statement"}],"recommendation":"major_revision","confidential_remarks":"The SPR validation and the transparent reporting of limitations are genuine strengths, and the paper is within the journal's scope. The main risks are the unresolved binder-count inconsistency and the incomplete control structure for attributing the 50% success rate to STAR; both are fixable with additional analyses or more careful claims, so I do not see a need to reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid application paper with a real resource attached—67 or so SPR-validated anti-GFRAL antibodies—and the experimental work is mostly careful. The new thing is not the STAR algorithm, which is prior work from the same group. It is the longitudinal bulk-plus-single-cell design applied to GFRAL, the cross-mouse convergent selection observation, and the validated binder catalog. I would send it to review.\n\nWhat the paper does well: SPR is an independent binding readout, so the pipeline is not self-validating. The authors are candid about the UMI limitation and about heavy/light pairing uncertainty. The logistic regression section includes a proper cross-mouse holdout (AUROC 0.73) rather than only the lineage-confounded 0.89. The AlphaFold3 section is honest about low predictive value. The citation pattern looks fine; STAR is their own method and is properly cited as prior work.\n\nThe soft spot is load-bearing for the efficiency claim. All 40 STAR hits that went to SPR were required to be present in the single-cell data. The only control that removes the single-cell presence effect—reconstructed antibodies from the same STAR clusters but absent from the single-cell data—bound at 19%, statistically indistinguishable from 20% for random single-cell antibodies. So the headline 50% success rate cannot actually be attributed to STAR's clustering signal. It may simply be a single-cell presence/frequency effect. The paper's statement that finding 19 binders in 40 has 'very low' probability under the null is not supported without a frequency-matched or presence-matched baseline. This is fixable: stratify by single-cell presence/frequency, or report STAR hit rates among sequences matched for single-cell presence.\n\nThere is also a numerical inconsistency that needs resolving: the abstract and discussion say 67 validated binders, Section II.D says 137 modeled antibodies split into 70 binders and 67 non-binders, and Section II.B enumerates 19+26+22. The catalog cannot be treated as reliable until these numbers reconcile. Minor point: the paper presents the 19% vs 20% comparison as meaningful, but it is evidence against the STAR effect, not for it.\n\nLight-chain pairing is a real limitation but less damaging, because SPR validates each expressed pair as a binder. The catalog is still a catalog of binders, just not necessarily the in vivo pairs.\n\nWho this is for: people working on antibody discovery pipelines and GFRAL/GDF-15 therapeutics. It deserves serious peer review despite the statistical weakness, because the resource is useful and the confound is addressable. Send it to review with a request to fix the baseline analysis and the binder-count discrepancy.","headline":"A useful GFRAL binder resource and a clean experimental validation setup, but the headline STAR enrichment claim is not established because the only control that isolates STAR's clustering signal performs no better than random.","tokens_in":15178,"tokens_out":2367,"would_cite":false,"duration_ms":24141,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a computational pipeline—STAR on bulk repertoires plus single-cell light-chain recovery—shortlists GFRAL-binding antibodies with a 50 percent experimental hit rate and delivers a catalog of 67 validated binders.","keywords":["antibody discovery","GFRAL","GDF-15","B cell receptor repertoire","STAR method","single-cell V(D)J sequencing","affinity maturation","convergent selection"],"falsifier":"Express each of the 40 STAR heavy chains with every light chain observed paired to it in the single-cell data, not only the representative chosen by the pipeline, and measure SPR binding; if the 50 percent hit rate collapses or the binding specificities change, the single-cell pairing choice rather than the heavy-chain clump is carrying the result.","tokens_in":14189,"feed_emoji":"🧬","tokens_out":6817,"duration_ms":66755,"temperature":0.7,"pith_summary":"The paper tries to establish a faster antibody-discovery route: instead of screening large libraries or sorting many single cells, run STAR on bulk B-cell receptor sequencing from immunized mice, then use single-cell data only to recover the light chains of the heavy-chain hits. Applied to GFRAL, the receptor for the appetite-regulating cytokine GDF-15, the pipeline selected 40 antibody candidates, and surface plasmon resonance confirmed that 19 of them (50 percent) bind the target. Additional selection strategies raised the total to a catalog of 67 validated anti-GFRAL antibodies. A sympathetic reader would care because GFRAL is the hub of the GDF-15 appetite pathway, making these binders starting points for drugs against cachexia, anorexia, obesity, and diabetes, and because the pipeline's hit rate suggests computational preselection can cut the time and cost of finding therapeutic antibodies.","feed_headline":"Half of computed antibody picks bind the appetite receptor GFRAL","feed_subtitle":"STAR marks affinity-matured clusters in bulk data; single-cell pairing turns them into 67 tested GFRAL binders.","key_machinery":"The load-bearing object is the STAR hit: a cluster of heavy-chain CDR3 nucleotide sequences that differ by one amino acid and contain more near-neighbors than expected by chance, with a cluster-level threshold of at least 10 over-threshold sequences. STAR scans each bulk time point independently, ranks sequences by neighbor count, and outputs clusters that bear the signature of affinity maturation. The second mechanism is the bulk-to-single-cell mapping: each STAR cluster is represented by its highest-neighbor sequence, and that heavy chain is paired with a light chain found in the single-cell data, yielding an expressible, testable antibody. The mapping is what turns a statistical clump in deep bulk data into a physical reagent.","core_discovery":"The authors claim that a two-step integration—deep bulk repertoire sequencing plus targeted single-cell pairing—can replace exhaustive single-cell screening as the primary engine of antibody discovery. In three humanized Trianni mice immunized with GFRAL, STAR identified 40 heavy-chain CDR3 clusters with statistically overrepresented affinity-maturation neighborhoods across time points; matching those heavy chains to paired light chains in single-cell data and expressing the reconstructed antibodies yielded 19 SPR-confirmed binders (50 percent), against a 20 percent success rate for randomly chosen single-cell antibodies. High-frequency single-cell sequences alone performed better (80 percent), and heavy chains from the same STAR clusters paired with single-cell light chains but absent from single-cell data bound at 19 percent, showing that reconstructed pairs can expand the candidate pool beyond observed sequences. The paper further reports convergent selection (13 binder sequences shared across mice), a weak positive correlation between single-cell abundance and affinity (Spearman rho = 0.26), an AlphaFold3 interface score (ipTM) that is higher on average for binders but too noisy to classify them, and a CDR3-sequence logistic-regression model that predicts binding with AUROC 0.89 within mice and 0.73 across mice.","pith_inferences":["This pipeline should transfer to other antigens with strong germinal-center responses, but its hit rate will likely depend on how densely the responding clones expand; weak or T-cell-independent responses may produce no STAR clusters above threshold.","A testable extension is to rank STAR hits by the logistic-regression CDR3 weights or AlphaFold3 ipTM before expression; if such ranking lifts the 50 percent validation above, say, 70 percent, the computational steps become a true pre-screen rather than a triage aid.","Because a heavy chain can pair with several light chains, the current catalog probably underestimates the true binder space; re-screening the same heavy chains against alternate observed light chains would reveal how many GFRAL specificities were lost to the single-cell pairing rule.","Convergent CDR3 motifs shared across mice suggest that some GFRAL epitopes are consistently targeted; mapping those motifs onto the AlphaFold3-predicted structures could nominate the dominant epitope before any competition-binding experiment."],"forward_implications":["If the 50 percent validation rate holds, computational preselection with STAR can replace the usual practice of screening hundreds of single-cell-derived antibodies, reserving single-cell sequencing for light-chain recovery only.","The 67 validated anti-GFRAL binders give drug development programs an immediate panel for testing GDF-15 pathway blockade (cachexia, anorexia) or activation (obesity, diabetes), including antibodies absent from single-cell data.","The 80 percent success rate for high-frequency single-cell sequences identifies a simple abundance threshold as a strong predictor of binding, while the weak KD-abundance correlation warns that abundance alone will miss high-affinity rare clones.","The cross-mouse convergent CDR3 sequences define reproducible public response motifs that could seed epitope-focused or germline-targeting vaccine designs.","Using ipTM as a pre-filter before SPR could raise the hit rate further, since binders score higher on average even though AlphaFold3 cannot reliably separate binders from non-binders."],"supporting_citations":[{"why":"Supplies the STAR computational method whose neighbor-density clusters identify responding heavy-chain CDR3 sequences in bulk repertoire data.","marker":"[42]"},{"why":"The humanized Trianni mouse platform generates antibodies with fully human variable regions, the experimental substrate of the pipeline.","marker":"[43]"},{"why":"Four independent identifications of GFRAL as the GDF-15 receptor; establishes the biological target and its therapeutic rationale.","marker":"[26–29]"},{"why":"Shows antibody-mediated inhibition of GDF15-GFRAL reverses cancer cachexia in mice, the key therapeutic motivation for the binder catalog.","marker":"[36]"},{"why":"AlphaFold3 provides the predicted structures and ipTM interface scores used to compare binders and non-binders.","marker":"[48]"},{"why":"The Left-Right one-hot encoding converts CDR3 sequences into positional features for the logistic-regression binding predictor.","marker":"[51]"}],"fun_headline_variants":["Half of STAR-picked antibodies bind GFRAL","67 GFRAL binders from bulk plus single-cell search","Convergent CDR3s mark GFRAL binders in multiple mice","Computational screen beats random for GFRAL antibodies","STAR maps GFRAL antibodies faster than exhaustive screening"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that the one light chain recovered from single-cell data for each heavy-chain hit is the functionally correct partner, even though a heavy chain can pair with several light chains.","fun_headline_variants_meta":{"raw":{"variants":["Half of STAR-picked antibodies bind GFRAL","67 GFRAL binders from bulk plus single-cell search","Convergent CDR3s mark GFRAL binders in multiple mice","Computational screen beats random for GFRAL antibodies","STAR maps GFRAL antibodies faster than exhaustive screening"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000225,"raw_usage":{"total_tokens":1482,"prompt_tokens":982,"completion_tokens":500,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":415}},"tokens_in":598,"tokens_out":500,"duration_ms":5095,"temperature":1.0,"reasoning_tokens":415,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:24:22.051121+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Express each of the 40 STAR heavy chains with every light chain observed paired to it in the single-cell data, not only the representative chosen by the pipeline, and measure SPR binding; if the 50 percent hit rate collapses or the binding specificities change, the single-cell pairing choice rather than the heavy-chain clump is carrying the result.","supporting_citations":[{"cited_title":"Computational detection of antigen-specific b cell receptors following immuniza- tion","cited_arxiv_id":null,"evidence_quote":"Supplies the STAR computational method whose neighbor-density clusters identify responding heavy-chain CDR3 sequences in bulk repertoire data."},{"cited_title":"The trianni mouse: The next genera- tion transgenic platform for the isolation of fully human monoclonal antibodies","cited_arxiv_id":null,"evidence_quote":"The humanized Trianni mouse platform generates antibodies with fully human variable regions, the experimental substrate of the pipeline."},{"cited_title":"Antibody-mediated inhibition of gdf15–gfral ac- tivity reverses cancer cachexia in mice","cited_arxiv_id":null,"evidence_quote":"Shows antibody-mediated inhibition of GDF15-GFRAL reverses cancer cachexia in mice, the key therapeutic motivation for the binder catalog."},{"cited_title":"Accurate structure prediction of biomolecular in- teractions with alphafold 3","cited_arxiv_id":null,"evidence_quote":"AlphaFold3 provides the predicted structures and ipTM interface scores used to compare binders and non-binders."},{"cited_title":"12 Population variability in the generation and selection of t-cell repertoires","cited_arxiv_id":null,"evidence_quote":"The Left-Right one-hot encoding converts CDR3 sequences into positional features for the logistic-regression binding predictor."}],"review_version":1}