{"id":"47e77049-53fa-46db-b37e-283e65fcf7b8","arxiv_id":"2411.17798","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"DapPep, built from ESM-2 with cross-attention and peptide-reconstruction pre-training, reports ROC-AUC 0.816 and PR-AUC 0.836 on unseen peptides, beating PanPep by about 9 to 11 percent.","lead":"A new model called DapPep predicts which T-cell receptors bind to antigen peptides, and it claims to beat current tools especially for peptides never seen during training. If confirmed, the model could help select tumor-fighting T cells for immunotherapy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Zero-shot gains may be artifacts of unverified negative controls and split leakage; Section III-A gives no construction details for the 60M control TCRs or for peptide-disjoint splits.","rationale":"The reader's weakest_assumption names exactly this combination of split leakage and unvalidated negative controls; I agree. My independent reading of Section III-A found no description of how the 60,333,379 control TCRs were generated, and the method section does not show that the asyAE pre-training peptides are disjoint from ZeroSet/FewSet peptides. I also note secondary issues: Table I text numbers differ from Table I entries for DLpTCR and pMTnet, and no variance or error bars are reported, but those are less likely to overturn the core claim than the control/split question. I do not see an internal inconsistency in the architecture itself; the method is plausible and the ESM-2 initialization is reasonable. The concern is not that the results disagree with consensus, but that the evidence as presented does not eliminate a plausible alternative explanation: the model may be detecting source-repertoire differences rather than peptide-specific binding. A concrete re-evaluation with documented peptide-disjoint splits and experimentally validated negatives would settle this. If the numbers survive that test, the central claim would be much stronger; if not, the reported margins are not trustworthy. Therefore I recommend keeping the reader's conditional verdict rather than moving to accept or reject.","tokens_in":8844,"tokens_out":5613,"duration_ms":52339,"concrete_test":"Obtain the exact split files (test_list, zero_test_list, few_test_list) and the training/pre-training peptide lists; compute the number of ZeroSet/FewSet peptides with exact sequence overlap. Then re-evaluate ZeroSet and FewSet with the same positive pairs but negatives restricted to experimentally validated non-binding TCR-peptide pairs (e.g., IEDB/negative assay data), keeping the same class ratio. If any overlap is nonzero, or if DapPep's AUC drops by more than about 0.05 while PanPep's does not, the generalization claim is not supported by the current evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that DapPep generalizes to unseen peptides better than PanPep—rests on the ZeroSet/FewSet evaluation described in Section III. The most load-bearing unknown is the negative-control construction: Section III-A states only that 'All the binding TCRs are balanced by a controlTCRset, where the controlTCRset contains 60,333,379 non-binding TCRs (negative samples).' It does not say how these 60M TCRs were obtained, whether they were experimentally confirmed non-binders, or whether they are simply unlabeled TCRs never tested against the peptide. If the latter, the model can separate curated binding TCRs from arbitrary repertoire TCRs using distributional cues (CDR3 length, V/J usage, source cohort) instead of peptide-specific binding, inflating ROC-AUC and especially PR-AUC. The zero-shot claim also requires that every ZeroSet/FewSet peptide is absent from the binding-affinity training set and from the Section II-A peptide-only asyAE pre-training set. The paper cites PanPep's datasets instead of documenting its own split construction, so no peptide-overlap statistics are given. Combined with the fact that each reported number is a single point with no variance or significance testing, the 9–11% margin over PanPep is not yet separable from an artifact of control sampling.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DapPep, a TCR-peptide binding affinity prediction framework that combines an ESM-2-initialized TCR encoder, a shallow peptide encoder, a cross-attention module, and an asymmetric autoencoder (asyAE) pre-training objective. The method is evaluated on PanPep's majority, few-shot, and zero-shot benchmarks, where it is claimed to outperform PanPep and other baselines, and on a gastrointestinal neoantigen T-cell sorting task. The central claim is that DapPep generalizes to unseen peptides and data-scarce settings better than existing tools, with ROC-AUC 0.816 and PR-AUC 0.836 on unseen peptides and 0.835/0.834 on the clinical sorting task.","tokens_in":9116,"tokens_out":3783,"duration_ms":31513,"significance":"If the reported results are correct, DapPep would be a practical sequence-only predictor for novel antigens, with direct relevance to neoantigen therapy and vaccine design. The architecture is lightweight and the two-stage training idea (pre-training a peptide reconstruction module before supervised binding learning) is interesting. However, the evidence as presented does not yet establish the claim: there is no code or data release, no error bars or significance tests, no documentation of negative-control construction or peptide-disjoint splits, and there are inconsistencies between the text and Table I. The central message is plausible but not rigorously supported.","major_comments":[{"comment":"The negative-control construction is underspecified and load-bearing. The text says only 'All the binding TCRs are balanced by a controlTCRset, where the controlTCRset contains 60,333,379 non-binding TCRs (negative samples).' It is not stated whether these 60M TCRs are experimentally confirmed non-binders or are unlabeled repertoire TCRs never tested against the relevant peptides. If they are unlabeled, the model can separate curated binders from background repertoire using distributional cues such as CDR3 length, V/J usage, or source cohort, rather than peptide-specific binding. This would inflate both ROC-AUC and PR-AUC and directly affect the zero-shot and few-shot claims. Please describe how the control set was constructed, whether any filtering or matching was applied, and provide overlap statistics between the positive and negative sets. Without this, the reported margins over PanPep are not interpretable.","section":"Section III-A"},{"comment":"The reported numbers are internally inconsistent. In Section III-B, DapPep's zero-shot results are ROC-AUC 0.787 and PR-AUC 0.815, while Table I (Unseen Peptides) reports 0.816 and 0.836. Baseline numbers also disagree: pMTnet is listed as 0.564/0.555 in Table I but 0.563/0.555 in the text; ERGO2 is 0.504/0.524 in Table I but 0.496/0.542 in the text; DLpTCR is 0.483/0.481 in Table I but 0.517/0.488 in the text. Because the central claim is the margin over PanPep, these discrepancies must be reconciled. The manuscript should present one consistent set of results and explain the source of the differences (e.g., different evaluation subsets or random seeds).","section":"Section III-B and Table I"},{"comment":"There is no evidence that the zero-shot and few-shot test peptides are truly unseen. The claim that 'the peptides in this dataset are not available in the training set of DapPep and other baseline tools' is not backed by any documentation of the split construction or peptide-level overlap statistics. In addition, the asyAE pre-training in Section II-A uses 'only the peptide sequences from the training set of TCR-peptide pairs,' so it is also necessary to confirm that ZeroSet/FewSet peptides are excluded from the pre-training corpus. Please provide peptide-overlap analysis between the training set, the pre-training set, and the ZeroSet/FewSet, and describe the exact provenance of the splits.","section":"Section III-C and Section II-A"},{"comment":"All results are single-point estimates with no variance, confidence intervals, or statistical significance tests. Given that the headline improvements over PanPep are roughly 9-11% in ROC-AUC, it is essential to know whether the difference is stable across random seeds or data subsamples. For the clinical T-cell sorting task, the PR-AUC difference between PanPep (0.781) and DapPep (0.834) is smaller relative to the apparent variability in the other comparisons. Please report at least three independent runs (or bootstrapped CIs) and, ideally, a paired significance test such as a McNemar or bootstrap test on the test pairs.","section":"Section III-B and III-D"}],"minor_comments":[{"comment":"Equation (4) contains an extra closing parenthesis: 'LDapPep = MSE(Scorepred, Scoretarget)).' should be 'LDapPep = MSE(Scorepred, Scoretarget).'","section":"Equation (4)"},{"comment":"The caption of Figure 2 includes 'Ref. Panpep(Fig. 2d)' followed by an unexplained code snippet in Chinese-style brackets; this appears to be a leftover artifact and should be removed.","section":"Figure 2 caption"},{"comment":"The manuscript says the dataset follows PanPep but gives no dataset statistics (e.g., number of positive TCR-peptide pairs per split, number of peptides, TCR chain type, length distributions). A data-availability and reproducibility statement would be helpful.","section":"Section III-A"},{"comment":"The conclusions list 'neonatal antigen' where 'neoantigen' is intended; the same typo appears in the abstract. Also, 'zero-setting' in Section III-B should be 'zero-shot setting.'","section":"Section IV"},{"comment":"The pre-training objective cites a long list of references ([7],[8],[10],[11],[13],[24],[28]-[30],[35]-[43]) that appear largely unrelated to the asyAE framework; these citations do not support the specific claim being made and should be replaced with relevant prior work on autoencoders or sequence reconstruction.","section":"Section II-A Pre-training Objective"}],"recommendation":"major_revision","confidential_remarks":"The paper is a methods paper with a clinically relevant claim, but the empirical evidence is currently insufficient. The inconsistent numbers between text and Table I, the lack of negative-control documentation, and the absence of error bars or code/data are all fixable within a revision. I would also note an unusual number of self-citations in the pre-training objective section; the authors should be asked to justify or trim these. Given the load-bearing reproducibility gaps, major revision is appropriate, not rejection at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nDapPep is a plausible extension of the PanPep benchmark with a sensible architecture, but the reported zero-shot gains over PanPep are not yet separable from artifact. The paper lacks code, data, error bars, a description of how the 60M negative control TCRs were assembled, and any demonstration that test peptides were truly absent from training or pre-training. It also reports inconsistent numbers for the same zero-shot setting in the text (0.787 ROC-AUC, 0.815 PR-AUC) and in Table I (0.816, 0.836).\n\nWhat is actually new: a combination of ESM-2 to initialize the TCR encoder, a cross-attention module between peptide query and TCR key/value, and an asymmetric autoencoder that reconstructs peptide sequences as a pre-training objective. That is a new combination, and the clinical validation on the gastrointestinal neoantigen dataset is a useful addition. The paper is clearly written and the method is not obviously wrong.\n\nThe main soft spot is the negative-control construction. Section III-A merely says that binding TCRs are balanced by a controlTCRset of 60,333,379 'non-binding TCRs.' If those are simply unlabeled repertoire TCRs never tested against the peptide, the model can separate curated binders from arbitrary repertoire sequences using CDR3 length, V/J usage, or cohort, inflating both ROC-AUC and especially PR-AUC. That is not a hypothetical concern with this data type. Similarly, the zero-shot claim depends on every ZeroSet/FewSet peptide being excluded from both the binding-affinity training set and the asyAE pre-training set; the paper states this but gives no overlap statistics or split construction details. Given single-point estimates and no significance tests, the 9–11% margin over PanPep is not yet convincing.\n\nThe inconsistent numbers are a concrete problem. A referee should ask for corrected, variance-aware reporting. There is no evidence of fraud, and the method could well work, but as presented the evidence does not establish the headline claim.\n\nWho gets value: anyone working on TCR-peptide prediction or benchmarking domain-adaptive immunology models. It deserves a serious referee, but only conditionally; require code/data release, detailed control and split descriptions, and corrected numbers.\n\nRecommendation: send to peer review, expect major revision.","headline":"A plausible but unverified zero-shot TCR-peptide predictor: the architecture is new, but undocumented control construction, missing leak checks, and inconsistent numbers leave the headline claim unproven.","tokens_in":15,"tokens_out":2871,"would_cite":false,"duration_ms":45380,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DapPep, a sequence-only model, predicts TCR-antigen binding for unseen peptides with ROC-AUC 0.816, outperforming prior tools.","keywords":["T-cell receptor","peptide binding prediction","zero-shot learning","domain adaptation","self-supervised pre-training","attention mechanism","neoantigen","immunoinformatics"],"falsifier":"One could train DapPep on the exact same data but with all ZeroSet peptides artificially added to the training set; if performance on the zero-shot benchmark jumps dramatically, the claimed generalization is mostly memorization. Similarly, if re-analysis shows any zero-shot test peptide overlapping the pre-training peptide set, or if experimentally validated non-binders among the control TCRs are shown to be rare, the reported AUC/PR-AUC advantage would need to be re-evaluated.","tokens_in":85,"feed_emoji":"🧬","tokens_out":4746,"duration_ms":40536,"temperature":0.7,"pith_summary":"This paper introduces DapPep, a sequence-only model that predicts whether a T-cell receptor (TCR) binds a given antigenic peptide, with a focus on peptides never seen in training. The authors claim DapPep consistently outperforms existing tools across majority, few-shot, and zero-shot settings, and that its advantage is largest for unseen peptides: ROC-AUC 0.816 and PR-AUC 0.836 on a zero-shot benchmark, versus 0.744/0.755 for the prior state of the art PanPep. DapPep also scores best on a clinical task of sorting reactive T cells for gastrointestinal neoantigens (0.835/0.834). If these results hold, the model gives clinicians and vaccine designers a fast, general predictor for novel antigens without per-peptide fine-tuning.","feed_headline":"Unseen-peptide TCR binding predicted at 0.816 AUC","feed_subtitle":"DapPep generalizes without fine-tuning, aiding neoantigen therapy and vaccine screening.","key_machinery":"The load-bearing component is the TCR-peptide cross-attention module (T PRepr). It receives TCR features from the ESM-2-initialized TCR encoder as key and value, peptide features as query, and produces a joint representation that a linear decoder maps to a binding probability. Before supervised training, the same module is pre-trained in an inner-loop, peptide-reconstruction objective using an asymmetric auto-encoder (asyAE) that treats peptides as both input and output; the authors argue this self-supervised step, rather than direct transfer learning, is what lets the model generalize to peptides with few or no known binding TCRs.","core_discovery":"The central claim is that a lightweight, peptide-agnostic architecture can learn TCR-peptide binding patterns that transfer to completely unseen peptides. DapPep encodes TCRs with an ESM-2-initialized transformer, encodes peptides with shallow self-attention, and combines them in a cross-attention module; before supervised binding training, that module is pre-trained as an asymmetric autoencoder to reconstruct peptide sequences, which the authors argue teaches it the essential characteristics of peptides. After transfer to the binding task with only an MSE regression loss, the model outperforms PanPep, pMTnet, ERGO2, and DLpTCR on the zero-shot benchmark and on a gastrointestinal neoantigen T-cell sorting task, with no fine-tuning on the target peptides.","pith_inferences":["Editorial: the paper's claimed advantage would be most convincing if the authors showed that the 60-million-TCR control set contains true non-binders rather than unobserved pairs; if many controls are simply unlabeled, the absolute AUC numbers may be inflated even if relative rankings hold.","Editorial: the zero-shot result implies that peptide identity is not memorized but rather reconstructed from composition and motif patterns; a direct test would be to mutate a known antigen at anchor positions and check whether predicted affinity drops as expected.","Editorial: because the pre-training uses only peptide sequences from the training pairs, the method's generalization could depend on the diversity of peptides in the training set; extending the asyAE to an external peptide corpus might further improve zero-shot transfer.","Editorial: the architecture could be adapted to paired alpha/beta TCR sequences, which are increasingly available, and might improve accuracy for peptides where the beta chain alone is insufficient."],"forward_implications":["DapPep can rank candidate neoantigens for T-cell reactivity without requiring known peptide-specific TCR data, potentially accelerating neoantigen vaccine and adoptive cell therapy pipelines.","The zero-shot and few-shot gains hold without task-specific fine-tuning, so the model can be deployed directly on new peptides in high-throughput screening.","Because the model is sequence-only and lightweight, it can be applied at the scale of full TCR repertoire screens to filter reactive T cells before experimental validation.","The same two-stage recipe—self-supervised peptide reconstruction followed by supervised affinity regression—could be transferred to other receptor-ligand binding problems with sparse training data.","Comparisons on the gastrointestinal cancer dataset suggest the model can sort tumor-infiltrating lymphocytes by neoantigen specificity, a step toward personalized immunotherapy."],"supporting_citations":[{"why":"Supplies the majority, zero-shot, and few-shot benchmark splits and the PanPep baseline that DapPep is compared against.","marker":"[4]"},{"why":"Provides the ESM-2 pretrained protein language model used to initialize the TCR representation module.","marker":"[16]"},{"why":"Provides the gastrointestinal neoantigen dataset used for the clinical T-cell sorting validation.","marker":"[25]"},{"why":"Defines the multi-head self-attention mechanism on which the TCR, peptide, and cross-attention modules are built.","marker":"[26]"},{"why":"pMTnet, a peptide-agnostic baseline that DapPep outperforms on unseen-peptide and TCR-sorting tasks.","marker":"[18]"},{"why":"ERGO2, a peptide-agnostic baseline that DapPep outperforms on both evaluation tasks.","marker":"[23]"},{"why":"DLpTCR, a baseline that DapPep outperforms on both evaluation tasks.","marker":"[31]"}],"fun_headline_variants":["DapPep predicts TCR-peptide binding for unseen antigens","Universal TCR-antigen binding prediction without fine-tuning","Peptide-agnostic model generalizes to novel antigens","Zero-shot TCR binding affinity for unseen peptides","DapPep: robust binding affinity beyond known antigens"],"cache_read_input_tokens":11776,"weakest_assumption_plain":"The reported generalization rests on the assumption that ZeroSet and FewSet peptides never appear in DapPep's training data and that the 60,333,379 control TCRs are true non-binders, not merely unobserved or unlabeled pairs.","fun_headline_variants_meta":{"raw":{"variants":["DapPep predicts TCR-peptide binding for unseen antigens","Universal TCR-antigen binding prediction without fine-tuning","Peptide-agnostic model generalizes to novel antigens","Zero-shot TCR binding affinity for unseen peptides","DapPep: robust binding affinity beyond known antigens"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000716,"raw_usage":{"total_tokens":3178,"prompt_tokens":865,"completion_tokens":2313,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":2246}},"tokens_in":481,"tokens_out":2313,"duration_ms":15105,"temperature":1.0,"reasoning_tokens":2246,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:52:51.234494+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One could train DapPep on the exact same data but with all ZeroSet peptides artificially added to the training set; if performance on the zero-shot benchmark jumps dramatically, the claimed generalization is mostly memorization. Similarly, if re-analysis shows any zero-shot test peptide overlapping the pre-training peptide set, or if experimentally validated non-binders among the control TCRs are shown to be rare, the reported AUC/PR-AUC advantage would need to be re-evaluated.","supporting_citations":[{"cited_title":"Pan- peptide meta learning for t-cell receptor–antigen binding recognition","cited_arxiv_id":null,"evidence_quote":"Supplies the majority, zero-shot, and few-shot benchmark splits and the PanPep baseline that DapPep is compared against."},{"cited_title":"Language models of protein sequences at the scale of evolution enable accurate structure prediction","cited_arxiv_id":null,"evidence_quote":"Provides the ESM-2 pretrained protein language model used to initialize the TCR representation module."},{"cited_title":"Immunogenicity of somatic mutations in human gastrointestinal cancers","cited_arxiv_id":null,"evidence_quote":"Provides the gastrointestinal neoantigen dataset used for the clinical T-cell sorting validation."},{"cited_title":"Deep learning-based prediction of the t cell receptor–antigen binding specificity","cited_arxiv_id":null,"evidence_quote":"pMTnet, a peptide-agnostic baseline that DapPep outperforms on unseen-peptide and TCR-sorting tasks."},{"cited_title":"Prediction of specific tcr-peptide binding from large dictionaries of tcr-peptide pairs","cited_arxiv_id":null,"evidence_quote":"ERGO2, a peptide-agnostic baseline that DapPep outperforms on both evaluation tasks."},{"cited_title":"Dlptcr: an ensemble deep learning framework for predicting immuno- genic peptide recognized by t cell receptor","cited_arxiv_id":null,"evidence_quote":"DLpTCR, a baseline that DapPep outperforms on both evaluation tasks."}],"review_version":1}