{"id":"0e8511a7-79de-424b-8342-ed0337d3b4cf","arxiv_id":"2507.00087","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A unified pre-trained model, pUniFind, jointly scores database search results and performs open de novo sequencing, reporting increased peptide identifications across diverse proteomics datasets.","lead":"pUniFind is a deep learning model that scores peptide-spectrum matches and performs open de novo sequencing, trained on over 100 million mass spectra. It reports more peptide identifications than existing tools on several proteomics datasets, especially immunopeptidomics, and claims the first deep learning open de novo sequencing with over 1,300 modification types.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"De novo recall and the '60% more PSMs' claim are measured against Open-pFind labels, the same engine that annotated the training set; learned annotator bias could inflate the central open de novo claim.","rationale":"The reader's weakest assumption - that Open-pFind annotations serve as ground truth for both training and evaluation - is exactly the load-bearing concern. The central claim of superior open de novo sequencing rests on the 21PTM evaluation, where the ground truth is derived from the same engine used to label the training set. This creates a direct circularity: pUniFind is trained to reproduce Open-pFind's ranking, and its de novo accuracy is then measured by agreement with Open-pFind. The paper's validation strategies (metabolic labeling, entrapment, mixed-species searches) address database search accuracy but do not provide independent ground truth for the de novo modified-peptide predictions on 21PTM. Because the reader already conditioned acceptance on independent de novo evaluation, the existing CONDITIONAL verdict is appropriate; no change is needed. The concrete test would settle whether the circularity actually inflates the headline numbers.","tokens_in":12909,"tokens_out":10295,"duration_ms":108602,"concrete_test":"Re-annotate the 21PTM spectra with an independent engine that does not share Open-pFind's training data or scoring (e.g., MSFragger open search with PTM-Shepherd, or use synthetic ProteomeTools peptides with known modifications) and recompute pUniFind's peptide-, modification-, and site-level recalls against those labels instead of Open-pFind. If the 63.8% peptide recall or the 60% improvement over pNovo drops by more than, say, 10 percentage points, the circularity concern is confirmed and the central de novo claim would need to be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that Open-pFind annotations used to build the 100M-PSM training set are accurate enough to serve as ground truth for both training and evaluation. The model is trained to rank Open-pFind's top candidate above its ranks 3-10 (Fig. 1), and the central de novo claims are then measured against Open-pFind-derived labels: the 21PTM peptide-level recall (63.8%) and the '60% more PSMs than pNovo' are computed by matching pUniFind's predictions to Open-pFind identifications, and the modification/site-level accuracies in Fig. 5c-e are defined by agreement with Open-pFind. This is circular: a model that has learned Open-pFind's systematic errors - especially for rare modifications, where the training set is sparse (fewer than 100 PMMINs for several lysine-targeted modifications) - will show artificially high recall. The problem is aggravated by the absence of any independent ground truth for the 21PTM set: no synthetic peptide validation, no cross-engine consensus with an engine not used in training, and no explicit statement that the 21PTM evaluation spectra were excluded from the 6,524 training files. If Open-pFind's false positives are concentrated in the same modification classes that pUniFind is claimed to handle, the reported 60% improvement over pNovo and the 63.8% recall are not evidence of accurate de novo sequencing; they are evidence of agreement with the annotator.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces pUniFind, a large pre-trained multimodal deep learning model that jointly performs peptide–spectrum match (PSM) rescoring for database search and open, zero-shot de novo sequencing. The model is trained on over 100 million spectra annotated by Open-pFind using open database search, and it is evaluated on nine-species datasets, Astral and timsTOF data, immunopeptidomics, metaproteomics, and the 21PTM dataset. The central claims are that pUniFind improves peptide identifications by up to 42.6% in immunopeptidomics relative to Open-pFind, identifies 60% more PSMs than existing de novo methods in a 300-fold larger search space, and enables modification-aware de novo sequencing with a peptide-level recall of 63.8% on the 21PTM benchmark. The paper also describes a deep learning-based quality control module that recovers additional peptides, including peptides mapped to the genome but absent from reference proteomes.","tokens_in":13217,"tokens_out":2400,"duration_ms":27537,"significance":"If the reported results are accurate, pUniFind would represent a substantial advance in computational proteomics: a single model that unifies database search rescoring and open de novo sequencing, with demonstrated gains across diverse instruments and applications. The scale of the training data, the breadth of validation strategies (entrapment, metabolic labeling, mixed-species searches, and genome-translated database matching), and the explicit design to avoid label leakage by decoupling representation learning from scoring are notable strengths. The claim of handling over 1,300 modifications in a zero-shot de novo setting is particularly significant and would address a long-standing limitation of existing de novo tools. However, the evaluation's reliance on Open-pFind as both the training annotator and the ground-truth labeler for de novo accuracy creates a substantial circularity risk that must be resolved before the central claims can be accepted.","major_comments":[{"comment":"The de novo accuracy evaluation is circular: the training set is annotated by Open-pFind, and the modification-level, site-level, and sequence-level accuracies in Fig. 5c–e are defined by agreement with Open-pFind identifications. The reported average peptide-level recall of 63.8% on the 21PTM dataset and the 60% improvement over pNovo are therefore measurements of agreement with the annotator, not independent evidence of correct de novo sequencing. The paper should include validation against an independent ground truth—for example, synthetic peptide spectra with known modifications, a cross-engine consensus that excludes Open-pFind, or a clearly documented exclusion of the 21PTM spectra from the training set. Without such a control, the central de novo claims remain unverified.","section":"Fig. 1 and training description"},{"comment":"The text does not explicitly state that the 21PTM evaluation spectra were excluded from the 6,524 files used for training. If any of these spectra or their near-identical counterparts appeared in training, the reported recall and accuracy numbers would be inflated by memorization. This is a load-bearing issue for the de novo claims; the authors must either provide an explicit exclusion statement or quantify the overlap.","section":"Results: 'Application of pUniFind in open de novo sequencing'"},{"comment":"The performance gains across the nine-species datasets are reported as a single value per dataset without error bars, replicate runs, or statistical significance testing. The reported improvements range from 2% to 18% over Open-pFind, and the variability across datasets is substantial. Given that the database search workflow uses Open-pFind as the candidate generator and pUniFind as the rescoring model, the paper should demonstrate that the observed gains are reproducible and not driven by a few spectra or by the specific choice of the top-k prefilter (10 or 20). At minimum, a per-dataset breakdown of the number of spectra and the variance across subsets should be provided.","section":"Results: 'Performance evaluation on MS/MS data from various species'"},{"comment":"The workflow rescores only spectra whose top-ranked candidate has a q-value below 0.1 from Open-pFind, and then applies target-decoy analysis to the final pUniFind scores. The validity of this two-step FDR control depends on the assumption that pUniFind's scores are well-calibrated and that the prefilter does not distort the target-decoy ratio. The paper should report the target-decoy score distributions for pUniFind (e.g., as shown for timsTOF in Fig. 3d) for the database search results, and should justify the q=0.1 prefilter threshold. Without this, the reported false discovery rates cannot be verified.","section":"Results: 'The pUniFind model and its integration into the database search workflow'"}],"minor_comments":[{"comment":"The abstract states that pUniFind is 'the first large-scale multimodal pre-trained model in proteomics' and 'the first deep learning-based open de novo sequencing method.' The claims of novelty should be qualified by placing them in the context of recent work such as DeepSearch, DDA-BERT, and Casanovo, and by clarifying the specific sense in which 'unified' is used (the model itself performs both tasks, but the database search workflow still relies on Open-pFind for candidate generation).","section":"Introduction"},{"comment":"The architecture description in Fig. 1 mentions a 'Peptide Length Aware (PLA) module' but the main text does not define PLA or explain how length conditioning is implemented. Please provide a clear description in the Methods section.","section":"Fig. 1"},{"comment":"The sentence 'Tesorai slightly outperformed conventional search engines' is vague; please specify which engines and datasets are compared, and provide the corresponding numbers in the supplement.","section":"Results"},{"comment":"The text says 'the target PTM ranked within the top four by number of identified PSMs, and was the most frequently identified modification in 81% of the datasets,' but the corresponding ranking for the remaining 19% is not shown. Please include the complete ranking table in the supplement.","section":"Fig. 5"},{"comment":"Several references contain placeholder question marks (e.g., 'SEQUEST ?', 'Alphapept ?', 'Tesorai ?', '8'). Please resolve these. Also, 'a a 60% improvement' on page 10 contains a duplicated article. These should be corrected before publication.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reports a large and potentially important advance, but the evaluation design is heavily dependent on Open-pFind, which is from the same group (Hao Chi, pFind Studio). This is not inherently disqualifying, but the circularity in the de novo evaluation (training labels and ground truth both derived from Open-pFind) is a genuine validity threat that needs to be addressed with independent ground truth. I would recommend requesting major revision rather than reject, because the central claims may be salvageable with additional validation. I also note that the paper lacks a proper Methods section in the provided text; if the full manuscript contains it, the authors should ensure that the training/evaluation overlap and FDR calibration details are explicitly included."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Briefly: pUniFind is a legitimate advance, and the database-search half of the paper stands up reasonably well. But the open de novo numbers are measured against the same Open-pFind labels that built the training set, so I would not take the 60% improvement over pNovo or the 63.8% recall at face value until there is independent ground truth.\n\nWhat is genuinely new: the first model I know of that does end-to-end PSM rescoring and open de novo sequencing in one architecture, trained on 100M open-search spectra. The coverage of 1,300+ modification types is unusual, and the paper backs the database-search claims with several independent checks: entrapment, metabolic labeling, mixed-species searches, plus nine-species and multi-instrument benchmarking. Those checks suggest the rescoring really does add sensitivity without obvious FDR inflation. Credit where due: this is a lot more evaluation than most proteomics papers ship with.\n\nSoft spots, in order of weight. First, circularity. The training set is labeled by Open-pFind, and the de novo evaluation in Fig. 5 and the '60% more PSMs' claim are computed by matching pUniFind's predictions to Open-pFind identifications. That is agreement with the annotator, not independent accuracy. The lack of any non-Open-pFind ground truth for the 21PTM set—no synthetic peptides, no external consensus—makes this a real problem, not a nitpick. It is somewhat mitigated by the fact that the database-search half has independent validation, and the modification-level accuracy does degrade for rare PTMs (which is what you would expect if the model learned annotator bias). But the central de novo claim as stated is not yet believably quantified.\n\nSecond, reproducibility: code is promised on acceptance, training data promised, no error bars, and a few baselines were excluded after preliminary checks (Open-pNovo). None of these are fatal, but they delay verification.\n\nThird, the target-decoy FDR for the rescored candidates is standard practice but not fully argued; I would want a supplementary note showing the score distributions.\n\nBottom line: this is a paper a serious proteomics journal should send to a referee, primarily because the architectural claim—unified open de novo plus rescoring—is new and the evaluation is extensive. But I would make acceptance conditional on public data/code and an independent de novo evaluation, e.g., synthetic peptide mixtures or at least a cross-engine consensus that does not include Open-pFind. If that lands, this could be a widely used tool.","headline":"A substantive advance in unified MS/MS interpretation, but the headline de novo gains are anchored to the same Open-pFind labels used for training, so they need independent verification.","tokens_in":13832,"tokens_out":2538,"would_cite":true,"duration_ms":27174,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"pUniFind unifies database search scoring and open de novo sequencing in one pretrained model trained on over 100 million spectra.","keywords":["mass spectrometry","peptide-spectrum matching","de novo sequencing","open search","multimodal pre-training","proteomics","deep learning","cross-modality prediction"],"falsifier":"Re-annotate or synthesize a benchmark where the true peptide sequence is known—for example, synthetic peptides or metabolic-labeling pairs—and compare pUniFind's de novo sequences and modification calls against that ground truth; if the 60% PSM gain and the 42.6% immunopeptidomics gain over baselines shrink or disappear under independent ground truth, the central claim of unified superiority is falsified.","tokens_in":12684,"feed_emoji":"🧬","tokens_out":6303,"duration_ms":61795,"temperature":0.7,"pith_summary":"pUniFind is a large pre-trained deep learning model that treats peptide identification and de novo peptide sequencing as a single multimodal problem, aligning spectra and peptide sequences through cross-modality prediction. The paper claims that, trained on over 100 million spectra annotated by open database search, it rescores database search candidates and performs open, zero-shot de novo sequencing over more than 1,300 modification types, identifying 60% more peptide-spectrum matches than existing de novo methods despite a 300-fold larger search space and 42.6% more peptides in immunopeptidomics. If correct, this would mean a single model can replace handcrafted, feature-based scoring and eliminate the need to prespecify modifications for modified-peptide sequencing. The paper also reports a deep-learning quality-control filter that recovers additional peptides, including 1,891 mapped in the human genome but absent from reference proteomes.","feed_headline":"One model rescores and sequences peptides from mass spectra","feed_subtitle":"Trained on 100M spectra, it outperforms classic engines and finds 60% more modified peptides.","key_machinery":"The machinery is cross-modality pre-training plus a joint scoring head. Separate encoders embed spectra and peptides; pre-training tasks include predicting the spectrum from a peptide, predicting the peptide length, amino-acid count and ion type (b/y, neutral losses) from each peak, and listwise candidate ranking. A joint modality scorer then produces the PSM score used to rerank Open-pFind's top-k candidates, with target-decoy analysis for FDR control. For de novo sequencing, a Peptide Length Aware module predicts a length within ±2 amino acids and generates the sequence token by token; a deep-learning feature filter with predicted spectra and retention times removes unreliable results, and for modification-enriched data a pFind search rescoring step is added.","core_discovery":"The central claim is that end-to-end deep learning, not feature engineering, is the right scoring framework for mass-spectrometry interpretation, and that database search and de novo sequencing share one underlying representation. pUniFind is presented as the first large-scale multimodal pre-trained model to integrate both tasks: it reranks Open-pFind candidate lists with a joint modality scorer and, in the same model, generates peptide sequences with modifications without being told which modifications to expect. The authors report consistent gains over existing engines across nine species, timsTOF and Astral instruments, metaproteomics, and immunopeptidomics, with accuracy checks through entrapment databases, metabolic labeling, and mixed-species searches. The reported 42.6% increase in immunopeptidomic peptide identifications and the 60% increase over de novo baselines are the concrete quantitative claims that carry the argument.","pith_inferences":["Because the training labels come from Open-pFind, the model may have learned Open-pFind's systematic blind spots; the reported de novo recall against Open-pFind labels could look better than against independent ground truth. That is my inference, not a claim in the paper.","The tokenized modification representation treats modification type without site; extending it with site prediction would likely improve site-level accuracy, which the paper already measures and which remains the hardest level.","The paper states that retention time is not used and DIA data are not fully exploited; adding either to the same cross-modal framework is the most direct next step and could widen the model's advantage on timsTOF and DIA datasets.","The same architecture could generalize to other molecule-spectrum matching problems, such as metabolomics or glycomics, where a spectrum must be aligned to a structured sequence with modifications."],"forward_implications":["Database search can drop handcrafted scoring features: rescoring candidates with the pretrained model improves peptide identifications, especially when open search expands the candidate space.","Modified-peptide de novo sequencing becomes feasible without prespecified modification lists, since modification types are treated as tokens the model can emit.","Hard applications benefit most: immunopeptidomics gains 42.6% more peptides than Open-pFind, and metaproteomics and non-tryptic searches also improve.","The model transfers to new instrument types with a single epoch of fine-tuning, as shown on timsTOF and Astral data.","A deep-learning QC filter can recover additional genome-derived peptides outside reference proteomes while keeping ion coverage, extending the reach of discovery proteomics."],"supporting_citations":[{"why":"Provides the open search annotations (100M PSMs) used to train pUniFind and serves as the retrieval kernel and main baseline.","marker":"[2]"},{"why":"MSFragger is a baseline open search engine that pUniFind outperforms, including MSBooster-enhanced results.","marker":"[4]"},{"why":"Defines the 300-fold expanded de novo search space and is the excluded low-performance baseline in 21PTM evaluation.","marker":"[5]"},{"why":"pNovo is the de novo baseline compared on 21PTM across modification, sequence, and site accuracy.","marker":"[6]"},{"why":"Casanovo V2 is the transformer de novo baseline that pUniFind outperforms on regular and immunopeptidomics data.","marker":"[12]"},{"why":"MSBooster supplies deep-learning features used by traditional scoring, the main non-end-to-end alternative compared.","marker":"[14]"},{"why":"Prosit established spectrum prediction, the peptide-side pre-training task pUniFind adapts.","marker":"[15]"},{"why":"Provides the 21PTM dataset of synthetic peptides with 21 post-translational modifications used to benchmark open de novo sequencing.","marker":"[17]"}],"fun_headline_variants":["One pretrained model does scoring and de novo sequencing from spectra","Unified AI ups peptide IDs by 60% over de novo sequencing","100M spectra train model that unifies search and de novo tasks","pUniFind: one model for peptide rescoring and zero-shot sequencing","Multimodal pretraining boosts mass spec interpretation and coverage"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Open-pFind annotations used to build the 100-million-spectrum training set are correct enough to serve as ground truth, and that target-decoy FDR estimation remains valid for pUniFind's scores; if the model learns the annotator's errors, the reported gains over Open-pFind and the de novo recall numbers measured against Open-pFind labels would be inflated.","fun_headline_variants_meta":{"raw":{"variants":["One pretrained model does scoring and de novo sequencing from spectra","Unified AI ups peptide IDs by 60% over de novo sequencing","100M spectra train model that unifies search and de novo tasks","pUniFind: one model for peptide rescoring and zero-shot sequencing","Multimodal pretraining boosts mass spec interpretation and coverage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000754,"raw_usage":{"total_tokens":3337,"prompt_tokens":909,"completion_tokens":2428,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":2337}},"tokens_in":525,"tokens_out":2428,"duration_ms":18478,"temperature":1.0,"reasoning_tokens":2337,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:36:28.119760+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate or synthesize a benchmark where the true peptide sequence is known—for example, synthetic peptides or metabolic-labeling pairs—and compare pUniFind's de novo sequences and modification calls against that ground truth; if the 60% PSM gain and the 42.6% immunopeptidomics gain over baselines shrink or disappear under independent ground truth, the central claim of unified superiority is falsified.","supporting_citations":[{"cited_title":"Nature biotechnology 36(11), 1059–1061 (2018)","cited_arxiv_id":null,"evidence_quote":"Provides the open search annotations (100M PSMs) used to train pUniFind and serves as the retrieval kernel and main baseline."},{"cited_title":"Nature methods 14(5), 513–520 (2017)","cited_arxiv_id":null,"evidence_quote":"MSFragger is a baseline open search engine that pUniFind outperforms, including MSBooster-enhanced results."},{"cited_title":"Journal of proteome research16(2), 645–654 (2017)","cited_arxiv_id":null,"evidence_quote":"Defines the 300-fold expanded de novo search space and is the excluded low-performance baseline in 21PTM evaluation."},{"cited_title":"Bioinformatics 35(14), 183–190 (2019)","cited_arxiv_id":null,"evidence_quote":"pNovo is the de novo baseline compared on 21PTM across modification, sequence, and site accuracy."},{"cited_title":"Nature communications 15(1), 6427 (2024)","cited_arxiv_id":null,"evidence_quote":"Casanovo V2 is the transformer de novo baseline that pUniFind outperforms on regular and immunopeptidomics data."},{"cited_title":"Nature Communications 14(1), 4539 (2023)","cited_arxiv_id":null,"evidence_quote":"MSBooster supplies deep-learning features used by traditional scoring, the main non-end-to-end alternative compared."},{"cited_title":"Nature methods 16(6), 509–518 (2019) 17","cited_arxiv_id":null,"evidence_quote":"Prosit established spectrum prediction, the peptide-side pre-training task pUniFind adapts."},{"cited_title":"Molecular & Cellular Proteomics 17(9), 1850–1863 (2018)","cited_arxiv_id":null,"evidence_quote":"Provides the 21PTM dataset of synthetic peptides with 21 post-translational modifications used to benchmark open de novo sequencing."}],"review_version":1}