REVIEW 4 major objections 5 minor 20 references
TumorHoPe2: An updated database for Tumor Homing Peptides
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read TumorHoPe2 reports 1,847 curated tumor-homing peptide entries in a public database with search, structure, and API tools.
desk verdict A genuinely useful database update with real new content, but the headline counts don't add up and the curation methods need to be spelled out before the 'experimentally validated' label can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the structured database entry: one peptide sequence linked to a source identifier, an experimental-validation record, a cancer cell line, and a tumor type, plus computed structural and physicochemical descriptors. The work is carried by a manual curation pipeline that extracts peptides from literature and patent sources, merges them with the legacy TumorHoPe set, and stores the result in a relational database with a web interface and a REST API. The same records power the search, browse, structure, and sequence-alignment tools, so the curation schema is what connects every reported count to a retrievable peptide.
What would settle it
Draw a random sample of entries and resolve each cited source identifier; if a substantial fraction of the sources do not actually report the peptide or do not report experimental validation of tumor homing, the database's counts and any predictions built on them would inherit the error.
Extended reading notes
Core claim
The central claim is that TumorHoPe2 constitutes a substantially expanded, manually curated repository of experimentally validated tumor-homing peptides. The authors assembled 487 new entries from the published literature and 652 from patent records, six of which overlap, and combined them with 744 entries carried over from TumorHoPe to reach 1,847 total entries and 1,297 unique peptide sequences. Each record carries the peptide sequence, terminal or chemical modifications, the cancer cell lines on which it was validated, and the tumor type it targets, with additional predicted secondary and tertiary structures, physicochemical properties, and motif annotations. The paper also reports that common homing motifs RGD, NGR, and CendR appear frequently in the curated set and that peptides of lengths 7, 9, 10, and 12 amino acids dominate.
Load-bearing premise
The headline counts and coverage figures rest on the assumption that every entry labeled 'experimentally validated' really was validated in the cited source and that manual curation applied that label consistently.
Editorial extensions
If this is right
- Researchers can retrieve nearly 2.5 times as many tumor-homing peptide entries from one platform as the previous version allowed.
- The REST API lets bioinformatics pipelines query the records programmatically, supporting automated screening and prediction workflows.
- The 172 cell lines and 37 tumor types give a broader experimental context for selecting peptides for targeted therapy and imaging studies.
- The inclusion of 594 chemically modified peptides, annotated with modification tags, makes engineered peptide data available for therapeutic design.
- Sequence, motif, structure, and physicochemical filters allow users to narrow the dataset by properties relevant to stability and binding.
Reading between the lines
- A natural extension the authors do not pursue is retraining machine-learning predictors of tumor-homing activity on this expanded set; the new distribution of sequences and modifications could change predictor performance, but the paper does not test this.
- Because the resource's value depends on the 'experimentally validated' label, a random audit that resolves each source identifier and confirms the reported validation would strengthen confidence in every downstream count.
- Coverage is uneven, with breast cancer and a few cell lines contributing a large share of entries, so cross-cancer comparisons should be interpreted with that imbalance in mind.
- The reported motif frequencies suggest that RGD, NGR, and CendR could serve as seeds for motif-based design tools, though the paper does not evaluate whether those motifs are sufficient or necessary for homing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. TumorHoPe2 is presented as a manually curated update of the TumorHoPe database, containing 1,847 entries and 1,297 unique tumor-homing peptides compiled from PubMed, Patent Lens, and the previous TumorHoPe version. The paper describes the database architecture, web interface, search/browse tools, REST API, and an analysis of peptide length, tumor type, cell line, modifications, motifs, and secondary-structure distributions. It also compares the updated resource with its 2012 predecessor, reporting increases in entries, unique peptides, cell lines, and tumor types. The central claim is that TumorHoPe2 is the most comprehensive current public resource for experimentally validated tumor-homing peptides.
Significance. If the counts and curation labels are correct, TumorHoPe2 is a valuable community resource: it consolidates data that were previously scattered, provides programmatic access via a REST API, and supplies structural and physicochemical annotations that can support downstream predictor development. The paper explicitly situates the resource within an established line of work (TumorHoPe, THPep, NEPTUNE, StackTHPred, LLM4THP, SCMTHP), which is appropriate for an update. The main strengths are the freely accessible web interface, the comparison with the previous version, and the inclusion of both PubMed and patent-derived entries. However, the manuscript currently contains multiple internal numerical inconsistencies and lacks a documented curation protocol, so the headline quantitative claims are not yet fully reproducible from the paper.
major comments (4)
- [Material and Methodology, Data collection; Results and Discussion, first paragraph] The entry counts are internally inconsistent. Data collection reports "487 from PubMed," while Results reports "457 from PubMed." Moreover, 487 + 652 + 744 - 6 = 1,877, not 1,847; the sum that yields 1,847 is 457 + 652 + 744 - 6. The Results sentence "bringing the total number of unique peptides in the updated database to 1,847" also contradicts the abstract and Table 2, which state 1,297 unique peptides. These are load-bearing numbers for a database paper, and every occurrence should be reconciled, with a clear distinction between total entries and unique sequences.
- [Data collection] The central premise that "All entries consist of manually curated, experimentally validated peptide sequences extracted from peer-reviewed articles and patents" is not supported by any curation protocol. There is no operational definition of "experimentally validated," no inclusion/exclusion criteria, no description of how Patent Lens records were screened, and no inter-annotator agreement or external audit. Because 652 of 1,847 entries come from Patent Lens, where patent disclosures often include predicted or broadly claimed peptides rather than demonstrated tumor-homing activity, the validation labels cannot be independently verified from the manuscript. The Limitation section only discusses tertiary-structure prediction and does not address curation provenance. Please provide a detailed protocol and, ideally, a mapping from each entry to its source record and the experimental evidence used.
- [Results and Discussion, cyclic/non-cyclic counts] The structural classification counts do not add up: the text states 1,789 non-cyclic, 58 cyclic, and 1 peptide with both linear and cyclic forms, which sums to 1,848, not the stated total of 1,847 entries. This is a concrete example of the counting inconsistencies that affect the paper's quantitative claims. I recommend auditing all aggregate distributions (tumor types, cell lines, terminal modifications, cyclic status, secondary-structure bins) against the downloadable dataset before resubmission.
- [Comparison with previous version; Results and Discussion] The relationship between the "1,103 newly curated" peptides and the "1,297 unique peptides" is not explained. Given the old database had 707 unique peptides, the new total of 1,297 unique peptides implies an undocumented overlap of 707 + 1,103 - 1,297 = 513 between the old unique set and the newly curated entries. The manuscript should state whether and how this overlap was identified and removed, since duplicate handling directly affects the uniqueness claim.
minor comments (5)
- [Table 2] The header uses "TumorHope" and "TumorHope2" instead of "TumorHoPe" and "TumorHoPe2"; please correct the spelling for consistency with the rest of the paper.
- [Figure 4 caption] The caption reads "Distribution of peptides validated in various cancer cell line" and should end with "cell lines" and a period.
- [Results and Discussion, first paragraph] The sentence "bringing the total number of unique peptides in the updated database to 1,847 The detailed length distribution" is missing a period after "1,847."
- [Comparison with previous version] The text says "more than 30 distinct tumors," while the abstract and Table 2 give a precise figure of 37; please use the precise number consistently.
- [Introduction, reference [2]] Reference [2] is about a large language model for pancreatic cancer transcriptomics and does not appear to support the statement about traditional chemotherapy and radiation therapy lacking selectivity; please verify the citation.
Circularity Check
No circularity: TumorHoPe2 is a curated data-resource paper with no fitted parameters, predictive derivations, or self-referential argument chain; its update legitimately reuses the earlier TumorHoPe database as one of three documented data sources.
full rationale
TumorHoPe2 is a database-compilation paper. It contains no equations, no fitted parameters, and no prediction that is subsequently tested. The central claim is that the database exists, is freely accessible, and contains manually curated, experimentally validated tumor-homing peptide entries compiled from PubMed, Patent Lens, and the previous TumorHoPe. There is no derivation chain in which an output is equivalent to an input by construction. The reuse of the authors' earlier TumorHoPe database (Kapoor et al., 2012) as one of three data sources is appropriate for an update and does not constitute load-bearing self-citation: the earlier database is contributing raw entries, not an unverified theorem or ansatz used to force a conclusion. The paper's 'experimentally validated' curation label is a data-quality premise; internal inconsistencies in the reported counts (487 vs. 457 PubMed entries; 1,847 total entries vs. 1,297 unique peptides; 707 old + 1,103 new requiring an undocumented 513-overlap) are accuracy and provenance concerns, not circularity. Similarly, the limitation passage about tertiary-structure prediction for complex modified peptides is a scope caveat, not a circular step. No circularity pattern enumerated in the rubric is present, so the appropriate score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption All entries labeled 'experimentally validated' are indeed tumor-homing peptides verified in experiments reported by the source papers and patents.
- domain assumption The deduplication process that yields 1,297 unique peptides correctly handles modified peptide variants and linear/cyclic forms.
- domain assumption Keyword searches for 'tumor-homing peptides' and 'tumor-targeting peptides' in PubMed and Patent Lens from 2012 to 2024 retrieve the relevant THP literature.
Cite this review
Pith. "Pith review of TumorHoPe2: An updated database for Tumor Homing Peptides." pith.science (2026). https://pith.science/paper/H4LBBTTW
@misc{pith2026250520913,
author = {Pith},
title = {Pith review of: TumorHoPe2: An updated database for Tumor Homing Peptides},
year = {2026},
howpublished = {\url{https://pith.science/paper/H4LBBTTW}},
note = {Machine review of arXiv:2505.20913}
}
read the original abstract
Addressing the growing need for organized data on tumor homing peptides (THPs), we present TumorHoPe2, a manually curated database offering extensive details on experimentally validated THPs. This represents a significant update to TumorHoPe, originally developed by our group in 2012. TumorHoPe2 now contains 1847 entries, representing 1297 unique tumor homing peptides, a substantial expansion from the 744 entries in its predecessor. For each peptide, the database provides critical information, including sequence, terminal or chemical modifications, corresponding cancer cell lines, and specific tumor types targeted. The database compiles data from two primary sources: phage display libraries, which are commonly used to identify peptide ligands targeting tumor-specific markers, and synthetic peptides, which are chemically modified to enhance properties such as stability, binding affinity, and specificity. Our dataset includes 594 chemically modified peptides, with 255 having N-terminal and 195 C-terminal modifications. These THPs have been validated against 172 cancer cell lines and demonstrate specificity for 37 distinct tumor types. To maximize utility for the research community, TumorHoPe2 is equipped with intuitive tools for data searching, filtering, and analysis, alongside a RESTful API for efficient programmatic access and integration into bioinformatics pipelines. It is freely available at https://webs.iiitd.edu.in/raghava/tumorhope2/
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
F. Bray, M. Laversanne, H. Sung, J. Ferlay, R.L. Siegel, I. Soerjomataram, A. Jemal, Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries, CA Cancer J. Clin. 74 (2024) 229–263
work page 2024
-
[2]
S. Choudhury, N.K. Mehta, G.P.S. Raghava, A large language model for predicting pancreatic ductal adenocarcinoma patients from blood-derived exosomal transcriptomics data, bioRxiv (2025). https://doi.org/10.1101/2025.03.06.641795
-
[3]
V.V. Veselov, A.E. Nosyrev, L. Jicsinszky, R.N. Alyautdin, G. Cravotto, Targeted delivery methods for anticancer drugs, Cancers (Basel) 14 (2022) 622
work page 2022
- [4]
- [5]
- [6]
- [7]
- [8]
Show all 20 references
-
[9]
Shoombuatong, N
W. Shoombuatong, N. Schaduangrat, R. Pratiwi, C. Nantasenamat, THPep: A machine learning-based approach for predicting tumor homing peptides, Comput. Biol. Chem. 80 (2019) 441–451
2019
-
[10]
R. Arif, S. Kanwal, S. Ahmed, M. Kabir, A computational predictor for accurate identification of tumor homing peptides by integrating sequential and deep BiLSTM features, Interdiscip. Sci. 16 (2024) 503–518
2024
-
[11]
Charoenkwan, N
P. Charoenkwan, N. Schaduangrat, P. Lio’, M.A. Moni, B. Manavalan, W. Shoombuatong, NEPTUNE: A novel computational approach for accurate and large-scale identification of tumor homing peptides, Comput. Biol. Med. 148 (2022) 105700
2022
-
[12]
J. Guan, L. Yao, C.-R. Chung, Y.-C. Chiang, T.-Y. Lee, StackTHPred: Identifying tumor- homing peptides through GBDT-based feature selection with stacking ensemble architecture, Int. J. Mol. Sci. 24 (2023) 10348
2023
-
[13]
S. Yang, P. Xu, LLM4THP: a computing tool to identify tumor homing peptides by molecular and sequence representation of large language model based on two-layer ensemble model strategy, Amino Acids 56 (2024) 62
2024
-
[14]
Charoenkwan, W
P. Charoenkwan, W. Chiangjong, C. Nantasenamat, M.A. Moni, P. Lio’, B. Manavalan, W. Shoombuatong, SCMTHP: A new approach for identifying and characterizing of tumor- homing peptides using estimated propensity scores of amino acids, Pharmaceutics 14 (2022) 122
2022
-
[15]
Gautam, P
A. Gautam, P. Kapoor, K. Chaudhary, R. Kumar, Open Source Drug Discovery Consortium, G.P.S. Raghava, Tumor homing peptides as molecular probes for cancer therapeutics, diagnostics and theranostics, Curr. Med. Chem. 21 (2014) 2367–2391
2014
-
[16]
Bhardwaj, V
A. Bhardwaj, V. Scaria, G.P.S. Raghava, A.M. Lynn, N. Chandra, S. Banerjee, M.V. Raghunandanan, V. Pandey, B. Taneja, J. Yadav, D. Dash, J. Bhattacharya, A. Misra, A. Kumar, S. Ramachandran, Z. Thomas, Open Source Drug Discovery Consortium, S.K. Brahmachari, Open source drug d...
2011
-
[17]
Rose, A.R
A.S. Rose, A.R. Bradley, Y. Valasatava, J.M. Duarte, A. Prlic, P.W. Rose, NGL viewer: web-based molecular graphics for large complexes, Bioinformatics 34 (2018) 3755–3758
2018
-
[18]
Altschul, W
S.F. Altschul, W. Gish, W. Miller, E.W. Myers, D.J. Lipman, Basic local alignment search tool, J. Mol. Biol. 215 (1990) 403–410
1990
-
[19]
Smith, M.S
T.F. Smith, M.S. Waterman, Identification of common molecular subsequences, J. Mol. Biol. 147 (1981) 195–197
1981
-
[20]
Shendre, N.K
A. Shendre, N.K. Mehta, A.S. Rathore, N. Kumar, S. Patiyal, G.P.S. Raghava, MAP format for representing chemical modifications, annotations, and mutations in protein sequences: An extension of the FASTA format, arXiv [q-bio.BM] (2025). https://doi.org/10.48550/ARXIV.2505.03403
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.