Pith. sign in

REVIEW 4 major objections 5 minor 20 references

TumorHoPe2: An updated database for Tumor Homing Peptides

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read TumorHoPe2 reports 1,847 curated tumor-homing peptide entries in a public database with search, structure, and API tools.

desk verdict A genuinely useful database update with real new content, but the headline counts don't add up and the curation methods need to be spelled out before the 'experimentally validated' label can be trusted. read the letter →

arxiv 2505.20913 v1 pith:H4LBBTTW submitted 2025-05-27 q-bio.BM

classification q-bio.BM
keywords Tumor-homingpeptidesTargeteddrugdeliveryPhagedisplayPeptidedatabaseCancertherapeuticsBioinformaticstoolsChemicallymodifiedRESTAPI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

TumorHoPe2 is a manually curated, freely accessible database of tumor-homing peptides, built by combining entries from the scientific literature, patent records, and the earlier TumorHoPe database. It expands the old resource from 744 entries to 1,847 entries, representing 1,297 unique peptide sequences, 594 chemically modified peptides, 172 cancer cell lines, and 37 tumor types. The paper's goal is to give cancer researchers one place to find experimentally validated peptides that bind tumors, together with their sequences, modifications, target cells, and predicted structures. If the curation label is accurate, this would be the largest public resource of its kind and a practical foundation for targeted drug delivery and imaging.

What carries the argument

The central object is the structured database entry: one peptide sequence linked to a source identifier, an experimental-validation record, a cancer cell line, and a tumor type, plus computed structural and physicochemical descriptors. The work is carried by a manual curation pipeline that extracts peptides from literature and patent sources, merges them with the legacy TumorHoPe set, and stores the result in a relational database with a web interface and a REST API. The same records power the search, browse, structure, and sequence-alignment tools, so the curation schema is what connects every reported count to a retrievable peptide.

What would settle it

Draw a random sample of entries and resolve each cited source identifier; if a substantial fraction of the sources do not actually report the peptide or do not report experimental validation of tumor homing, the database's counts and any predictions built on them would inherit the error.

Watch

Extended reading notes

Core claim

The central claim is that TumorHoPe2 constitutes a substantially expanded, manually curated repository of experimentally validated tumor-homing peptides. The authors assembled 487 new entries from the published literature and 652 from patent records, six of which overlap, and combined them with 744 entries carried over from TumorHoPe to reach 1,847 total entries and 1,297 unique peptide sequences. Each record carries the peptide sequence, terminal or chemical modifications, the cancer cell lines on which it was validated, and the tumor type it targets, with additional predicted secondary and tertiary structures, physicochemical properties, and motif annotations. The paper also reports that common homing motifs RGD, NGR, and CendR appear frequently in the curated set and that peptides of lengths 7, 9, 10, and 12 amino acids dominate.

Load-bearing premise

The headline counts and coverage figures rest on the assumption that every entry labeled 'experimentally validated' really was validated in the cited source and that manual curation applied that label consistently.

Editorial extensions

If this is right

  • Researchers can retrieve nearly 2.5 times as many tumor-homing peptide entries from one platform as the previous version allowed.
  • The REST API lets bioinformatics pipelines query the records programmatically, supporting automated screening and prediction workflows.
  • The 172 cell lines and 37 tumor types give a broader experimental context for selecting peptides for targeted therapy and imaging studies.
  • The inclusion of 594 chemically modified peptides, annotated with modification tags, makes engineered peptide data available for therapeutic design.
  • Sequence, motif, structure, and physicochemical filters allow users to narrow the dataset by properties relevant to stability and binding.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the authors do not pursue is retraining machine-learning predictors of tumor-homing activity on this expanded set; the new distribution of sequences and modifications could change predictor performance, but the paper does not test this.
  • Because the resource's value depends on the 'experimentally validated' label, a random audit that resolves each source identifier and confirms the reported validation would strengthen confidence in every downstream count.
  • Coverage is uneven, with breast cancer and a few cell lines contributing a large share of entries, so cross-cancer comparisons should be interpreted with that imbalance in mind.
  • The reported motif frequencies suggest that RGD, NGR, and CendR could serve as seeds for motif-based design tools, though the paper does not evaluate whether those motifs are sufficient or necessary for homing.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. TumorHoPe2 is presented as a manually curated update of the TumorHoPe database, containing 1,847 entries and 1,297 unique tumor-homing peptides compiled from PubMed, Patent Lens, and the previous TumorHoPe version. The paper describes the database architecture, web interface, search/browse tools, REST API, and an analysis of peptide length, tumor type, cell line, modifications, motifs, and secondary-structure distributions. It also compares the updated resource with its 2012 predecessor, reporting increases in entries, unique peptides, cell lines, and tumor types. The central claim is that TumorHoPe2 is the most comprehensive current public resource for experimentally validated tumor-homing peptides.

Significance. If the counts and curation labels are correct, TumorHoPe2 is a valuable community resource: it consolidates data that were previously scattered, provides programmatic access via a REST API, and supplies structural and physicochemical annotations that can support downstream predictor development. The paper explicitly situates the resource within an established line of work (TumorHoPe, THPep, NEPTUNE, StackTHPred, LLM4THP, SCMTHP), which is appropriate for an update. The main strengths are the freely accessible web interface, the comparison with the previous version, and the inclusion of both PubMed and patent-derived entries. However, the manuscript currently contains multiple internal numerical inconsistencies and lacks a documented curation protocol, so the headline quantitative claims are not yet fully reproducible from the paper.

major comments (4)
  1. [Material and Methodology, Data collection; Results and Discussion, first paragraph] The entry counts are internally inconsistent. Data collection reports "487 from PubMed," while Results reports "457 from PubMed." Moreover, 487 + 652 + 744 - 6 = 1,877, not 1,847; the sum that yields 1,847 is 457 + 652 + 744 - 6. The Results sentence "bringing the total number of unique peptides in the updated database to 1,847" also contradicts the abstract and Table 2, which state 1,297 unique peptides. These are load-bearing numbers for a database paper, and every occurrence should be reconciled, with a clear distinction between total entries and unique sequences.
  2. [Data collection] The central premise that "All entries consist of manually curated, experimentally validated peptide sequences extracted from peer-reviewed articles and patents" is not supported by any curation protocol. There is no operational definition of "experimentally validated," no inclusion/exclusion criteria, no description of how Patent Lens records were screened, and no inter-annotator agreement or external audit. Because 652 of 1,847 entries come from Patent Lens, where patent disclosures often include predicted or broadly claimed peptides rather than demonstrated tumor-homing activity, the validation labels cannot be independently verified from the manuscript. The Limitation section only discusses tertiary-structure prediction and does not address curation provenance. Please provide a detailed protocol and, ideally, a mapping from each entry to its source record and the experimental evidence used.
  3. [Results and Discussion, cyclic/non-cyclic counts] The structural classification counts do not add up: the text states 1,789 non-cyclic, 58 cyclic, and 1 peptide with both linear and cyclic forms, which sums to 1,848, not the stated total of 1,847 entries. This is a concrete example of the counting inconsistencies that affect the paper's quantitative claims. I recommend auditing all aggregate distributions (tumor types, cell lines, terminal modifications, cyclic status, secondary-structure bins) against the downloadable dataset before resubmission.
  4. [Comparison with previous version; Results and Discussion] The relationship between the "1,103 newly curated" peptides and the "1,297 unique peptides" is not explained. Given the old database had 707 unique peptides, the new total of 1,297 unique peptides implies an undocumented overlap of 707 + 1,103 - 1,297 = 513 between the old unique set and the newly curated entries. The manuscript should state whether and how this overlap was identified and removed, since duplicate handling directly affects the uniqueness claim.
minor comments (5)
  1. [Table 2] The header uses "TumorHope" and "TumorHope2" instead of "TumorHoPe" and "TumorHoPe2"; please correct the spelling for consistency with the rest of the paper.
  2. [Figure 4 caption] The caption reads "Distribution of peptides validated in various cancer cell line" and should end with "cell lines" and a period.
  3. [Results and Discussion, first paragraph] The sentence "bringing the total number of unique peptides in the updated database to 1,847 The detailed length distribution" is missing a period after "1,847."
  4. [Comparison with previous version] The text says "more than 30 distinct tumors," while the abstract and Table 2 give a precise figure of 37; please use the precise number consistently.
  5. [Introduction, reference [2]] Reference [2] is about a large language model for pancreatic cancer transcriptomics and does not appear to support the statement about traditional chemotherapy and radiation therapy lacking selectivity; please verify the citation.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: TumorHoPe2 is a curated data-resource paper with no fitted parameters, predictive derivations, or self-referential argument chain; its update legitimately reuses the earlier TumorHoPe database as one of three documented data sources.

full rationale

TumorHoPe2 is a database-compilation paper. It contains no equations, no fitted parameters, and no prediction that is subsequently tested. The central claim is that the database exists, is freely accessible, and contains manually curated, experimentally validated tumor-homing peptide entries compiled from PubMed, Patent Lens, and the previous TumorHoPe. There is no derivation chain in which an output is equivalent to an input by construction. The reuse of the authors' earlier TumorHoPe database (Kapoor et al., 2012) as one of three data sources is appropriate for an update and does not constitute load-bearing self-citation: the earlier database is contributing raw entries, not an unverified theorem or ansatz used to force a conclusion. The paper's 'experimentally validated' curation label is a data-quality premise; internal inconsistencies in the reported counts (487 vs. 457 PubMed entries; 1,847 total entries vs. 1,297 unique peptides; 707 old + 1,103 new requiring an undocumented 513-overlap) are accuracy and provenance concerns, not circularity. Similarly, the limitation passage about tertiary-structure prediction for complex modified peptides is a scope caveat, not a circular step. No circularity pattern enumerated in the rubric is present, so the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The database's claims rest primarily on curation assumptions rather than mathematical axioms. There are no fitted numerical parameters or invented entities; the key reliance is that the source literature is correctly indexed and that 'experimentally validated' classification is reliable.

assumptions (3)
  • domain assumption All entries labeled 'experimentally validated' are indeed tumor-homing peptides verified in experiments reported by the source papers and patents.
    The database's central value depends on accepting the curation team's assignment of the 'experimentally validated' label. Section 'Data collection' states all entries are manually curated, experimentally validated sequences, but no inclusion criteria or quality grading are defined.
  • domain assumption The deduplication process that yields 1,297 unique peptides correctly handles modified peptide variants and linear/cyclic forms.
    The database reports 1,847 entries but only 1,297 unique sequences; no explicit deduplication rule or sequence-normalization algorithm is described in the paper, yet the uniqueness count is a headline statistic.
  • domain assumption Keyword searches for 'tumor-homing peptides' and 'tumor-targeting peptides' in PubMed and Patent Lens from 2012 to 2024 retrieve the relevant THP literature.
    Section 'Data collection' describes this retrieval strategy. If relevant peptides are described with other terms or in other patent databases, the curated set may be incomplete without the paper acknowledging that scope.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TumorHoPe2: An updated database for Tumor Homing Peptides." pith.science (2026). https://pith.science/paper/H4LBBTTW

@misc{pith2026250520913,
  author       = {Pith},
  title        = {Pith review of: TumorHoPe2: An updated database for Tumor Homing Peptides},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H4LBBTTW}},
  note         = {Machine review of arXiv:2505.20913}
}
read the original abstract

Addressing the growing need for organized data on tumor homing peptides (THPs), we present TumorHoPe2, a manually curated database offering extensive details on experimentally validated THPs. This represents a significant update to TumorHoPe, originally developed by our group in 2012. TumorHoPe2 now contains 1847 entries, representing 1297 unique tumor homing peptides, a substantial expansion from the 744 entries in its predecessor. For each peptide, the database provides critical information, including sequence, terminal or chemical modifications, corresponding cancer cell lines, and specific tumor types targeted. The database compiles data from two primary sources: phage display libraries, which are commonly used to identify peptide ligands targeting tumor-specific markers, and synthetic peptides, which are chemically modified to enhance properties such as stability, binding affinity, and specificity. Our dataset includes 594 chemically modified peptides, with 255 having N-terminal and 195 C-terminal modifications. These THPs have been validated against 172 cancer cell lines and demonstrate specificity for 37 distinct tumor types. To maximize utility for the research community, TumorHoPe2 is equipped with intuitive tools for data searching, filtering, and analysis, alongside a RESTful API for efficient programmatic access and integration into bioinformatics pipelines. It is freely available at https://webs.iiitd.edu.in/raghava/tumorhope2/

Figures

Figures reproduced from arXiv: 2505.20913 by the authors.

Figure 1
Figure 1. Architecture of TumorHoPe2 database. Web interface and tools implementation Search Tools TumorHoPe2 integrates a suite of enhanced search tools designed to enable efficient and flexible retrieval of tumor-homing peptide data. Users can explore the database through keyword, advanced, and peptide-based search options. For general exploration, the keyword search allows simple input of any relevant term or phrase to sca… view at source ↗
Figure 2
Figure 2. Length wise distribution of the number of unique tumor-homing peptides. The distribution shows that peptides of lengths 7, 9, 10, and 12 amino acids were the most frequent, with 9-residue peptides being the most abundant. Peptides with lengths greater than 20 amino acids were relatively rare. The newly curated dataset spans a diverse range of cancer types, with a significant proportion of peptides associated with br… view at source ↗
Figure 3
Figure 3. Distribution of number of entries of peptide across different cancer types. The updated dataset includes peptides validated in various cancer cell lines, with the highest number of peptides associated with MDA-MB-435 (292 peptides), 4T1 (233), B16F1 (214), PPC1 (206), and PC-3 (209). Additionally, a considerable number of peptides were tested on MCF-7 (146), HeLa (71) and A549 (55) cell lines. The distribution of ca… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Distribution of peptides validated in various cancer cell line [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Distribution of the number of entries of cyclic and non-cyclic peptides Searching for relevant peptides has also become easier, with an improved interface that supports keyword, peptide, and advanced search options, making data retrieval more intuitive and efficient. A…
Figure 6
Figure 6. Figure 6: Visual representation highlighting the advancements in TumorHoPe2.0 The complete comparison entries between TumorHoPe and TumorHoPe2 are provided in [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

20 extracted references · 20 canonical work pages

  1. [1]

    F. Bray, M. Laversanne, H. Sung, J. Ferlay, R.L. Siegel, I. Soerjomataram, A. Jemal, Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries, CA Cancer J. Clin. 74 (2024) 229–263

  2. [2]

    Choudhury, N.K

    S. Choudhury, N.K. Mehta, G.P.S. Raghava, A large language model for predicting pancreatic ductal adenocarcinoma patients from blood-derived exosomal transcriptomics data, bioRxiv (2025). https://doi.org/10.1101/2025.03.06.641795

  3. [3]

    Veselov, A.E

    V.V. Veselov, A.E. Nosyrev, L. Jicsinszky, R.N. Alyautdin, G. Cravotto, Targeted delivery methods for anticancer drugs, Cancers (Basel) 14 (2022) 622

  4. [4]

    Sharma, P

    A. Sharma, P. Kapoor, A. Gautam, K. Chaudhary, R. Kumar, J.S. Chauhan, A. Tyagi, G.P.S. Raghava, Computational approach for designing tumor homing peptides, Sci. Rep. 3 (2013) 1607

  5. [5]

    Kapoor, H

    P. Kapoor, H. Singh, A. Gautam, K. Chaudhary, R. Kumar, G.P.S. Raghava, TumorHoPe: a database of tumor homing peptides, PLoS One 7 (2012) e35187

  6. [6]

    Gupta, K

    S. Gupta, K. Chaudhary, R. Kumar, A. Gautam, J.S. Nanda, S.K. Dhanda, S.K. Brahmachari, G.P.S. Raghava, Prioritization of anticancer drugs against a cancer using genomic features of cancer cells: A step towards personalized medicine, Sci. Rep. 6 (2016) 23857

  7. [7]

    Singh, S

    H. Singh, S. Singh, D. Singla, S.M. Agarwal, G.P.S. Raghava, QSAR based model for discriminating EGFR inhibitors and non-inhibitors using Random forest, Biol. Direct 10 (2015) 10

  8. [8]

    Gautam, K

    A. Gautam, K. Chaudhary, R. Kumar, G.P.S. Raghava, Computer-aided virtual screening and designing of cell-penetrating peptides, Methods Mol. Biol. 1324 (2015) 59–69

Show all 20 references
  1. [9]

    Shoombuatong, N

    W. Shoombuatong, N. Schaduangrat, R. Pratiwi, C. Nantasenamat, THPep: A machine learning-based approach for predicting tumor homing peptides, Comput. Biol. Chem. 80 (2019) 441–451

  2. [10]

    R. Arif, S. Kanwal, S. Ahmed, M. Kabir, A computational predictor for accurate identification of tumor homing peptides by integrating sequential and deep BiLSTM features, Interdiscip. Sci. 16 (2024) 503–518

  3. [11]

    Charoenkwan, N

    P. Charoenkwan, N. Schaduangrat, P. Lio’, M.A. Moni, B. Manavalan, W. Shoombuatong, NEPTUNE: A novel computational approach for accurate and large-scale identification of tumor homing peptides, Comput. Biol. Med. 148 (2022) 105700

  4. [12]

    J. Guan, L. Yao, C.-R. Chung, Y.-C. Chiang, T.-Y. Lee, StackTHPred: Identifying tumor- homing peptides through GBDT-based feature selection with stacking ensemble architecture, Int. J. Mol. Sci. 24 (2023) 10348

  5. [13]

    S. Yang, P. Xu, LLM4THP: a computing tool to identify tumor homing peptides by molecular and sequence representation of large language model based on two-layer ensemble model strategy, Amino Acids 56 (2024) 62

  6. [14]

    Charoenkwan, W

    P. Charoenkwan, W. Chiangjong, C. Nantasenamat, M.A. Moni, P. Lio’, B. Manavalan, W. Shoombuatong, SCMTHP: A new approach for identifying and characterizing of tumor- homing peptides using estimated propensity scores of amino acids, Pharmaceutics 14 (2022) 122

  7. [15]

    Gautam, P

    A. Gautam, P. Kapoor, K. Chaudhary, R. Kumar, Open Source Drug Discovery Consortium, G.P.S. Raghava, Tumor homing peptides as molecular probes for cancer therapeutics, diagnostics and theranostics, Curr. Med. Chem. 21 (2014) 2367–2391

  8. [16]

    Bhardwaj, V

    A. Bhardwaj, V. Scaria, G.P.S. Raghava, A.M. Lynn, N. Chandra, S. Banerjee, M.V. Raghunandanan, V. Pandey, B. Taneja, J. Yadav, D. Dash, J. Bhattacharya, A. Misra, A. Kumar, S. Ramachandran, Z. Thomas, Open Source Drug Discovery Consortium, S.K. Brahmachari, Open source drug d...

  9. [17]

    Rose, A.R

    A.S. Rose, A.R. Bradley, Y. Valasatava, J.M. Duarte, A. Prlic, P.W. Rose, NGL viewer: web-based molecular graphics for large complexes, Bioinformatics 34 (2018) 3755–3758

  10. [18]

    Altschul, W

    S.F. Altschul, W. Gish, W. Miller, E.W. Myers, D.J. Lipman, Basic local alignment search tool, J. Mol. Biol. 215 (1990) 403–410

  11. [19]

    Smith, M.S

    T.F. Smith, M.S. Waterman, Identification of common molecular subsequences, J. Mol. Biol. 147 (1981) 195–197

  12. [20]

    Shendre, N.K

    A. Shendre, N.K. Mehta, A.S. Rathore, N. Kumar, S. Patiyal, G.P.S. Raghava, MAP format for representing chemical modifications, annotations, and mutations in protein sequences: An extension of the FASTA format, arXiv [q-bio.BM] (2025). https://doi.org/10.48550/ARXIV.2505.03403

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.