REVIEW 2 major objections 5 minor 17 references
A public multi-magnification image set now supplies WHO three-tier grades for colorectal adenocarcinoma so machines can learn histologic grading.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-15 03:37 UTC pith:ETNX3ET2
load-bearing objection Solid public multi-mag WHO-grade CRC resource that fills a real gap; label noise and missing IRB/scanner metadata are the main soft spots, not the claim itself. the 2 major comments →
CRC-HGD: A Histopathological Image Dataset for Grading Colorectal Cancer
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
No publicly available histopathological microscopy image dataset specifically targets the WHO three-tier grading of colorectal adenocarcinoma with multiple magnification levels per specimen; CRC-HGD fills that gap with 1,914 labeled 800×800 RGB images from 214 patients (Grade I 106, Grade II 75, Grade III 33) at 4×, 10×, 20× and 40×.
What carries the argument
The CRC-HGD dataset: patient- and grade-organized JPEG images at four magnifications, labeled by WHO gland-formation criteria (well-, moderately-, and poorly-differentiated), which supplies the supervised multi-scale resource required for training and benchmarking automated grade classifiers.
Load-bearing premise
The single-source WHO grade labels on the archived slides are treated as correct ground truth even though no multi-reader agreement statistics are reported, so every model inherits whatever noise those labels contain.
What would settle it
An independent multi-pathologist re-grading of a random subset of the released images that shows low concordance with the published labels, or the documented existence of an earlier public dataset that already supplies WHO three-tier grades at multiple magnifications per colorectal adenocarcinoma specimen.
If this is right
- Supervised models can be trained and scored directly on WHO Grade I/II/III labels rather than on tissue-type proxies.
- Multi-magnification training becomes routine, letting algorithms fuse low-power architecture with high-power nuclear detail.
- Public release enables reproducible head-to-head comparison of grading algorithms across groups.
- Normal and mixed slides supply contrast for tumor-boundary and normal-versus-tumor tasks.
- Systems that flag high-grade (Grade III) cases for aggressive therapy can be developed and validated on this resource.
Where Pith is reading between the lines
- Severe class imbalance (fewest images in Grade III) will force oversampling or cost-sensitive losses before models reach clinically useful sensitivity for the highest-risk tumors.
- Single-center origin implies domain-shift tests against other scanners and stain protocols will be a necessary next experiment for any model trained here.
- The four-magnification design invites explicit multi-resolution fusion networks that may transfer to grading other gland-forming carcinomas.
- If later multi-reader studies reveal substantial label noise, the set remains useful as a controlled study of noisy supervision in computational pathology.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces CRC-HGD, a publicly released histopathological image dataset of 1,914 RGB JPEG images (800 imes800, 96 DPI) from 214 colorectal adenocarcinoma patients (Grade I: 106, Grade II: 75, Grade III: 33) diagnosed 2014–2019 at a single Iranian center. Specimens are H&E-stained and labeled by WHO three-tier criteria (well-/moderately-/poorly-differentiated) based on gland-formation proportion; each specimen is imaged at four magnifications (4×, 10×, 20×, 40×). The authors survey existing CRC histopathology resources (NCT-CRC-HE-100K, Kather texture, EBHI-Seg, and two non-public grading sets) and argue that none simultaneously supplies WHO grade labels and multi-magnification views of the same specimens. Tables 1–3, a naming convention, representative figures, and public DOIs (Mendeley + databiox.com) document the resource; a short “Value of the Data” section lists intended uses for supervised multi-scale grading models.
Significance. If the gap claim holds, CRC-HGD supplies a missing public benchmark for automated WHO-grade classification of colorectal adenocarcinoma—an independent prognostic factor that currently suffers from inter-observer variability. The multi-magnification design is a genuine practical strength for multi-scale CNN/transformer work, and the open Mendeley DOI plus stratified counts make the resource immediately usable. The contribution is therefore a solid, if incremental, data-release paper rather than a methodological advance; its long-term value will depend on community uptake and on whether later users can quantify or mitigate the label noise inherent in single-center grading.
major comments (2)
- Section 3 (labeling paragraph) and the five-step preparation list treat the WHO grade labels as ground truth without any reported inter-observer agreement (κ, percentage concordance) or multi-reader adjudication. The introduction itself cites substantial inter-observer variability for CRC grading; without a reliability statistic the labels remain single-center, single-team annotations whose noise will be inherited by every model trained on CRC-HGD. At minimum the authors should state how many pathologists graded each case, whether consensus was required, and any available agreement numbers; if none exist, this limitation must be stated explicitly in Sections 4–5.
- Table 1 and Section 3 omit all acquisition metadata (microscope/scanner model, camera, white-balance or color-calibration protocol, physical pixel size). Because the images are fixed 800×800 JPEGs rather than whole-slide images, downstream users cannot assess stain/color variation or perform stain normalization without this information; it is load-bearing for any claim of “comprehensive computational analysis.”
minor comments (5)
- Table 3 and the accompanying text list 7 “mixed normal-tumoral” images but give no per-magnification breakdown (shown as dashes); either supply the counts or explain why they are omitted.
- No IRB/ethics-approval or patient-consent statement appears; for a clinical archive dataset this is standard documentation and should be added.
- Class imbalance is severe (Grade III: only 33 patients / 327 images). While not fatal for a resource paper, a brief note on recommended stratified splits or re-weighting strategies would help users.
- Figure 2 caption and the text refer to “Figure 2” before any earlier figures are numbered; renumber or reorder for sequential presentation.
- Reference [17] (authors’ prior breast-cancer dataset) is cited only for context; ensure the citation style is consistent with the rest of the list.
Circularity Check
No circularity: CRC-HGD is a dataset-release paper with no derivation, prediction, or load-bearing self-citation chain.
full rationale
The manuscript is a resource paper that introduces a labeled histopathological image collection (1,914 images from 214 patients, WHO three-tier grades, four magnifications). Its central claim is empirical and descriptive: that no prior public dataset specifically targets WHO Grade I/II/III colorectal adenocarcinoma with multi-magnification images per specimen, and that CRC-HGD fills that gap. Section 2 surveys external datasets (NCT-CRC-HE-100K, Kather_texture_2016, EBHI-Seg, and two non-public institution-specific collections) and correctly notes they lack the WHO three-tier multi-magnification structure; none of those citations are by the present authors and none are used to force the CRC-HGD labels or counts. The single self-citation ([17], the authors’ prior breast-cancer grading dataset) appears only in the reference list and author biographies as contextual precedent; it is not invoked to define grades, justify uniqueness, or derive any numerical claim about CRC-HGD. Labeling follows the external WHO criteria (gland-formation thresholds), which are independent of the authors. There are no equations, fitted parameters, uniqueness theorems, or “predictions” that could reduce to inputs by construction. Consequently the derivation chain is empty and circularity score is zero.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption WHO three-tier grading thresholds (>95 %, 50–95 %, <50 % gland formation) correctly define Grade I/II/III for colorectal adenocarcinoma.
- domain assumption Archival H&E slides from 2014–2019 at one Iranian center, once re-photographed, constitute a representative and usable training resource.
- ad hoc to paper JPEG 800×800 RGB images at 96 DPI preserve the morphological features needed for grade classification.
read the original abstract
Colorectal cancer (CRC) is the third most common cancer worldwide and the second leading cause of cancer-related deaths globally, with approximately 1,926,425 new cases and 904,019 deaths reported in 2022. Accurate histologic grading plays a critical role in prognosis and treatment planning for colorectal adenocarcinoma. In recent years, artificial intelligence and its subcategories, including machine learning and deep learning, have been increasingly employed for automated cancer detection and classification. An appropriate and well-organized dataset is the essential first step to achieve this goal. This paper introduces CRC-HGD, a histopathological microscopy image dataset of 1,914 images obtained from 214 colorectal adenocarcinoma patients (Grade I: 106, Grade II: 75, Grade III: 33). The specimens are H&E-stained colorectal tissue sections acquired at the Poursina Hakim Research Center of Isfahan University of Medical Sciences, Iran, diagnosed between 2014 and 2019, and graded according to the World Health Organization (WHO) criteria into three grades: well-differentiated (Grade I), moderately differentiated (Grade II), and poorly differentiated (Grade III). For each specimen, four magnification levels are provided: 4x, 10x, 20x, and 40x. The dataset is accessible via Mendeley Data (https://doi.org/10.17632/yfp5sfj47m.4) and at http://databiox.com, where the latest version is also available. The distinctive feature of this dataset is the provision of labeled specimens across all three differentiation grades at multiple magnification levels, enabling comprehensive computational analysis of colorectal cancer grading.
Figures
Reference graph
Works this paper leans on
-
[1]
Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries
Bray F, Laversanne M, Sung H, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229 –263
2022
-
[2]
Global colorectal cancer burden in 2022 and projections to 2050: incidence and mortality estimates from GLOBOCAN
Xi Y, Xu P. Global colorectal cancer burden in 2022 and projections to 2050: incidence and mortality estimates from GLOBOCAN. Cancer Epidemiol. 2023;87:102–155
2022
-
[3]
Global, regional, and national burden of colorectal cancer and its attributable risk factors in 204 countries and territories, 1990–2021
GBD 2021 Colorectal Cancer Collaborators. Global, regional, and national burden of colorectal cancer and its attributable risk factors in 204 countries and territories, 1990–2021. Lancet Gastroenterol Hepatol. 2024
2021
-
[4]
Colorectal carcinoma: pathologic aspects
Fleming M, Ravula S, Tatishchev SF, Wang HL. Colorectal carcinoma: pathologic aspects. J Gastrointest Oncol. 2012;3(3):153–173
2012
-
[5]
Prognostic factors in colorectal cancer: College of American Pathologists Consensus Statement 1999
Compton CC, Fielding LP, Burgart LJ, et al. Prognostic factors in colorectal cancer: College of American Pathologists Consensus Statement 1999. Arch Pathol Lab Med. 2000;124(7):979–994
1999
-
[6]
Has the new TNM classification for colorectal cancer improved care? Nat Rev Clin Oncol
Nagtegaal ID, Quirke P, Schmoll HJ. Has the new TNM classification for colorectal cancer improved care? Nat Rev Clin Oncol. 2012;9(2):119–123
2012
-
[7]
Deep learning models for poorly differentiated colorectal adenocarcinoma classification in whole slide images
Yoshida H, et al. Deep learning models for poorly differentiated colorectal adenocarcinoma classification in whole slide images. Sci Rep. 2021;11:23041
2021
-
[8]
Association between molecular subtypes of colorectal tumors and patient survival
Phipps AI, Alwers E, Harrison T, et al. Association between molecular subtypes of colorectal tumors and patient survival. Gastroenterology. 2020;158(8):2158–2168
2020
-
[9]
WHO Classification of Tumours of the Digestive System
Nagtegaal ID, et al. WHO Classification of Tumours of the Digestive System. 5th ed. Lyon: IARC Press; 2019
2019
-
[10]
100,000 histological images of human colorectal cancer and healthy tissue
Kather JN, Halama N, Marx A. 100,000 histological images of human colorectal cancer and healthy tissue. Zenodo. 2018. https://doi.org/10.5281/zenodo.1214456
-
[11]
Multi -class texture analysis in colorectal cancer histology
Kather JN, Weis CA, Bianconi F, et al. Multi -class texture analysis in colorectal cancer histology. Sci Rep. 2016;6:27988
2016
-
[12]
Color-CADx: a deep learning approach for colorectal cancer classification
Sharkas M, Attallah O. Color-CADx: a deep learning approach for colorectal cancer classification. PeerJ Comput Sci. 2024;10:e1866
2024
-
[13]
CRCCN-Net: automated framework for classification of colorectal tissue using histopathological images
Alkassar S, et al. CRCCN-Net: automated framework for classification of colorectal tissue using histopathological images. Biomed Signal Process Control. 2022;78:104002
2022
-
[14]
EBHI -Seg: a novel enteroscopy biopsy histological image dataset for image segmentation tasks
Shi W, et al. EBHI -Seg: a novel enteroscopy biopsy histological image dataset for image segmentation tasks. Front Med. 2023;10:1114673
2023
-
[15]
HCCANet: histopathological image grading of colorectal cancer using CNN
Liang Y, et al. HCCANet: histopathological image grading of colorectal cancer using CNN. Sci Rep. 2022;12:15103
2022
-
[16]
Deep learning models for poorly differentiated colorectal adenocarcinoma classification
Yoshida H, et al. Deep learning models for poorly differentiated colorectal adenocarcinoma classification. Sci Rep. 2021;11:23041
2021
-
[17]
A histopathological image dataset for grading breast invasive ductal carcinomas
Bolhasani H, Amjadi E, Tabatabaeian M, Jassbi SJ. A histopathological image dataset for grading breast invasive ductal carcinomas. Informatics Med Unlocked. 2020;19:100341. Author Biographies Elham Amjadi, MD is a Pathologist with expertise in gastrointestinal and breast pathology. She received her medical degree and completed her residency in pathology a...
2020
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.