Pith. sign in

REVIEW 2 major objections 5 minor 17 references

A public multi-magnification image set now supplies WHO three-tier grades for colorectal adenocarcinoma so machines can learn histologic grading.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-15 03:37 UTC pith:ETNX3ET2

load-bearing objection Solid public multi-mag WHO-grade CRC resource that fills a real gap; label noise and missing IRB/scanner metadata are the main soft spots, not the claim itself. the 2 major comments →

arxiv 2607.12750 v1 pith:ETNX3ET2 submitted 2026-07-14 cs.CV

CRC-HGD: A Histopathological Image Dataset for Grading Colorectal Cancer

classification cs.CV
keywords colorectal canceradenocarcinomahistopathologydigital pathologygradingimage datasetH&E staining
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper establishes that a labeled collection of 1,914 H&E microscopy images from 214 colorectal adenocarcinoma patients, spanning WHO Grades I–III and four standard magnifications, is now openly available. Existing public CRC histology resources cover tissue types or binary malignancy but omit the three-tier differentiation grade that drives prognosis and treatment. By releasing multi-scale, grade-labeled images plus a few normal and mixed slides, the work gives supervised algorithms a direct training and evaluation target for automated grading. A reader cares because histologic grade is an independent prognostic factor; faster, more consistent identification of poorly differentiated tumors can influence chemotherapy choices and monitoring.

Core claim

No publicly available histopathological microscopy image dataset specifically targets the WHO three-tier grading of colorectal adenocarcinoma with multiple magnification levels per specimen; CRC-HGD fills that gap with 1,914 labeled 800×800 RGB images from 214 patients (Grade I 106, Grade II 75, Grade III 33) at 4×, 10×, 20× and 40×.

What carries the argument

The CRC-HGD dataset: patient- and grade-organized JPEG images at four magnifications, labeled by WHO gland-formation criteria (well-, moderately-, and poorly-differentiated), which supplies the supervised multi-scale resource required for training and benchmarking automated grade classifiers.

Load-bearing premise

The single-source WHO grade labels on the archived slides are treated as correct ground truth even though no multi-reader agreement statistics are reported, so every model inherits whatever noise those labels contain.

What would settle it

An independent multi-pathologist re-grading of a random subset of the released images that shows low concordance with the published labels, or the documented existence of an earlier public dataset that already supplies WHO three-tier grades at multiple magnifications per colorectal adenocarcinoma specimen.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Supervised models can be trained and scored directly on WHO Grade I/II/III labels rather than on tissue-type proxies.
  • Multi-magnification training becomes routine, letting algorithms fuse low-power architecture with high-power nuclear detail.
  • Public release enables reproducible head-to-head comparison of grading algorithms across groups.
  • Normal and mixed slides supply contrast for tumor-boundary and normal-versus-tumor tasks.
  • Systems that flag high-grade (Grade III) cases for aggressive therapy can be developed and validated on this resource.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Severe class imbalance (fewest images in Grade III) will force oversampling or cost-sensitive losses before models reach clinically useful sensitivity for the highest-risk tumors.
  • Single-center origin implies domain-shift tests against other scanners and stain protocols will be a necessary next experiment for any model trained here.
  • The four-magnification design invites explicit multi-resolution fusion networks that may transfer to grading other gland-forming carcinomas.
  • If later multi-reader studies reveal substantial label noise, the set remains useful as a controlled study of noisy supervision in computational pathology.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The manuscript introduces CRC-HGD, a publicly released histopathological image dataset of 1,914 RGB JPEG images (800 imes800, 96 DPI) from 214 colorectal adenocarcinoma patients (Grade I: 106, Grade II: 75, Grade III: 33) diagnosed 2014–2019 at a single Iranian center. Specimens are H&E-stained and labeled by WHO three-tier criteria (well-/moderately-/poorly-differentiated) based on gland-formation proportion; each specimen is imaged at four magnifications (4×, 10×, 20×, 40×). The authors survey existing CRC histopathology resources (NCT-CRC-HE-100K, Kather texture, EBHI-Seg, and two non-public grading sets) and argue that none simultaneously supplies WHO grade labels and multi-magnification views of the same specimens. Tables 1–3, a naming convention, representative figures, and public DOIs (Mendeley + databiox.com) document the resource; a short “Value of the Data” section lists intended uses for supervised multi-scale grading models.

Significance. If the gap claim holds, CRC-HGD supplies a missing public benchmark for automated WHO-grade classification of colorectal adenocarcinoma—an independent prognostic factor that currently suffers from inter-observer variability. The multi-magnification design is a genuine practical strength for multi-scale CNN/transformer work, and the open Mendeley DOI plus stratified counts make the resource immediately usable. The contribution is therefore a solid, if incremental, data-release paper rather than a methodological advance; its long-term value will depend on community uptake and on whether later users can quantify or mitigate the label noise inherent in single-center grading.

major comments (2)
  1. Section 3 (labeling paragraph) and the five-step preparation list treat the WHO grade labels as ground truth without any reported inter-observer agreement (κ, percentage concordance) or multi-reader adjudication. The introduction itself cites substantial inter-observer variability for CRC grading; without a reliability statistic the labels remain single-center, single-team annotations whose noise will be inherited by every model trained on CRC-HGD. At minimum the authors should state how many pathologists graded each case, whether consensus was required, and any available agreement numbers; if none exist, this limitation must be stated explicitly in Sections 4–5.
  2. Table 1 and Section 3 omit all acquisition metadata (microscope/scanner model, camera, white-balance or color-calibration protocol, physical pixel size). Because the images are fixed 800×800 JPEGs rather than whole-slide images, downstream users cannot assess stain/color variation or perform stain normalization without this information; it is load-bearing for any claim of “comprehensive computational analysis.”
minor comments (5)
  1. Table 3 and the accompanying text list 7 “mixed normal-tumoral” images but give no per-magnification breakdown (shown as dashes); either supply the counts or explain why they are omitted.
  2. No IRB/ethics-approval or patient-consent statement appears; for a clinical archive dataset this is standard documentation and should be added.
  3. Class imbalance is severe (Grade III: only 33 patients / 327 images). While not fatal for a resource paper, a brief note on recommended stratified splits or re-weighting strategies would help users.
  4. Figure 2 caption and the text refer to “Figure 2” before any earlier figures are numbered; renumber or reorder for sequential presentation.
  5. Reference [17] (authors’ prior breast-cancer dataset) is cited only for context; ensure the citation style is consistent with the rest of the list.

Circularity Check

0 steps flagged

No circularity: CRC-HGD is a dataset-release paper with no derivation, prediction, or load-bearing self-citation chain.

full rationale

The manuscript is a resource paper that introduces a labeled histopathological image collection (1,914 images from 214 patients, WHO three-tier grades, four magnifications). Its central claim is empirical and descriptive: that no prior public dataset specifically targets WHO Grade I/II/III colorectal adenocarcinoma with multi-magnification images per specimen, and that CRC-HGD fills that gap. Section 2 surveys external datasets (NCT-CRC-HE-100K, Kather_texture_2016, EBHI-Seg, and two non-public institution-specific collections) and correctly notes they lack the WHO three-tier multi-magnification structure; none of those citations are by the present authors and none are used to force the CRC-HGD labels or counts. The single self-citation ([17], the authors’ prior breast-cancer grading dataset) appears only in the reference list and author biographies as contextual precedent; it is not invoked to define grades, justify uniqueness, or derive any numerical claim about CRC-HGD. Labeling follows the external WHO criteria (gland-formation thresholds), which are independent of the authors. There are no equations, fitted parameters, uniqueness theorems, or “predictions” that could reduce to inputs by construction. Consequently the derivation chain is empty and circularity score is zero.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

As a dataset paper the central claim rests almost entirely on standard domain conventions (WHO grading thresholds, H&E staining) and on the authors’ own archival collection process. No free parameters are fitted; no new physical or mathematical entities are postulated. The only non-standard premise is that the single-center labels are sufficiently accurate to serve as public ground truth.

axioms (3)
  • domain assumption WHO three-tier grading thresholds (>95 %, 50–95 %, <50 % gland formation) correctly define Grade I/II/III for colorectal adenocarcinoma.
    Invoked throughout Section 3 and the abstract as the sole labeling rule; taken from references [4,5] without re-derivation.
  • domain assumption Archival H&E slides from 2014–2019 at one Iranian center, once re-photographed, constitute a representative and usable training resource.
    Stated in Section 3 patient-selection paragraph; no multi-center or multi-scanner validation is supplied.
  • ad hoc to paper JPEG 800×800 RGB images at 96 DPI preserve the morphological features needed for grade classification.
    Technical specification chosen by the authors (Table 1); no ablation of resolution or compression is reported.

pith-pipeline@v1.1.0-grok45 · 12148 in / 2527 out tokens · 26707 ms · 2026-07-15T03:37:46.426982+00:00 · methodology

0 comments
read the original abstract

Colorectal cancer (CRC) is the third most common cancer worldwide and the second leading cause of cancer-related deaths globally, with approximately 1,926,425 new cases and 904,019 deaths reported in 2022. Accurate histologic grading plays a critical role in prognosis and treatment planning for colorectal adenocarcinoma. In recent years, artificial intelligence and its subcategories, including machine learning and deep learning, have been increasingly employed for automated cancer detection and classification. An appropriate and well-organized dataset is the essential first step to achieve this goal. This paper introduces CRC-HGD, a histopathological microscopy image dataset of 1,914 images obtained from 214 colorectal adenocarcinoma patients (Grade I: 106, Grade II: 75, Grade III: 33). The specimens are H&E-stained colorectal tissue sections acquired at the Poursina Hakim Research Center of Isfahan University of Medical Sciences, Iran, diagnosed between 2014 and 2019, and graded according to the World Health Organization (WHO) criteria into three grades: well-differentiated (Grade I), moderately differentiated (Grade II), and poorly differentiated (Grade III). For each specimen, four magnification levels are provided: 4x, 10x, 20x, and 40x. The dataset is accessible via Mendeley Data (https://doi.org/10.17632/yfp5sfj47m.4) and at http://databiox.com, where the latest version is also available. The distinctive feature of this dataset is the provision of labeled specimens across all three differentiation grades at multiple magnification levels, enabling comprehensive computational analysis of colorectal cancer grading.

Figures

Figures reproduced from arXiv: 2607.12750 by Alireza Fahim, Amin Bahreini, Elham Amjadi, Hamidreza Bolhasani, Hojjatollah Rahimi, Sayed Mohammad Hasan Emami, Sayyed Mohammadreza Hakimian.

Figure 2
Figure 2. Figure 2: Representative H&E-stained colorectal tissue images from the CRC-HGD dataset at four magnification levels (4x, 10x, 20x, 40x) for Grade I (well-differentiated), Grade II (moderately differentiated), and Grade III (poorly differentiated) adenocarcinoma. What distinguishes this dataset from existing ones is the provision of histologically graded specimens across all three WHO differentiation levels, each ima… view at source ↗
Figure 1
Figure 1. Figure 1: Number of patients and total images per grade in the CRC [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 3
Figure 3. Figure 3: Distribution of Grade I (well-differentiated) images per magnification level: 4x=106, 10x=244, 20x=251, 40x=259 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Distribution of Grade II (moderately differentiated) images per magnification level: 4x=75, 10x=210, [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Distribution of Grade III (poorly differentiated) images per magnification level: 4x=33, 10x=97, 20x=99, [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Proportional distribution of (left) patients and (right) images across all categories in the CRC [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Stacked bar chart showing total image counts per magnification level stratified by grade (4x=216, [PITH_FULL_IMAGE:figures/full_fig_p007_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

17 extracted references

  1. [1]

    Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries

    Bray F, Laversanne M, Sung H, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229 –263

  2. [2]

    Global colorectal cancer burden in 2022 and projections to 2050: incidence and mortality estimates from GLOBOCAN

    Xi Y, Xu P. Global colorectal cancer burden in 2022 and projections to 2050: incidence and mortality estimates from GLOBOCAN. Cancer Epidemiol. 2023;87:102–155

  3. [3]

    Global, regional, and national burden of colorectal cancer and its attributable risk factors in 204 countries and territories, 1990–2021

    GBD 2021 Colorectal Cancer Collaborators. Global, regional, and national burden of colorectal cancer and its attributable risk factors in 204 countries and territories, 1990–2021. Lancet Gastroenterol Hepatol. 2024

  4. [4]

    Colorectal carcinoma: pathologic aspects

    Fleming M, Ravula S, Tatishchev SF, Wang HL. Colorectal carcinoma: pathologic aspects. J Gastrointest Oncol. 2012;3(3):153–173

  5. [5]

    Prognostic factors in colorectal cancer: College of American Pathologists Consensus Statement 1999

    Compton CC, Fielding LP, Burgart LJ, et al. Prognostic factors in colorectal cancer: College of American Pathologists Consensus Statement 1999. Arch Pathol Lab Med. 2000;124(7):979–994

  6. [6]

    Has the new TNM classification for colorectal cancer improved care? Nat Rev Clin Oncol

    Nagtegaal ID, Quirke P, Schmoll HJ. Has the new TNM classification for colorectal cancer improved care? Nat Rev Clin Oncol. 2012;9(2):119–123

  7. [7]

    Deep learning models for poorly differentiated colorectal adenocarcinoma classification in whole slide images

    Yoshida H, et al. Deep learning models for poorly differentiated colorectal adenocarcinoma classification in whole slide images. Sci Rep. 2021;11:23041

  8. [8]

    Association between molecular subtypes of colorectal tumors and patient survival

    Phipps AI, Alwers E, Harrison T, et al. Association between molecular subtypes of colorectal tumors and patient survival. Gastroenterology. 2020;158(8):2158–2168

  9. [9]

    WHO Classification of Tumours of the Digestive System

    Nagtegaal ID, et al. WHO Classification of Tumours of the Digestive System. 5th ed. Lyon: IARC Press; 2019

  10. [10]

    100,000 histological images of human colorectal cancer and healthy tissue

    Kather JN, Halama N, Marx A. 100,000 histological images of human colorectal cancer and healthy tissue. Zenodo. 2018. https://doi.org/10.5281/zenodo.1214456

  11. [11]

    Multi -class texture analysis in colorectal cancer histology

    Kather JN, Weis CA, Bianconi F, et al. Multi -class texture analysis in colorectal cancer histology. Sci Rep. 2016;6:27988

  12. [12]

    Color-CADx: a deep learning approach for colorectal cancer classification

    Sharkas M, Attallah O. Color-CADx: a deep learning approach for colorectal cancer classification. PeerJ Comput Sci. 2024;10:e1866

  13. [13]

    CRCCN-Net: automated framework for classification of colorectal tissue using histopathological images

    Alkassar S, et al. CRCCN-Net: automated framework for classification of colorectal tissue using histopathological images. Biomed Signal Process Control. 2022;78:104002

  14. [14]

    EBHI -Seg: a novel enteroscopy biopsy histological image dataset for image segmentation tasks

    Shi W, et al. EBHI -Seg: a novel enteroscopy biopsy histological image dataset for image segmentation tasks. Front Med. 2023;10:1114673

  15. [15]

    HCCANet: histopathological image grading of colorectal cancer using CNN

    Liang Y, et al. HCCANet: histopathological image grading of colorectal cancer using CNN. Sci Rep. 2022;12:15103

  16. [16]

    Deep learning models for poorly differentiated colorectal adenocarcinoma classification

    Yoshida H, et al. Deep learning models for poorly differentiated colorectal adenocarcinoma classification. Sci Rep. 2021;11:23041

  17. [17]

    A histopathological image dataset for grading breast invasive ductal carcinomas

    Bolhasani H, Amjadi E, Tabatabaeian M, Jassbi SJ. A histopathological image dataset for grading breast invasive ductal carcinomas. Informatics Med Unlocked. 2020;19:100341. Author Biographies Elham Amjadi, MD is a Pathologist with expertise in gastrointestinal and breast pathology. She received her medical degree and completed her residency in pathology a...