{"id":"89d321f2-9ef8-494a-9ee9-d0dd510d5d1d","arxiv_id":"2507.05656","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"ADPv2 is a new public 32-label hierarchical histology dataset for healthy colon tissue, plus a VMamba classifier and an exploratory confidence-shift analysis on polyp subtypes.","lead":"The authors release ADPv2, a public dataset of 20,004 colon biopsy image patches annotated with a hierarchical taxonomy of 32 tissue types. They also train a VMamba classifier that reaches 0.88 mean average precision on 14 tissue labels and report confidence-score shifts on diseased tissue that align with two known colorectal cancer pathways.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The §5.2 t-tests treat each patch as independent, ignoring clustering within slides and RoIs, so the pathway-confirming p-values in Table 4 may reflect pseudoreplication rather than biology.","rationale":"The paper's primary contribution is the ADPv2 dataset, and the dataset release plus the two-stage VMamba pipeline appear useful and reproducible in principle. The reader's CONDITIONAL verdict is reasonable. The strongest over-claim is in the abstract and Section 5.2, where patch-level confidence shifts are said to confirm the two pathological pathways of colon cancer. The reader identified the domain-shift-versus-biology assumption as the weak point; my stress-test finds a more concrete, internal statistical flaw that undermines the quantitative evidence before that assumption is even considered. Welch t-tests on tens of thousands of patches from only a few hundred slides treat every patch as an independent observation, which inflates significance whenever patches within a slide are correlated—something that is almost certain in histopathology tiles from the same biopsy and RoI. A slide-level reanalysis is a decisive, inexpensive check and would settle whether the reported p-values are meaningful. Even if that check survives, the domain-shift concern remains, so the reader's verdict should not move upward; but the dataset itself can stand after the biomarker claim is reframed as exploratory rather than confirmatory. Therefore I keep the verdict at CONDITIONAL, matching the reader's assessment.","tokens_in":22977,"tokens_out":12787,"duration_ms":154287,"concrete_test":"Recompute Table 4 at the slide level: for each slide in Table B.8, average the logit-transformed GD confidence over all extracted RoI patches, then rerun the three Welch t-tests on these slide-level means. Alternatively, fit a linear mixed model with random intercepts for slide and RoI on the per-patch values. Report ICCs and cluster-robust p-values. If TT-vs-Normal or TT-vs-SH no longer reach p<0.05, the claimed two-pathway confirmation is not supported by the data as analyzed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central biomarker claim—that confidence-score shifts on gland areas confirm the two colorectal cancer pathways—rests on the three Welch t-tests in Table 4, computed on logit-transformed GD confidence scores for individual patches. But those patches are not independent. Table B.8 reports up to 26,258 patches from 217 TVA slides (~121 patches/slide) and 6,510 patches from 201 SSL slides (~32 patches/slide), with many patches likely coming from the same RoI annotation. If slide- or RoI-level effects exist (patient, staining batch, scanning session, tissue processing), the effective sample size is the number of slides or RoIs, not the number of patches. With an intraclass correlation of roughly 0.3–0.5, the design effect for TVA is about 22–38, reducing the effective sample from thousands to hundreds. The marginal TT-vs-Normal comparison (t = −1.99, p = 0.047) would very likely lose significance, and the TT-vs-SH contrast (t = 3.88) could also become non-significant. The paper reports no ICCs and fits no mixed-effects or cluster-robust models. This is a sharper, more immediate failure than the general domain-shift concern: even if the model's confidence reflects morphology, the statistical test used to 'confirm' the two pathways is invalid as reported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ADPv2, a publicly released dataset of 20,004 image patches extracted from 461 healthy colon biopsy slides, annotated with a hierarchical taxonomy of 32 histological tissue types (HTTs) across three levels. The authors train a VMamba encoder with Barlow Twins self-supervised pretraining and asymmetric-loss fine-tuning on a pruned 14-label subset, reporting a multilabel mAP of 0.879. They then apply this healthy-tissue-only model to RoI patches from four colorectal polyp subtypes (HP, SSL, TA, TVA) and analyze shifts in the model's confidence for the glands (GD) label, reporting three Welch t-tests that are interpreted as confirming the two canonical pathways of colorectal cancer development, namely the serrated pathway and the classical adenoma-carcinoma sequence. The paper's claimed contributions are the dataset, the two-stage model pipeline, and the confidence-distribution-based biomarker analysis.","tokens_in":23278,"tokens_out":22183,"duration_ms":226383,"significance":"If the results hold, the dataset contribution is genuinely useful: a public, hierarchically annotated healthy-colon patch dataset with a 32-label taxonomy is scarce, the data are released on Zenodo, the hyperparameters are fully tabulated in Tables B.5-B.7, and the train/test split is performed at the slide level, which prevents the most common form of data leakage. The classification result (mAP 0.879) provides a usable community baseline, and the confidence-shift probe is an interesting design. However, the most striking claim, that the confidence shifts confirm the two pathological pathways of colorectal cancer, is currently not supported by the statistical treatment as reported because the t-tests ignore clustering and the analysis lacks domain-shift controls; the specific problems and the path to fixing them are detailed in the major comments.","major_comments":[{"comment":"The three Welch t-tests in Table 4 treat individual patches as independent observations, but the data are nested within slides and RoIs. Table B.8 reports 26,258 patches from 217 TVA slides (about 121 patches per slide) and 6,510 patches from 201 SSL slides, with multiple patches per RoI annotation. Slide- and RoI-level factors (patient, staining batch, scanning session, tissue processing) very plausibly induce intra-cluster correlation, so the effective sample sizes are much smaller than the reported patch counts. Even with a modest intraclass correlation of 0.3, the design effect for TVA is about 37, reducing the effective TVA sample from 26,258 to roughly 700 patches, and the reported t-statistics are inflated in precision by approximately the square root of the design effect; for the TVA-dominated TT-vs-Normal contrast this factor is on the order of 6-9, so the marginal t = -1.99, padj = 0.047 would very likely cease to be significant, and the TT-vs-SH contrast (t = 3.88) is also at risk. The paper reports no ICCs and no mixed-effects or cluster-robust analysis, so, as reported, the Table 4 p-values do not provide valid quantitative support for the claim that the patterns confirm the two cancer pathways. This analysis should be re-run with slide- and RoI-level random effects (or cluster-robust standard errors), and the ICCs and effect sizes should be reported.","section":"Section 5.2, Table 4, Table B.8"},{"comment":"The interpretation of the confidence-score shifts as biologically meaningful morphology changes is not supported without a domain-shift baseline. Because the model was trained exclusively on healthy tissue, lower confidence on any diseased input is expected a priori, and differences between disease groups could reflect staining, scanning, tissue composition, or annotation-density differences rather than pathway biology. The disease slides in Table B.8 are all listed at 20x/0.4 mpp, but their staining and source institution are not stated, and the normal-side distribution appears to be computed on the study's own healthy slides, some of which the model encountered during pretraining on 115,413 unlabelled patches and during fine-tuning; this confounds disease status with slide familiarity. The authors should add controls, for example confidence on non-gland HTTs in the same RoI patches, normal mucosa adjacent to the lesions, a held-out healthy cohort from a different institution, or stain-normalized versions of the same slides. In the absence of such controls, the abstract's claim that the analysis 'confirms the two pathological pathways' overstates what the evidence supports; at most the data are consistent with such an interpretation.","section":"Sections 4.2 and 5.2"},{"comment":"The 'strong classification performance' claim should be qualified in light of the per-class metrics in Table 3. LC and RBC achieve TPR above 0.99 with FPR of 0.913 and 0.961 and TNR of 0.087 and 0.039, meaning these two classes are predicted positively in nearly all patches, so their high F1 scores (0.928 and 0.954) largely reflect an 82% class prevalence rather than discriminative ability. Several diagnostically relevant inflammatory classes (ES, MA, PC, LY) have F1 between 0.62 and 0.79. The paper should report per-class average precision (the components of the mAP), state the mAP after excluding or down-weighting the near-degenerate classes, and temper the statement that the model performs strongly on 'HTTs pathologists consider diagnostically relevant.'","section":"Section 5.1, Table 3"}],"minor_comments":[{"comment":"The manuscript says the label set is reduced 'from 32 labels to 14 labels' after merging surface epithelium into glands, yet Table 3 reports per-class metrics for 15 HTTs, including SE and GD as separate rows, and Figure 9 shows separate SE and GD heatmaps; the authors should clarify the actual training label set and how the SE row was computed if SE labels were merged into GD.","section":"Section 4.1.3, Table 3"},{"comment":"The sentence 'EF has a high FPR, low TNR, and low accuracy' contradicts the EF row of Table 3 (FPR 0.034, TNR 0.966, accuracy 0.950); the surrounding text also calls neutrophils the sibling of EF under POC, which is inconsistent with Table 2 where EF is placed in the Glands branch, and the cross-reference to 'section 4.1.2' for neutrophils appears to point to the wrong section.","section":"Section 5.1.1"},{"comment":"The t-tests are reported only as t-statistics and adjusted p-values; with thousands of patches per group, even trivially small distribution shifts will reach p < 0.001, so effect sizes (e.g., Cohen's d or a distribution-overlap measure) are needed to support the qualitative description of 'pronounced' versus 'modest' shifts.","section":"Table 4"},{"comment":"The epsilon in the logit transform is printed as '1-6', which should read 10^-6; please fix the typesetting and state the histogram binning and normalization used for Figure 11.","section":"Section 4.2"},{"comment":"The annotation process relies on a single annotator trained under pathologist supervision, with subsequent pathologist evaluation, but no inter-annotator agreement statistic is reported; for a dataset paper, a kappa or F1 agreement measure on a re-annotated subset would substantially strengthen confidence in the 32-label taxonomy.","section":"Section 3.3"},{"comment":"The claim of demonstrating the merit of the VMamba architecture and the two-stage SSL pipeline would be strengthened by at least one baseline comparison (e.g., ImageNet-initialized VMamba, a ResNet or ViT with the same pretraining, or BCE instead of ASL); as it stands the mAP is reported without any comparator.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the dataset contribution and the public Zenodo release with detailed hyperparameter documentation are genuine strengths that should be credited. The main risk of this submission is the gap between the headline claim of confirming the two pathological pathways and the evidence as analyzed: patch-level Welch t-tests on data that are evidently clustered by slide and RoI, with no domain-shift baseline. The pseudoreplication concern lands substantively because Table B.8 reports roughly 121 patches per TVA slide, and the marginal TT-vs-Normal contrast (p = 0.047) is exactly the comparison most likely to collapse under clustering-aware inference. I would ask the authors to re-run the biomarker analysis with mixed-effects models or cluster-robust standard errors, report ICCs and effect sizes, and add healthy-side and stain controls before the paper is accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"ADPv2 is worth your time, mostly for the data. The new contribution is real: 20,004 patches from 461 healthy colon biopsies, a 32-label hierarchical GI taxonomy, pathologist-reviewed annotation, and a public download. Patch size is normalized to 544x544 microns across scanners, and the annotation protocol is described in enough detail to reproduce. That fills a genuine gap between multi-organ atlases like ADP and narrow patch-level sets.\n\nThe modeling is competent but standard: Barlow Twins pretraining on 115k unlabeled patches, VMamba fine-tuning on 14 pruned labels, mAP 0.879. The pruning is disclosed, but the abstract's 'mAP 0.88' should be read with that caveat; Table 3 shows real weak spots (LC and RBC FPRs above 0.91, EF and MA poor). It's fine, but the paper's value isn't the architecture.\n\nThe soft spot is the biomarker analysis. The claim that confidence shifts 'confirm the two pathological pathways' is too strong. There is no domain-shift baseline: the model only saw healthy tissue, so a lower-confidence bulge on any diseased slide is expected, even if the gland morphology is irrelevant. The TT-vs-SH contrast is more interesting, since it's a comparison between two disease groups, but it still lacks a control for staining, scanner, or patch composition.\n\nThe statistical objection lands harder. The Welch t-tests in Table 4 treat each patch as independent, but patches are nested in slides and RoIs. Table B.8 gives 26,258 patches from 217 TVA slides and 6,510 from 201 SSL slides. With that clustering, the effective sample size is slides, not patches. No ICC, no mixed-effects model, no cluster-robust error. The marginal TT-vs-Normal p=0.047 likely would not survive. This is the right thing to fix in revision, not just a footnote.\n\nThe citation pattern is fine, including the ADP atlas citation; the taxonomy deliberately extends that earlier work. The generative-AI disclosure is honest and irrelevant.\n\nWho is this for? Anyone building organ-specific histology datasets or studying colon tissue representations; the dataset itself deserves use. The biomarker section is not ready as evidence, but it is a reasonable exploratory idea. I'd send this to peer review. A serious referee should push for: release model weights, add a stain/domain control, and re-analyze the t-tests at slide or RoI level. With that, the dataset claim stands.","headline":"ADPv2 is a solid, publicly released colon histology dataset, but the 'confirms two cancer pathways' claim reads stronger than the statistics support.","tokens_in":23838,"tokens_out":3631,"would_cite":true,"duration_ms":43529,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A model trained only on healthy colon tissue separates serrated lesions from conventional adenomas through confidence-score shifts on glands, matching the two recognized colorectal cancer pathways.","keywords":["Deep Learning","Multilabel Representation","Computational Pathology","ADPv2 Dataset","Biomarker Discovery","Self-Supervised Learning","Colorectal Cancer","Colon Histopathology"],"falsifier":"A decisive control would run the same RoI confidence analysis on healthy tissue only, comparing two groups that differ in stain or scanner rather than disease status (for example, HPS-stained versus H&E-stained normal colon biopsies): if a comparable leftward shift and TT-versus-SH separation appears, the biomarker signature tracks imaging domain rather than biology. A cheaper check is a permutation test that shuffles disease labels across RoI patches and recomputes the Welch t-statistics; if randomized groupings produce separations as large as t = 3.88, the reported pattern is not specific to the two carcinogenesis pathways.","tokens_in":1788,"feed_emoji":"🔬","tokens_out":2160,"duration_ms":217602,"temperature":0.7,"pith_summary":"This paper introduces ADPv2, a publicly available dataset of 20,004 image patches from healthy colon biopsies annotated with a hierarchical taxonomy of 32 histological tissue types across three levels, and trains a multilabel classifier on it using a two-stage procedure that reaches a mean average precision of 0.88. The central claim is that a model trained exclusively on healthy tissue can act as a probe for disease: applied to gland regions of polyps, its confidence scores shift in ways that separate serrated lesions (sessile serrated lesions and hyperplastic polyps) from conventional adenomas (tubular and tubulovillous adenomas). The authors interpret the shift patterns as statistical confirmation of the two established pathological pathways of colorectal cancer development and propose them as candidate image-based biomarkers. If the claim holds, biomarker discovery would no longer require large collections of annotated diseased tissue; a well-annotated baseline of normality and a confident model would suffice to expose disease signatures.","feed_headline":"Confirms two colon cancer pathways with a healthy-only model","feed_subtitle":"A model that never saw a polyp reads the serrated-versus-adenoma split in gland confidence drops.","key_machinery":"The load-bearing mechanism is the confidence score distribution analysis: a model that has only seen healthy colon tissue is applied to pathologist-annotated regions of interest on diseased and normal slides, and the logit-transformed confidence scores for each histological tissue type are compared across disease groups with Welch's two-sample t-tests and Holm-Bonferroni correction. The premise is that morphological deviation from normal tissue lowers and reshapes the model's confidence, so the shift pattern itself is the candidate biomarker. The representations come from a two-stage pipeline: Barlow Twins self-supervised pretraining, which decorrelates features by driving the cross-correlation matrix of two augmented views toward the identity matrix, followed by fine-tuning with the asymmetric loss on 14 pruned labels using the VMamba visual state-space model, a token-scanning architecture whose gating mechanism keeps time and memory linear in the number of image tokens.","core_discovery":"The paper's central claim is that confidence-score distributions from a model trained only on healthy colon tissue carry disease-specific biological information. On gland areas, diseased patches shift the model's logit-transformed confidence relative to normal tissue, with lower peak sharpness, leftward displacement, and broader spread. Quantified with Welch's two-sample t-tests, sessile serrated lesions and hyperplastic polyps together produce a pronounced reduction in gland confidence versus normal tissue (t = -6.98, adjusted p < 0.001), tubular and tubulovillous adenomas produce a more modest reduction (t = -1.99, adjusted p = 0.047), and the two disease groups differ significantly from each other (t = 3.88, adjusted p < 0.001). GradCAM visualizations show the model attending to sawtooth luminal borders and branching crypt outlines in serrated lesions but to atypical, pseudostratified nuclei in adenomas. The authors read these signatures as evidence that the healthy-only model detects biologically meaningful deviations from normal glandular architecture that mirror the two known colorectal carcinogenesis pathways, and they position the shift patterns as potential computational biomarkers for precursor lesion characterization.","pith_inferences":["If the confidence-shift signature were validated against molecular data such as BRAF or KRAS mutation status and CpG island methylator phenotype, the same healthy-only probe could extend to molecular subtyping of colorectal lesions; the paper does not run that test.","The healthy-baseline-as-probe strategy should transfer to other organs and diseases: any tissue with a granular normal-state taxonomy could have its confidence-shift patterns mined for disease signatures without collecting large diseased cohorts.","The per-HTT histograms could be compressed into scalar shift features (location, sharpness, spread per tissue type), turning the qualitative description into a quantitative multi-tissue fingerprint that a downstream classifier could use for polyp subtyping; the paper stops at univariate t-tests.","The diseased and normal RoI patches in the distribution analysis are matched in resolution and patch size, so a multi-site replication study across stains and scanners is the natural next test of whether the separation survives acquisition noise."],"forward_implications":["Confidence-score shift patterns on specific HTTs such as glands can serve as candidate image-based biomarkers for separating serrated lesions from conventional adenomas, a distinction where even expert pathologists show notable inter-observer variability.","A model trained entirely on healthy tissue can be used as a screening probe, reducing the need for large, expensively annotated diseased-tissue datasets in early biomarker studies.","The ADPv2 dataset supplies the community with a colon-specific resource of 20,004 patches, each carrying a hierarchical multilabel annotation with an average of 10 labels, for training and benchmarking tissue-type classifiers.","The two-stage Barlow Twins plus VMamba pipeline with the asymmetric loss reaches a mean average precision of 0.88 across 14 HTTs, indicating the architecture scales to gigapixel pathology without down-sampling.","The co-occurrence structure of the taxonomy reproduces known features of colon wall architecture, such as the tight coupling between glands and inflammatory infiltrates, suggesting the label hierarchy captures biologically meaningful tissue relationships."],"supporting_citations":[{"why":"The parent Atlas of Digital Pathology dataset and annotation platform; ADPv2 extends its hierarchical tissue-type taxonomy from multi-organ to colon-specific.","marker":"[18]"},{"why":"Supplies the Barlow Twins redundancy-reduction objective used for self-supervised pretraining on 115,413 unlabeled patches.","marker":"[42]"},{"why":"Supplies the VMamba visual state-space architecture whose linear-time token scanning carries the multilabel representation.","marker":"[13]"},{"why":"Supplies the asymmetric loss used in fine-tuning to counter extreme label imbalance in the multilabel HTT task.","marker":"[51]"},{"why":"Provides the premise that a model trained on healthy tissue shows informative confidence shifts on unseen tissue; the biomarker analysis rests on this transferability claim.","marker":"[52]"},{"why":"Defines the two colorectal carcinogenesis pathways (adenoma-carcinoma sequence versus serrated pathway) that the confidence analysis claims to confirm.","marker":"[3]"},{"why":"Supplies the histological description of serrated lesions used to interpret the gland-confidence shift of the SSL plus HP group.","marker":"[54]"},{"why":"Supplies the RandStainNA stain augmentation used during pretraining to make learned features robust across institutions and staining protocols.","marker":"[50]"}],"fun_headline_variants":["Healthy-only model flags two distinct colon cancer paths","Gland confidence shifts tease apart colon cancer pathways","Trained on normal tissue, still sees cancer path signatures","No polyps in training, yet model confirms two cancer routes"],"cache_read_input_tokens":25984,"weakest_assumption_plain":"The analysis assumes the confidence-score shifts on diseased gland patches are caused by biologically meaningful morphological deviations from normal tissue, not by trivial out-of-distribution artifacts such as different staining, scanner resolution, or patch composition; this interpretation is needed because the model never saw diseased tissue during training, so some drop in confidence on any diseased input is expected by default.","fun_headline_variants_meta":{"raw":{"variants":["Healthy-only model flags two distinct colon cancer paths","Gland confidence shifts tease apart colon cancer pathways","Trained on normal tissue, still sees cancer path signatures","No polyps in training, yet model confirms two cancer routes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000745,"raw_usage":{"total_tokens":3369,"prompt_tokens":1042,"completion_tokens":2327,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":2263}},"tokens_in":658,"tokens_out":2327,"duration_ms":22136,"temperature":1.0,"reasoning_tokens":2263,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:20:48.726975+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive control would run the same RoI confidence analysis on healthy tissue only, comparing two groups that differ in stain or scanner rather than disease status (for example, HPS-stained versus H&E-stained normal colon biopsies): if a comparable leftward shift and TT-versus-SH separation appears, the biomarker signature tracks imaging domain rather than biology. A cheaper check is a permutation test that shuffles disease labels across RoI patches and recomputes the Welch t-statistics; if randomized groupings produce separations as large as t = 3.88, the reported pattern is not specific to the two carcinogenesis pathways.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The parent Atlas of Digital Pathology dataset and annotation platform; ADPv2 extends its hierarchical tissue-type taxonomy from multi-organ to colon-specific."},{"cited_title":"Zbontar, L","cited_arxiv_id":null,"evidence_quote":"Supplies the Barlow Twins redundancy-reduction objective used for self-supervised pretraining on 115,413 unlabeled patches."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the premise that a model trained on healthy tissue shows informative confidence shifts on unseen tissue; the biomarker analysis rests on this transferability claim."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the two colorectal carcinogenesis pathways (adenoma-carcinoma sequence versus serrated pathway) that the confidence analysis claims to confirm."},{"cited_title":"Mezzapesa, G","cited_arxiv_id":null,"evidence_quote":"Supplies the histological description of serrated lesions used to interpret the gland-confidence shift of the SSL plus HP group."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the RandStainNA stain augmentation used during pretraining to make learned features robust across institutions and staining protocols."}],"review_version":1}