{"id":"33b5071b-2167-407d-853b-595a0441057c","arxiv_id":"2607.04020","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new publicly available dataset pairs uterine whole-slide images with case- and slide-level pathology reports in German and English.","lead":"TUM-Uteria adds 216 uterine pathology cases that link whole-slide images to case- and slide-level diagnostic pathology reports in German and English. It gives computational pathology researchers a paired image–text resource for training and evaluating multimodal models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Slide–report pairing reliability lacks quantitative validation; 71/216 cases required manual judgment with no inter-rater agreement reported.","rationale":"The paper's strongest claim is that the dataset provides reliable slide-level WSI–report pairs, validated by pathologists. The reader's weakest assumption identifies slide–report pairing as the key risk, and I agree. The manuscript describes a multi-stage pipeline with expert review, but the only quantitative information is the count of 71 cases needing manual judgment. No quantitative measure of agreement, error rate, or reproducibility of the manual assignments is provided. Without such evidence, the central utility of the dataset for multimodal learning is not established: if a nontrivial fraction of pairs are misaligned, models trained on them will learn spurious correlations between image features and unrelated text. The paper also does not report how many WSIs were excluded due to low pairing confidence, which could bias the dataset toward easier cases. These gaps are addressable—an inter-annotator study on a random sample would directly test pairing reliability. I also note a minor internal inconsistency: the scanner resolution is 0.25 µm/pixel in Section 2.1 but 0.26 µm/pixel in Figure 2D, and the Data Records section contains a placeholder link instead of a direct URL; these are secondary but should be fixed. Given the dataset is otherwise clearly described and potentially valuable, the appropriate verdict is CONDITIONAL: accept once quantitative pairing validation and a direct repository link are provided.","tokens_in":7161,"tokens_out":3852,"duration_ms":45123,"concrete_test":"Select a random sample of 50 cases (e.g., 30 from the 71 manually judged and 20 from the automatically paired subgroups; include both 1-WSI and multi-WSI cases). Have two board-certified pathologists, blinded to the released assignments, independently map each WSI to the appropriate text block(s) in the original German report, or mark 'cannot determine.' Compute Cohen's kappa between their assignments and against the released mapping. Also have them rate semantic fidelity of the LLM-extracted English slide-level description to the original report on a 1–5 scale. If kappa is below 0.8 or median fidelity below 4, the pairing reliability claim is unsupported and the dataset should be released with confidence scores or restricted to the automatically paired subset.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim of TUM-Uteria is that each WSI is paired with a faithful, slide-level textual description, enabling multimodal training. This claim rests almost entirely on a manual pairing process whose reliability is not quantitatively established. Section 2.2.1 and Section 5 state that 71 of 216 cases required manual judgment to establish WSI–text correspondences, and that cases with insufficient confidence were excluded—but no exclusion count, no inter-rater agreement, and no repeat-annotation results are reported. For the remaining 145 cases, two pathologists 'reviewed' the automatically extracted pairs, but again no agreement metric or error rate is given. Because LLM-extracted slide descriptions from case-level reports can be subtly wrong (e.g., assigning text about block B to slide A when identifiers are absent), and because the pathologists' manual resolution is subjective, the dataset's core utility—learnable image–text alignment—is not demonstrated. The manuscript itself acknowledges 'sufficient confidence' as the criterion, which is undefined.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TUM-Uteria, a publicly released uterine pathology dataset containing 216 clinical cases and 455 H&E-stained whole-slide images (WSIs) paired with diagnostic pathology reports at both case and slide levels. The dataset was collected from routine diagnostic workflows at a tertiary medical center, with German pathology reports translated into English. The manuscript describes a multi-stage pipeline that includes specimen collection and WSI scanning, three-step report anonymization using an offline LLM (Qwen-30B), slide-level report extraction with LLM assistance, manual adjudication for ambiguous cases, expert review by pathologists, and final quality control. It also reports dataset statistics such as age distribution, WSI counts per case, and diagnostic category frequencies. The authors claim the dataset supports multimodal computational pathology tasks such as automated report generation and AI-assisted diagnosis.","tokens_in":7378,"tokens_out":4077,"duration_ms":51493,"significance":"If the claimed pairing reliability holds, TUM-Uteria fills a genuine gap: it provides slide-level WSI–text pairs drawn from routine clinical practice, including a realistic mix of benign, precancerous, and malignant uterine pathology, and includes both German and English report versions. The construction pipeline is described in sufficient detail to be reproduced, the dataset is made available through a gated Hugging Face repository, and the validation workflow involves domain experts, which strengthens the resource. However, the central claim that each WSI is paired with a faithful slide-level text description is not backed by quantitative evidence. Since the dataset's utility for multimodal learning depends directly on the correctness of these pairs, the absence of inter-rater agreement, exclusion counts, LLM extraction accuracy, and audit metrics is a load-bearing weakness rather than a cosmetic issue.","major_comments":[{"comment":"The paper states that 71 of 216 cases required manual judgment to determine WSI–text correspondences, but reports no exclusion count, no inter-rater agreement statistic (e.g., Cohen’s kappa), and no repeat-annotation results for these assignments. Because these pairs are the core deliverable of the dataset, the reader cannot verify the reliability of the alignment. Please report the number of cases/WSIs that were excluded after manual review, the number of assignments independently adjudicated by both pathologists, and an agreement metric, or at minimum a quantitative audit of a random sample.","section":"§2.2.1 / §5 (third validation paragraph)"},{"comment":"The criterion for inclusion, 'sufficient confidence,' is never operationalized. In addition, the two pathology experts reviewed all automatically extracted pairs, but the manuscript reports no outcome of that review: no count of corrected assignments, no error rate for the automatic LLM-based slide-level extraction, and no characterization of the types or frequency of errors found. Without these numbers, the claim that the released pairs are 'high-quality' and 'reliable' is unsupported. Please define the confidence threshold or describe the adjudication protocol, and report how many pairs were corrected or rejected during expert review.","section":"§2.2.1 / §5 (expert review step)"},{"comment":"The text states there were 'no significant differences in age or diagnostic category distribution' between cases with 1–3 WSIs and those with 4 or more WSIs, but no statistical test, effect size, or p-value is reported. This comparison is used to justify the retention of cases with at most 3 WSIs, which directly shapes the released dataset. Please provide the test used, the test statistic, and the p-values (or confidence intervals) for the age and category comparisons.","section":"§3 / Fig. 2C"},{"comment":"The Qwen-30B model is used for de-identification, residual-PHI screening, and slide-level text extraction, but no accuracy metrics are given for any of these steps. For a medical dataset, unanswered questions include: how many reports were flagged by the second LLM screen and what fraction required revision; and what the error rate of the slide-level extraction was relative to a gold standard. Even a small labeled evaluation set with precision/recall for extraction and a privacy audit on a held-out sample would substantiate the pipeline's reliability and mitigate the risk that subtle PHI leakage or text–slide misassignment is present.","section":"§2.2 / §2.2.1 / §5 (LLM-based anonymization and extraction)"}],"minor_comments":[{"comment":"The resolution is reported as 0.25 µm/pixel in the text and 0.26 µm/pixel in Figure 2D. Please reconcile this inconsistency.","section":"§2.1 / Fig. 2D"},{"comment":"The Data Records section contains the placeholder '[DATA REPOSITORY LINK]' even though the abstract and Data Availability name the Hugging Face repository. Provide the actual URL or DOI for the dataset.","section":"§4"},{"comment":"The WSI filename structure example appears garbled ('TUM_Uterus s001 T1-A-2 .svsc001 HE'). The intended pattern should be typeset clearly, e.g., 'TUM_Uterus_s001_T1-A-2_HE.svs'.","section":"Fig. 2E"},{"comment":"The LLM is referred to only as 'Qwen-30B.' For reproducibility, please specify the exact model version, prompt template, decoding parameters, and the date/software environment used.","section":"§2.2 and §2.2.1"},{"comment":"The paper notes that 'no evidence of malignancy' reports were not automatically classified as normal, which is a useful and careful design choice. It might help to state explicitly how the six diagnostic categories were assigned in ambiguous cases, e.g., whether the final label came from the critical finding section alone or from the full report.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"This is a dataset paper, so the dataset's construction and validation are the main scientific content. The central gap is quantitative: no inter-rater agreement, exclusion counts, LLM extraction accuracy, or audit metrics are reported for the slide–report pairing, which the entire resource depends on. I would not require clinical gold-standard validation, but I would require the authors to document the reliability of the pairing process with numbers. If they can provide these in a revision, the paper is likely acceptable; without them, the 'high-quality pairs' claim is not verifiable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know: TUM-Uteria is a real, novel resource—216 uterine cases, 455 slide-level WSI–report pairs, German and English, routine clinical spectrum, built with a serious multi-stage pipeline. I don't know of another public dataset that pairs WSIs with full diagnostic reports at slide level for uterine pathology. That alone makes it worth a look for anyone working on multimodal computational pathology.\n\nWhat's good: The construction is described concretely—Leica scanner, 40x, H&E only, anonymization with an offline LLM plus manual review, and explicit handling of reports without slide identifiers. The decision to keep only cases with ≤3 WSIs to reduce alignment ambiguity is sensible, and they checked age and diagnostic category balance against the excluded group. The HuggingFace gated release with automatic access is pragmatic and reviewer-friendly.\n\nWhere it's soft: The stress-test note is on target. The central claim is that each WSI is paired with a faithful slide-level text description. For 71 of 216 cases, that pairing required manual judgment, and the rest were \"reviewed\" by two pathologists—but nowhere do we get numbers: no inter-rater agreement, no count of unassignable cases, no error rate for the LLM extraction, no second review. \"Sufficient confidence\" is undefined. The repository link appears in Data Availability, not Data Records—a minor inconvenience but odd for a dataset paper. None of this means the dataset is bad; it means the authors are asking readers to take the alignment on faith, and for a resource whose entire value is correctness of those pairs, that's a real gap.\n\nAlso worth noting: the scale is modest (455 pairs), so it won't serve as a foundation-model training corpus, but it's well-suited for benchmarking report generation and vision-language evaluation.\n\nBottom line: This is a competent, honest dataset paper with a fixable weakness. I'd send it to peer review, but I'd ask for quantitative validation of the pairing—at minimum inter-observer agreement on a sample and a count of excluded cases. With that added, it's a solid contribution.","headline":"A genuinely useful new multimodal pathology dataset, but the slide–report pairing validation is under-reported; I'd want revision before fully trusting the alignment.","tokens_in":7907,"tokens_out":2528,"would_cite":false,"duration_ms":28277,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Uteria dataset pairs 455 uterine whole-slide images with diagnostic reports, giving multimodal pathology AI a clinically realistic training resource.","keywords":["uterine pathology","whole-slide images","pathology reports","multimodal computational pathology","visual-language learning","report generation","dataset","H&E staining"],"falsifier":"Take a random sample of the 455 pairs, or all 71 manually resolved cases, and ask two independent pathologists to match each slide-level report to the correct WSI among the case's slides without knowing the released assignment. If their agreement with the released pairing is at or near chance, or if the two pathologists disagree substantially, the paper's central reliability claim is not supported.","tokens_in":7083,"feed_emoji":"🔬","tokens_out":3752,"duration_ms":44394,"temperature":0.7,"pith_summary":"This paper introduces Uteria, a public dataset of 216 uterine pathology cases built from routine diagnostic service, containing 455 whole-slide image–report pairs. The authors' central claim is that these slide-level image–text pairs are reliable enough to train and evaluate multimodal computational pathology systems, including automated report generation and AI-assisted diagnosis. Unlike most existing pathology datasets, which offer only patch- or slide-level classification labels, Uteria preserves the structure of real clinical reports and reflects the natural mix of benign, precancerous, and malignant findings seen in a tertiary medical center. A sympathetic reader would see this as addressing a concrete scarcity: paired slide-level image–text data with meaningful diagnostic language.","feed_headline":"New dataset pairs 455 uterine slide images with pathology reports","feed_subtitle":"Routine clinical cases give multimodal AI the image–text training data needed for automated pathology report generation.","key_machinery":"The central object is the slide-level WSI–report pair. Since each case yields one report but up to three slides, the pipeline converts case-level reports into slide-level descriptions using an offline large language model, separates descriptions that carry explicit slide/block identifiers, flags reports without identifiers for manual review, and subjects all pairs to a multi-stage expert validation. The six diagnostic categories—normal, benign tumor, malignant tumor, precancerous lesion, inflammatory or reactive, and insufficient or uncertain—are assigned from the written findings and provide the supervised structure that makes the pairs usable for multimodal training and benchmarking.","core_discovery":"The paper presents Uteria, a publicly released dataset of 216 clinical cases comprising 455 H&E-stained whole-slide images, each paired at slide level with a diagnostic text extracted from the original German pathology report and translated into English. The core claim is that the slide–report pairing is clinically dependable: cases were collected from routine workflows rather than enriched for cancer, slide-level descriptions were produced by an LLM-assisted extraction pipeline with explicit flagging of ambiguous cases, 71 of 216 cases were manually resolved by experts, and the final pairs were reviewed by two pathologists and then by three pathology-AI researchers. Each pair also carries o","pith_inferences":["A natural benchmark the paper does not build: hold out the 71 manually resolved cases as a hard pairing test, and check whether a vision–language model can distinguish true slide–report pairs from shuffled mismatches on those cases. If it cannot, the manual assignments may encode information not present in the images.","The routine-clinical class imbalance—normal cases dominating—implies that the practical bottleneck for report generation on this dataset is producing accurate, conservative 'no abnormality' language, not recognizing malignancy.","Future extensions could add immunohistochemistry slides and frozen sections, and report inter-pathologist agreement on the manual assignments; either would directly strengthen or bound the central reliability claim.","The paired structure could support contrastive pretraining that links visual morphology to diagnostic wording, a recipe that might transfer to other gynecologic or organ-specific sites if the same report format is used."],"forward_implications":["If the pairing is trustworthy, models can be trained end-to-end to generate diagnostic narratives from whole-slide images and evaluated against these paired reports.","Because the cohort mirrors routine clinical composition—56.7% normal or no significant abnormality and only 8.35% malignant—models trained here may generalize better to everyday pathology workloads than models trained on cancer-enriched datasets.","Case-level identifiers allow patient-level data partitioning, so slides from the same patient can be kept together during training and evaluation, supporting multi-slide reasoning without leakage.","The bilingual German–English reports enable cross-language report generation and a direct check of whether diagnostic meaning survives translation.","The six diagnostic categories supply slide-level supervised labels for classification baselines, giving a common benchmark for future multimodal methods."],"fun_headline_variants":["455 uterine slides paired with text for AI pathology","New dataset: 455 uterine slide-report pairs for multimodal AI","Uterine WSI + reports: 455 pairs fuel AI report generation","Slide-level text pairs 455 uterine images for AI diagnosis"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that each slide-level text genuinely corresponds to its paired whole-slide image; for 71 of the 216 cases that correspondence was assigned by manual judgment, and the paper reports no quantitative inter-rater agreement or independent check of those assignments.","fun_headline_variants_meta":{"raw":{"variants":["455 uterine slides paired with text for AI pathology","New dataset: 455 uterine slide-report pairs for multimodal AI","Uterine WSI + reports: 455 pairs fuel AI report generation","Slide-level text pairs 455 uterine images for AI diagnosis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000383,"raw_usage":{"total_tokens":1863,"prompt_tokens":741,"completion_tokens":1122,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":1052}},"tokens_in":485,"tokens_out":1122,"duration_ms":10314,"temperature":1.0,"reasoning_tokens":1052,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T08:40:45.495511+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the 455 pairs, or all 71 manually resolved cases, and ask two independent pathologists to match each slide-level report to the correct WSI among the case's slides without knowing the released assignment. If their agreement with the released pairing is at or near chance, or if the two pathologists disagree substantially, the paper's central reliability claim is not supported.","supporting_citations":[],"review_version":2}