{"id":"a62d4797-d5d6-4794-ac88-0db8fe1a1118","arxiv_id":"2411.10752","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A cleaned Camelyon+ dataset re-labels breast cancer lymph node slides into four classes (negative, micro, macro, ITC) and provides MIL benchmark results.","lead":"This paper presents a cleaned and re-annotated version of the Camelyon breast cancer lymph node datasets, with a new four-class metastasis-size classification. It also benchmarks 12 multiple instance learning methods and 6 feature extractors on the new Camelyon+ dataset, which may become a standard testbed for AI in pathology.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unstated size thresholds and absent validation leave Camelyon+ four-class labels unverifiable; an AJCC-consistency check would resolve this.","rationale":"The reader's conditional verdict identifies the right soft spot, and I agree with it. Reading the paper in good faith, the authors do describe their exclusion criteria, release pre-extracted feature files, and blind the re-annotation of Camelyon-16 by renaming slides, which are good practices. But the four-class conversion is the crux: the paper contains no stated cutoff values, no agreement metric, and no comparison against any held-out reference. The claim that Camelyon+ is clinically more relevant requires that the classes mean what they mean clinically. The proposed AJCC-consistency check would settle whether the labels are clinically consistent; if it passes, the benchmark is likely usable, and if it fails, the labels and rankings need revision. Therefore I do not change the reader's conditional verdict; the concern is exactly why the verdict should remain conditional until the check is performed.","tokens_in":16564,"tokens_out":6320,"duration_ms":66233,"concrete_test":"Download the released XLSX labels and XML annotations from the ScienceDB/GitHub repository; compute the maximum tumor extent per WSI from polygon annotations at the stated 20x magnification; assign classes using AJCC 8th-edition cutoffs (ITC <=0.2 mm or <200 cells, micro >0.2-2 mm, macro >2 mm); and compare with the Camelyon+ labels. A mismatch on any slide, or an inability to reproduce the stated class counts (871/174/251/54), would show the four-class labels were not derived from clinically consistent size thresholds and would require relabeling before the benchmark rankings can be trusted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Camelyon+'s central contribution is a four-class label set (negative/micro/macro/ITC). The paper says these labels were assigned based on the sizes of re-annotated tumor regions, but it never states the numeric thresholds used, and no inter-observer or external validation is reported. Since the authors also relabeled the Camelyon-17 test set whose official labels are not publicly available, there is no independent reference check against their corrections. If the thresholds deviate from accepted clinical definitions (AJCC 8th: ITC <=0.2 mm or <200 cells; micro >0.2-2 mm; macro >2 mm), or if the corrections are not reproducible, the class labels and every comparison in Tables 3-5 inherit systematic error. This is load-bearing because the four-class labels are the dataset's main contribution; the exclusion criteria, while documented, do not establish that the remaining labels are ground truth.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces Camelyon+, a cleaned and re-annotated combination of the Camelyon-16 and Camelyon-17 datasets for breast cancer lymph node metastasis detection in whole slide images. The authors removed 49 low-quality slides, corrected slide-level labels, added pixel-level annotations to positive slides, and upgraded the binary classification task to a four-class task (negative, micro-metastasis, macro-metastasis, ITC) based on the sizes of re-annotated tumor regions. They then benchmarked 12 multiple instance learning (MIL) methods with 6 feature extractors, reporting metrics on the corrected Camelyon-17 and on the merged Camelyon+ dataset. The central claim is that the cleaned dataset provides a more reliable and clinically relevant benchmark than the original Camelyon series.","tokens_in":16761,"tokens_out":4628,"duration_ms":54008,"significance":"If the re-annotation and label corrections are validated, Camelyon+ would be a valuable community resource that merges the two Camelyon datasets, adds a clinically meaningful four-class label set, and provides extracted features and code to facilitate reproducible benchmarking. The paper also raises a substantive question about whether MIL is the right paradigm for size-based classification tasks. However, the dataset's core contribution depends on label corrections and threshold definitions that are not yet documented with sufficient protocol detail, validation, or inter-observer agreement measures, and the provenance of the Camelyon-17 test set labels is unclear. These gaps currently prevent the benchmark from being fully reproducible and trustworthy.","major_comments":[{"comment":"The four-class definitions are never operationalized. The paper states only that classes were assigned \"based on the sizes of re-annotated tumor regions,\" but gives no numeric thresholds. Please specify the exact criteria used (e.g., AJCC 8th edition: ITC ≤0.2 mm or <200 cells; micro >0.2–2 mm; macro >2 mm), how tumor region size was measured (e.g., largest focus diameter, total area, or cell count), and how slides with multiple foci of different sizes were assigned to a class. Without this information, the class labels and all downstream comparisons in Tables 3–5 are not reproducible.","section":"Methods / Dataset Overview and Abstract"},{"comment":"No inter-observer agreement or external validation is reported for the label corrections or the pixel-level annotations. Please provide the number and experience of the pathologists involved, the blinding procedure, the consensus method, and agreement statistics (e.g., Cohen's kappa) for slide-level classes and for pixel-level annotation overlap. An independent validation step, such as a second pathologist panel or comparison with a reference standard, is essential to support the claim that the original labels were \"erroneous\" and that the corrections are correct.","section":"Methods / Technical Validation and Data Records"},{"comment":"The paper states that the Camelyon-17 test set labels are \"not publicly available,\" yet Table 1 reports performance metrics on the Camelyon-17-Origin test set. Please clarify how the official labels for this test set were obtained. If they were acquired through the challenge organizers, state this explicitly. If they were generated by the authors' re-annotation, then the Camelyon-17-Origin results are not against the official ground truth, and the \"Origin vs. Refine\" comparison in Tables 1–2 and Figures 3–4 must be reinterpreted accordingly.","section":"Methods / Dataset Overview and Technical Validation / Camelyon-17 Comparative Experiment"},{"comment":"Several entries in Table 2 show very large standard deviations (e.g., PLIP Max-MIL AUC 72.7 ± 11.09, UNI Max-MIL AUC 77.9 ± 11.95, and several F1 values with ±8–10), while other entries are much more stable. The paper does not discuss this instability, despite using these numbers to argue that dataset refinement improves the accuracy and fairness of model rankings. Please comment on the stability of the results and consider reporting additional random seeds or per-fold performance to ensure that the comparative claims are robust.","section":"Technical Validation / Camelyon-17 Comparative Experiment (Table 2)"}],"minor_comments":[{"comment":"The word \"dastaset\" is a typo for \"dataset\" in the captions of Tables 3, 4, and 5.","section":"Tables 3–5 captions"},{"comment":"The phrase \"Camelon+ Dataset\" should read \"Camelyon+ Dataset.\"","section":"Usage Notes"},{"comment":"The feature extractor name \"PILP\" in the first paragraph of the benchmark experiment section appears to be a typo for \"PLIP.\" Also, the paper inconsistently uses \"VIT-S\" and \"ViT-S\" for the same model; please standardize the notation.","section":"Methods / Benchmark Experiment"},{"comment":"The sentence \"The original WSI data can be downloaded from the official websites of Camelyon161 and Camelyon-172\" is missing hyphens and spaces: it should read \"Camelyon-16\" and \"Camelyon-17.\"","section":"Data Records"},{"comment":"Figure 5 contains many small text labels that appear garbled or overlapping in the manuscript version; please provide a higher-resolution figure and check the readability of the feature encoder names.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely destined for a data-descriptor or benchmark-oriented venue. The reader's report and my own assessment agree that the core dataset contribution is promising but not yet fully documented. The inclusion of AMD-MIL (reference 27, from the same first author) is a minor self-citation concern but does not, by itself, compromise the benchmark; the more pressing issues are the missing threshold definitions, lack of inter-observer validation, and unclear provenance of the Camelyon-17 test set labels. These are fixable within the scope of a revision, so I recommend major_revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful thing here is the resource: a cleaned, relabeled Camelyon+ dataset with four-class slide labels (negative, micro, macro, ITC), expert pixel annotations for the previously unreleased test set, and extracted features for six encoders. The benchmark itself is standard MIL evaluation, but the dataset artifact is real and the community will likely use it. The paper also does a decent job documenting why slides were removed (49 slides, with reasons like treatment response and blur), and the Camelyon-17-Origin vs Refine comparison gives a sensible demonstration that cleaning changes model rankings. Credit where due: shipping the cleaned labels, the pixel annotations, and the extracted features is more than most benchmark papers do. The self-citation of AMD-MIL is noticeable but not a problem; no result is forced by its inclusion.\n\nThe soft spot is the one the stress-test flags, and it lands. The four-class labels are the paper's central contribution, but the paper never states the numeric size thresholds used to convert re-annotated tumor regions into micro, macro, and ITC. It says \"based on the sizes of re-annotated tumor regions\" and stops. No AJCC cutoff check, no inter-observer agreement, no validation against any external reference. Since the Camelyon-17 test labels were never public, there is no independent check on their corrections either. If the thresholds deviate from clinical definitions, every label and every ranking in Tables 3-5 inherits the error. The exclusion criteria are documented, but they do not establish that the remaining labels are right. The code is also promised, not yet released, which is a minor issue for reproducibility but fixable.\n\nThe paper is honest in its usage notes, saying the dataset is not for diagnosis-focused algorithms. That helps. But the missing annotation protocol is load-bearing. I would not trust the four-class labels as ground truth until the thresholds are stated and some agreement stats are reported. The benchmark numbers themselves look plausible and are reported with standard deviations; they just inherit whatever error is in the labels.\n\nWho this is for: anyone working on MIL for pathology who wants a size-based four-class benchmark, and anyone who cares about label quality in Camelyon. It deserves a serious referee, but the referee should require the threshold details and a validation step before acceptance. As written, it is a conditional pass.","headline":"A genuinely useful cleaned Camelyon benchmark, but the four-class labels are unverifiable until the authors state their size thresholds and show some validation of the relabeling.","tokens_in":17227,"tokens_out":1564,"would_cite":true,"duration_ms":22885,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Cleaning and re-annotating the Camelyon slides yields a four-class benchmark that changes how models are ranked and exposes weak performance on isolated tumor cells.","keywords":["breast cancer","lymph node metastasis","whole slide images","benchmark dataset","multiple instance learning","foundation models","isolated tumor cells","histopathology"],"falsifier":"Independently re-review a random sample of the 1,350 Camelyon+ slides with multiple pathologists and check whether the four-class labels match; if a substantial fraction, for example more than 10 percent, of positive slides are assigned different categories or reclassified as negative, the benchmark's validity fails. A more direct test is to recover the size thresholds from the released pixel annotations and verify them against the published ITC, micro, and macro definitions.","tokens_in":16410,"feed_emoji":"🔬","tokens_out":8949,"duration_ms":75835,"temperature":0.7,"pith_summary":"The paper argues that the widely used Camelyon datasets for breast cancer lymph node metastasis contain enough low-quality slides, labeling errors, and annotation gaps to distort model evaluation, and that a cleaned, merged four-class version, Camelyon+, provides a more reliable benchmark. It reports constructing this dataset from 1,399 original slides, removing 49, correcting labels, and adding expert pixel annotations, yielding 1,350 whole-slide images labeled negative, micro-metastasis, macro-metastasis, or isolated tumor cells. On this benchmark the paper re-evaluates twelve multiple-instance-learning methods and six feature extractors, finding that pathology-specific features outperform natural-image features and that model rankings shift after cleaning. A sympathetic reader would care because the field's most common evaluation set is claimed to be measurably cleaner, and the benchmark exposes that current methods label the rarest category, isolated tumor cells, poorly.","feed_headline":"1,350 cleaned slides form a four-class breast cancer benchmark","feed_subtitle":"Re-annotating Camelyon slides corrects labels and exposes a weak spot on isolated tumor cells.","key_machinery":"The central mechanism is a dataset re-annotation pipeline: professional pathologists review each whole-slide image, exclude low-quality or treatment-altered slides, correct slide-level labels, and draw pixel-level tumor annotations where they were missing or wrong; from the size of these re-annotated regions the binary positive label is upgraded into the four classes negative, micro-metastasis, macro-metastasis, and isolated tumor cells. The evaluation machinery is embedding-based multiple instance learning: each slide is tiled into 256x256 patches at 20x magnification, features are extracted by a pre-trained encoder, and a MIL aggregator pools the patch features for slide classification.","core_discovery":"The authors establish that Camelyon+, assembled from corrected Camelyon-16 and Camelyon-17 slides, constitutes a valid four-class benchmark for lymph node metastasis classification, and that on this benchmark the standard multiple-instance-learning pipeline performs well on negative, micro, and macro classes but fails on isolated tumor cells. They report removing slides with blur, poor staining, treatment artifacts, or ambiguous positivity; re-labeling slides based on the size of re-annotated tumor regions; and providing pixel annotations missed in the original data. Their experiments show that after this refinement, the ranking of MIL models changes, with CLAM-MB remaining the strongest, and that pathology-pretrained feature extractors, especially the contrastively trained CONCH model, match or exceed much larger feature extractors. The paper thereby positions Camelyon+ as a resource that should replace raw Camelyon data for fair evaluation of histopathology models.","pith_inferences":["Because the paper does not publish the exact size thresholds that separate ITC, micro, and macro metastases, a natural next step is for the authors or others to state these cutoffs and validate them against the international tumor-node-metastasis criteria.","The released pixel annotations could be reused for a separate task: measuring tumor burden or metastasis size directly, which may be a more natural framing than four-class classification.","The observed ranking shift after cleaning implies that any future method claiming state-of-the-art on Camelyon-17 should be re-checked on Camelyon+ before the claim is trusted.","The benchmark's long-tailed class distribution mimics real clinical data, so it could also serve as a testbed for long-tail learning methods beyond multiple instance learning."],"forward_implications":["Previous model comparisons run on the raw Camelyon-17 splits may rank methods differently once low-quality slides and label errors are removed.","Pathology-specific feature extractors consistently beat natural-image encoders for this four-class task, so future work should use them as the default.","The contrastive image-text model CONCH performs comparably to much larger visual encoders, indicating that training data quality can compensate for model scale.","The poor F1 score on isolated tumor cells, driven by a 16:1 imbalance and the size-based definition of the class, makes the ITC category a targeted challenge for future MIL research."],"supporting_citations":[{"why":"Supplies the original Camelyon-16 whole-slide images and binary labels that the paper reprocesses and corrects.","marker":"[1]"},{"why":"Supplies the original Camelyon-17 slides and four-class labels that the paper merges and refines.","marker":"[2]"},{"why":"Provides the attention-based multiple instance learning architecture used as a baseline.","marker":"[17]"},{"why":"Provides the transformer-based MIL aggregator evaluated as a baseline.","marker":"[18]"},{"why":"Provides the clustering-constrained attention MIL model that remains top-ranked after cleaning.","marker":"[19]"},{"why":"Supplies the large-scale pathology visual encoder whose features the benchmark evaluates.","marker":"[6]"},{"why":"Supplies the whole-slide pathology foundation model features evaluated in the benchmark.","marker":"[7]"},{"why":"Supplies the contrastively trained pathology language-image model whose small-scale features match larger models.","marker":"[9]"},{"why":"Supplies the image-text pathology model evaluated as a contrastive-learning baseline.","marker":"[5]"}],"fun_headline_variants":["Camelyon+ fix: four-class benchmark exposes ITC blind spot","Corrected Camelyon reorders MIL models, CONCH matches big ones","Cleaned slides upgrade binary diagnosis to four-class task","Re-annotated Camelyon reveals ITC weakness in MIL pipeline","1,350 vetted slides form four-class breast cancer benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the pathologists' corrected slide labels and the size cutoffs used to assign micro, macro, and ITC classes are accurate and consistent; the paper does not state the cutoffs or report inter-observer agreement.","fun_headline_variants_meta":{"raw":{"variants":["Camelyon+ fix: four-class benchmark exposes ITC blind spot","Corrected Camelyon reorders MIL models, CONCH matches big ones","Cleaned slides upgrade binary diagnosis to four-class task","Re-annotated Camelyon reveals ITC weakness in MIL pipeline","1,350 vetted slides form four-class breast cancer benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000277,"raw_usage":{"total_tokens":1645,"prompt_tokens":937,"completion_tokens":708,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":617}},"tokens_in":553,"tokens_out":708,"duration_ms":6502,"temperature":1.0,"reasoning_tokens":617,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:19:57.866303+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently re-review a random sample of the 1,350 Camelyon+ slides with multiple pathologists and check whether the four-class labels match; if a substantial fraction, for example more than 10 percent, of positive slides are assigned different categories or reclassified as negative, the benchmark's validity fails. A more direct test is to recover the size thresholds from the released pixel annotations and verify them against the published ITC, micro, and macro definitions.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the transformer-based MIL aggregator evaluated as a baseline."},{"cited_title":"Y .et al","cited_arxiv_id":null,"evidence_quote":"Supplies the contrastively trained pathology language-image model whose small-scale features match larger models."},{"cited_title":"& Welling, M","cited_arxiv_id":null,"evidence_quote":"Provides the attention-based multiple instance learning architecture used as a baseline."}],"review_version":1}