{"id":"cd0b165b-0953-431c-8134-df4eae5f6006","arxiv_id":"2608.10657","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A two-stage pipeline using pretrained vision encoders with retrieval augmentation achieves 0.9423 held-out accuracy for leukemia versus normal detection, but ALL versus AML subtyping collapses under domain shift, showing within-domain performance comes largely from dataset artifacts.","lead":"A leukemia-detection benchmark finds that a domain-specific vision model with retrieval augmentation reaches 94% accuracy on an unseen dataset for binary leukemia versus normal classification. The same protocol reveals that ALL versus AML subtyping collapses to the majority class, showing models exploited background artifacts rather than cell morphology.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Stage 2 collapse is convincingly demonstrated, but the causal attribution to a synthetic-background shortcut is not isolated from other source-level confounds, so the paper's 'robustness' framing is overstated.","rationale":"The reader's conditional verdict is appropriate, and this stress-test does not move it. The paper's core experimental contrast is real and valuable: a random within-domain split produces near-perfect ALL recall, while a strict held-out protocol collapses ALL recall to near zero. That contrast convincingly demonstrates that within-domain performance is driven by source-specific artifacts rather than transferable morphology, which is an important negative diagnostic result. My concern targets only the more specific mechanistic conclusion, repeated in the Conclusion and Future Work, that the artifact is primarily the synthetic versus natural smear background. Because all training ALL and all training AML come from different datasets with multiple simultaneously varying attributes, and because Section 4.2's background replacement was applied only to D1, background is collinear with dataset, stain, segmentation pipeline, and label noise. The retrieval-neighbor evidence is suggestive but is not imbalance-corrected, and the alpha sweep does not separate retrieval-bank composition from embedding geometry. A targeted background-restoration experiment on the held-out ALL-IDB2 images would settle the mechanism. The paper deserves credit for its transparency, for the control experiment, and for stating the background explanation as only 'a possible explanation' in Section 5.4; the problem is that the abstract and conclusion generalize beyond that hedge. The conditional verdict should stand, with the requested revision focusing on rewording the mechanistic claim and adding the proposed ablation or clearly labeling the background mechanism as an untested hypothesis.","tokens_in":18856,"tokens_out":5925,"duration_ms":63954,"concrete_test":"Apply the same synthetic-background pipeline used for C-NMC in Section 4.2 to the 130 ALL-IDB2 test images (replace the natural background with a sampled healthy peripheral-blood-smear region with random rotation and scaling, then Reinhard color normalization), and re-run the best Stage 2 configurations, e.g., DinoBloom linear probe and DinoBloom linear probe + RAC, with all other settings unchanged. If ALL recall rises materially from its current 0.000-0.023 while AML recall stays near 1, the background-shortcut mechanism is confirmed. If ALL recall remains near zero, an unisolated source artifact other than background is the dominant cause, and the paper's diagnostic claim and future-work direction would need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing diagnostic claim is that the ALL/AML collapse occurs because the encoder learns smear-background shortcuts rather than morphology. The evidence—near-perfect within-domain ALL recall versus near-zero held-out recall, 96.5% AML retrieval neighbors for ALL queries, and failure to recover ALL even at alpha=8—establishes that performance does not transfer across sources and that source identity is a strong predictor. It does not isolate background as the mechanism. In Stage 2, all training ALL images come from C-NMC (D1), with synthetic background replacement and Reinhard normalization, while all training AML images come from D4, with natural backgrounds, patient-level labels, four genetic AML subtypes, and some normal cells labeled AML. The held-out test is ALL-IDB2 (D5, natural background) versus AML-Cytomorphology (D2). Thus acquisition protocol, stain, magnification, cell-selection pipeline, label noise, and background all vary together with the ALL/AML label, and the paper performs no ablation separating them. The retrieval-neighbor statistic is also not corrected for the Stage 2 bank imbalance of 87.8% AML versus 12.2% ALL; random retrieval alone would already yield roughly 87.8% AML neighbors, so the observed 96.5% is an excess but does not identify the shortcut. Failure to recover ALL at alpha=8 is consistent with a top-20 retrieval set that is dominated by AML under many embedding geometries. Section 5.4 itself calls the background explanation 'a possible explanation,' yet the Conclusion states the collapse 'is attributable to dataset-specific artifacts' and the Future Work recommends background normalization as the remedy. The dataset-specific-artifact finding is well supported; the background-specific mechanism is not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript benchmarks three vision encoders (DinoBloom, BiomedCLIP, CLIP) under linear probing, LoRA, and a retrieval-augmented classification (RAC) module for two sequential tasks: Stage 1 binary leukemia detection and Stage 2 ALL/AML subtyping. Five public single-cell microscopy datasets are harmonized into three classes, and a held-out dataset protocol (ALL-IDB2 for Stage 1; ALL-IDB2 plus AML-Cytomorphology for Stage 2) is used to measure domain-shift generalization. Stage 1 reports best accuracy of 0.9423 for DinoBloom with RAC and shows LoRA narrowing the pretraining gap between CLIP and DinoBloom. Stage 2 collapses to the AML class under the held-out protocol (ALL recall near zero in Table 8) while a within-domain random 80/20 split achieves near-perfect ALL recall (Table 9), a contrast the paper attributes to dataset-specific background shortcuts learned during training. The paper concludes that cost-effective adaptation can compensate for most of the advantage of domain-specialized pretraining, and that the held-out protocol is a diagnostic tool for exposure of dataset-specific artifacts.","tokens_in":19139,"tokens_out":12937,"duration_ms":116896,"significance":"The most valuable element of this paper is the Stage 2 negative result: Table 8 versus Table 9 shows that near-perfect in-domain performance (ALL recall up to 1.0000 for frozen CLIP) coexists with near-zero held-out ALL recall across all three encoders, a clean and honestly reported demonstration that within-split evaluation overstates robustness. The label harmonization protocol of Section 4.3 with its explicit inclusion criteria is a concrete reusable contribution, and the paper ships a sensible control design (the random-split experiment) and a candid limitations statement (Section 5.6). If the Stage 1 claims were statistically supported, the paper would provide a useful controlled comparison of pretraining specialization, PEFT, and retrieval augmentation in hematology imaging. As it stands, the Stage 1 comparisons rest on a 260-image test set without confidence intervals, and the Stage 2 causal mechanism (synthetic-background shortcut) is not isolated from other source-level confounds; these two load-bearing points need work before the robustness framing can be accepted.","major_comments":[{"comment":"All Stage 1 conclusions about relative gains rest on the n=260 ALL-IDB2 held-out set, and no confidence intervals or significance tests are reported. At this sample size, the 0.0077 difference between DinoBloom LoRA and DinoBloom LoRA+RAC is about 2 images, the claimed closing of the pretraining gap from 0.2577 to 0.0193 is about 5 images, and the headline RAC gain of 0.0500 for DinoBloom linear probing is about 13 images; with binomial standard errors of roughly 0.016-0.028 at this n, most paired comparisons in Table 7 are indistinguishable from sampling noise. The paper should report bootstrap or exact confidence intervals for Tables 4 and 6 and restrict the conclusions in Section 6 (\"surpassed\", \"viable alternative\") to contrasts that survive interval comparison.","section":"Section 5.2-5.3, Tables 4, 6, 7"},{"comment":"The causal attribution of the Stage 2 collapse to a synthetic-background shortcut is not isolated from other confounds that vary with the subtype label. In training, ALL images come exclusively from D1 (black background replaced synthetically, Reinhard-normalized) and AML images exclusively from D4 (natural background, patient-level labels, four WHO 2022 genetic subtypes, and acknowledged instance-level label noise); in testing, ALL (D5) and AML (D2) differ in acquisition, stain, magnification, and cell-selection pipeline as well as background. No experiment separates background from these confounds, so Table 8 and Figure 12 establish that performance does not transfer and that source identity predicts the label, but not that background is the mechanism. The conclusive wording in Section 6 (\"This is attributable to each subtype class coming from a single dataset...\") exceeds the hedged \"a possible explanation\" in Section 5.4. A testable fix within the paper's scope would be to apply the Section 4.2 background-replacement procedure to the D5 ALL images at inference and check whether ALL recall recovers.","section":"Section 5.4 and Section 6"},{"comment":"The 96.5% AML-neighbor retrieval statistic is reported without correcting for the Stage 2 retrieval-bank composition: the bank contains 60,909 AML and 8,491 ALL training images (87.8% AML), so random retrieval alone would yield roughly 87.8% AML neighbors. The observed 96.5% is therefore a modest excess over the base rate, no per-query distribution or confidence bound is reported, and the statistic cannot discriminate the background-shortcut hypothesis from other dataset-specific cues that organize the embedding space. The alpha=8 sweep is likewise a consistency check rather than a mechanism test, since a top-20 retrieval set dominated by AML under many embedding geometries would behave identically.","section":"Section 5.4"},{"comment":"The held-out protocol is only partially held out for the domain-specialized encoder. DinoBloom's pretraining corpus is reported to include D2, D3, and D4, so the normal-class and AML appearance distributions in both training and evaluation have been seen during pretraining; only the ALL class (D1 excluded from pretraining, D5 held out) is genuinely unseen. The authors acknowledge this in Section 5.6, but it means the headline pretraining-specialization gap of 0.2577 in Table 5 conflates domain-specialized pretraining with direct exposure to the evaluation domains, and the abstract's \"robust... across multiple datasets\" framing should be tempered accordingly.","section":"Sections 4.7, 5.1, 5.6"}],"minor_comments":[{"comment":"The reported triple (accuracy 0.9423, recall 0.9423, F1 0.9307) is not jointly realizable on the balanced 130/130 held-out test under the macro-averaged definitions in Equations (3)-(5); please verify the confusion matrix or correct the rounding.","section":"Table 4, DinoBloom linear probe + RAC row"},{"comment":"The retrieval top-k is fixed at k=20 with no sensitivity analysis, and alpha is tuned on an internal validation split for linear probing but fixed at 0.5 for LoRA; the paper acknowledges the alpha inconsistency in Section 5.3, but Table 7 should mark the linear-probe and LoRA comparisons as not directly comparable.","section":"Section 4.6"},{"comment":"References [6] and [32] are traffic-sign classification papers cited to support statements about domain shift and VFM reusability in medical imaging; a domain-shift survey or a medical-imaging-specific citation would be more appropriate.","section":"Sections 1 and 3.2"},{"comment":"Although the paper describes a two-stage pipeline, the reported metrics are per-stage; no end-to-end evaluation propagates Stage 1 errors into Stage 2, so the performance claims should be labeled as per-stage to avoid implying system-level evaluation.","section":"Section 4.7"},{"comment":"The statement that \"training loss reached 0.0000\" should report the actual final loss values and their precision; an exactly zero cross-entropy value is unusual and may be a rounding artifact.","section":"Section 5.4"}],"recommendation":"major_revision","confidential_remarks":"The Stage 2 negative result is the genuinely valuable part of this manuscript and is reported honestly, but the paper is framed as a 'robust framework' benchmark when its own central finding is that subtype classification fails under domain shift; the authors should reposition the contribution as a diagnostic study. No code or model weights are released, which limits the reproducibility value of a benchmark-oriented paper. The novelty claim that no prior work combines cross-dataset training with held-out evaluation for leukemia subtyping is contestable given Wang et al. (ref. [17]) report cross-platform training with held-out evaluation for acute leukemia morphology, and Ng et al. (ref. [23]) already combine DinoBloom, LoRA, and retrieval-based inference; the genuinely new combination here is the two-stage design plus the systematic encoder comparison, which is sufficient if framed correctly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: the paper's Stage 2 negative result is real and worth taking seriously. Under a held-out protocol, ALL/AML subtyping collapses to the AML class (ALL recall near zero) even though the same models hit near-perfect ALL recall on random within-domain splits. That contrast is a clean demonstration that high in-domain performance on this task can be driven by source-specific artifacts.\n\nWhat's new: they harmonize labels across five public single-cell leukemia datasets, run three encoders (DinoBloom, BiomedCLIP, CLIP) under linear probe, LoRA, and retrieval-augmented classification, and use a held-out dataset protocol instead of an inner split. The closest prior benchmark (Ben Rabah and Serag) explicitly excluded leukemia from cross-dataset evaluation because of label inconsistencies; this paper fills that gap. The control experiment (random 80/20 vs held-out) is the right kind of diagnostic and strengthens the interpretation. The citation pattern is sound, with the most relevant prior work covered.\n\nThe Stage 1 claim that pretraining specialization matters is also plausible: DinoBloom beats CLIP by 0.2577 accuracy under linear probing, and LoRA shrinks the gap to 0.0193. But the held-out set is 260 images, so differences like 0.9308 versus 0.9385 are two or three cells. There are no confidence intervals, and no code is released, so the numbers cannot be independently checked without reimplementation.\n\nSoft spots, in proportion. First, the title and abstract say \"robust framework,\" but the paper's own Stage 2 result is a collapse; that framing overstates. Second, the causal story for the collapse—that the model learned the synthetic smear background in the ALL training source (C-NMC) rather than morphology—is presented as likely in Section 5.4 but then treated as established in the conclusion. The evidence supports a source-level shortcut, not specifically the background: acquisition, stain, magnification, and label noise all vary with the ALL/AML label. The 96.5% AML-neighbor retrieval statistic is also not corrected for the 87.8% AML majority in the retrieval bank, though the excess over the base rate still points somewhere. The paper itself hedges with \"a possible explanation,\" so this is a wording problem as much as an evidence problem.\n\nWho this is for: anyone working on medical image classification under domain shift, and especially computational hematopathology. The negative result is a useful calibration for the field. It deserves a serious referee; the concerns are addressable in revision, not fatal.","headline":"The Stage 2 ALL/AML collapse under held-out evaluation is a real and useful negative result, but the paper overstates its causal mechanism and its 'robust' framing; still worth refereeing.","tokens_in":19733,"tokens_out":2637,"would_cite":true,"duration_ms":26114,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-stage leukemia classification pipeline reaches 94.23% held-out accuracy for detection, but ALL/AML subtyping collapses to the majority class under the same protocol, which the paper attributes to dataset-specific background shortcuts.","keywords":["leukemia","single-cell microscopy","vision foundation models","domain shift","retrieval-augmented classification","LoRA","held-out evaluation","ALL/AML subtyping"],"falsifier":"Replace the backgrounds of the 130 held-out ALL-IDB2 cells with the same synthetic healthy-smear background used for C-NMC 2019 training images and retest the Stage 2 classifiers: if ALL recall stays near zero the background-mismatch explanation is unsupported, whereas a large jump would confirm it; alternatively, a held-out set containing ALL and AML cells imaged on the same platform would settle whether subtype classification can transfer once the background confound is removed.","tokens_in":18631,"feed_emoji":"🔬","tokens_out":11373,"duration_ms":98802,"temperature":0.7,"pith_summary":"Leukemia cell classification is often trained and tested on one dataset, so impressive accuracy may reflect dataset-specific artifacts rather than real morphology. This paper constructs a two-stage pipeline over five heterogeneous single-cell microscopy datasets and evaluates it on held-out datasets that are never seen during training or validation. Stage 1, leukemia versus healthy cells, reaches 94.23% accuracy with a domain-specialized encoder plus retrieval augmentation, and the paper shows that a general-purpose encoder adapted with LoRA can nearly close the pretraining gap. Stage 2, ALL versus AML, collapses under the same protocol: ALL recall falls to between 0 and 2.3% even though the same models reach 99.7–100% ALL recall on random within-domain splits. The paper argues that this contrast is evidence that in-domain performance is driven by dataset-specific shortcuts, and that a held-out protocol is therefore a necessary diagnostic for robustness in medical image classification.","feed_headline":"Leukemia subtype AI fails on unseen data despite 99% in-domain scores","feed_subtitle":"Within-domain splits near 100% ALL recall, but held-out recall drops to 0–2%, exposing dataset shortcuts.","key_machinery":"The load-bearing mechanism is the held-out dataset evaluation protocol: each stage trains on four (Stage 1) or two (Stage 2) datasets and is tested on a dataset never seen in training or validation, so that any surviving accuracy must transfer across acquisition, staining, and background differences. Within that protocol, retrieval-augmented classification (RAC) is the mechanism that attempts to ground predictions in cytomorphology: the query image's embedding is compared by cosine similarity to a bank of labeled training embeddings, the top-$k=20$ neighbors cast a similarity-weighted vote $p_{\\mathrm{ret}}$, and the final logits are $\\operatorname{logits}_{\\mathrm{final}} = \\log p_{\\mathrm{probe}} + \\alpha \\log p_{\\mathrm{ret}}$, with $\\alpha$ selected on an internal validation split for linear-probe configurations and fixed at $0.5$ for LoRA configurations. Low-Rank Adaptation (LoRA) is the cost-effective adaptation mechanism: it freezes the pretrained encoder and trains two low-rank matrices $B$ and $A$ whose product $BA$ is added to the attention projection weights, so only about 0.3 to 0.5 million parameters are trained instead of roughly 86 million. Label harmonization, collapsing five heterogeneous datasets into normal/ALL/AML classes under a uniform inclusion criterion, is what makes cross-dataset training possible, and the synthetic-background replacement of C-NMC 2019 images is the preprocessing step that the paper identifies as the source of the Stage 2 shortcut.","core_discovery":"The paper's central claim is that under a held-out dataset protocol, the value of domain-specialized pretraining is large but mostly replaceable, and the protocol itself exposes a failure that random splits hide. With frozen encoders and linear probing, the hematology-pretrained encoder beats the general-purpose encoder by 0.2577 accuracy on Stage 1 (0.8923 versus 0.6346), but LoRA adaptation shrinks that gap to 0.0193 and lifts the general-purpose encoder to 0.9115, above the specialized encoder's frozen 0.8923; adding retrieval to the specialized encoder gives the best Stage 1 result, 0.9423. In Stage 2, the held-out protocol reveals complete collapse: on a test set of 3,424 cells (130 ALL, 3,294 AML), ALL recall is 0.0000 for the two biomedical encoders and at most 0.0231 for the general-purpose encoder, despite a within-domain random split yielding ALL recall of 0.997 to 1.000 and training loss converging to 0.0000. The paper attributes this to shortcut learning on the synthetic background of the training ALL source, and reports that 96.5% of retrieved neighbors for query ALL images belong to the AML class, consistent with the encoder grouping cells by acquisition source rather than morphology.","pith_inferences":["The paper does not test this, but the same held-out logic implies that any classifier whose classes are confounded with acquisition source is vulnerable to collapse, so the protocol could serve as a general shortcut detector across medical imaging tasks.","A decisive experiment would be to apply the synthetic-background normalization to the held-out ALL-IDB2 cells: if ALL recall recovers substantially, the background is confirmed as the shortcut; if not, the bottleneck lies in genuine cytomorphological similarity between ALL and AML blasts.","Because RAC's benefit tracked encoder specialization, retrieval neighbor purity on a validation set could be used as a cheap predictor of whether to fuse retrieval at inference time, avoiding the observed accuracy drops on general-purpose encoders.","The paper's framework suggests that multi-center datasets that include both subtypes under shared acquisition protocols would be the next enabling resource; until such data exist, ALL/AML single-cell classification should be reported with a domain-shift evaluation rather than only a random split."],"forward_implications":["A general-purpose vision encoder fine-tuned with LoRA can substitute for a domain-specialized encoder in leukemia detection, so expensive hematology-specific pretraining is not strictly required when a modest fine-tuning budget exists.","Retrieval augmentation is only useful when the encoder's embedding space already reflects cytomorphology; on general-purpose encoders it can reduce accuracy, so RAC should be enabled based on retrieval neighbor purity rather than by default.","Within-domain accuracy in subtype classification is not a reliable measure of clinical readiness: the same classifiers that score 99.7–100% ALL recall on random splits collapse to 0–2.3% on a held-out dataset.","Label harmonization across heterogeneous single-cell datasets enables cross-dataset training, but when each subtype comes from a single imaging source, the model can still learn source-specific shortcuts instead of morphology.","A practical consequence for screening is that Stage 1 can operate as a first-line triage with 94.23% held-out accuracy, while Stage 2 should not be deployed until datasets with both subtypes under one imaging protocol are available."],"supporting_citations":[{"why":"Supplies DinoBloom, the domain-specialized single-cell encoder whose pretraining advantage is the central quantity measured in both stages.","marker":"[7]"},{"why":"Supplies BiomedCLIP, the biomedical image-text encoder used as the intermediate pretraining-specialization baseline.","marker":"[8]"},{"why":"Supplies CLIP, the general-purpose encoder whose gap to DinoBloom quantifies the value of specialized pretraining.","marker":"[9]"},{"why":"Supplies ALL-IDB2, the held-out dataset used for Stage 1 evaluation and for the ALL half of the Stage 2 test set.","marker":"[31]"},{"why":"Supplies the retrieval-augmented classification method (top-k neighbor voting) that the paper adapts and benchmarks.","marker":"[40]"},{"why":"Supplies Low-Rank Adaptation (LoRA), the parameter-efficient fine-tuning method that closes most of the pretraining gap.","marker":"[36]"},{"why":"Supplies C-NMC 2019, the only training source of ALL cells, whose synthetic background replacement is central to the shortcut explanation for Stage 2 collapse.","marker":"[42]"},{"why":"Supplies the AML-Cytomorphology dataset, used as the AML source for Stage 2 training and as the AML half of the held-out test set.","marker":"[43]"},{"why":"The closest prior benchmark quantifying pretraining specialization; the paper extends it by harmonizing leukemia labels and adding a held-out protocol for hematology.","marker":"[24]"}],"fun_headline_variants":["Leukemia AI fails cross-dataset: held-out ALL recall drops to 0-2%","Dataset shortcuts expose leukemia classifier collapse under domain shift","Cheap adaptation beats specialized pretraining for leukemia domain shift","Retrieval helps stage 1 but not stage 2: leukemia subtype recall near zero","Held-out protocol reveals leukemia AI relies on dataset artifacts, not morphology"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The explanation of the Stage 2 collapse assumes that the main difference between the ALL cells used for training (with a synthetic replaced background) and the ALL cells used for testing (with a natural smear background) is the background itself, rather than differences in staining, magnification, cell selection, or the fact that the test set has 130 ALL cells against 3,294 AML cells.","fun_headline_variants_meta":{"raw":{"variants":["Leukemia AI fails cross-dataset: held-out ALL recall drops to 0-2%","Dataset shortcuts expose leukemia classifier collapse under domain shift","Cheap adaptation beats specialized pretraining for leukemia domain shift","Retrieval helps stage 1 but not stage 2: leukemia subtype recall near zero","Held-out protocol reveals leukemia AI relies on dataset artifacts, not morphology"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000782,"raw_usage":{"total_tokens":3547,"prompt_tokens":1129,"completion_tokens":2418,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":745,"completion_tokens_details":{"reasoning_tokens":2320}},"tokens_in":745,"tokens_out":2418,"duration_ms":17238,"temperature":1.0,"reasoning_tokens":2320,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:59:41.973982+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace the backgrounds of the 130 held-out ALL-IDB2 cells with the same synthetic healthy-smear background used for C-NMC 2019 training images and retest the Stage 2 classifiers: if ALL recall stays near zero the background-mismatch explanation is unsupported, whereas a large jump would confirm it; alternatively, a held-out set containing ALL and AML cells imaged on the same platform would settle whether subtype classification can transfer once the background confound is removed.","supporting_citations":[{"cited_title":"Dinobloom: A foundation model for generalizable cell embeddings in hematology,","cited_arxiv_id":null,"evidence_quote":"Supplies DinoBloom, the domain-specialized single-cell encoder whose pretraining advantage is the central quantity measured in both stages."},{"cited_title":"A multimodal biomedical foundation model trained from fifteen million image–text pairs,","cited_arxiv_id":null,"evidence_quote":"Supplies BiomedCLIP, the biomedical image-text encoder used as the intermediate pretraining-specialization baseline."},{"cited_title":"Learning transferable visual models from natural language super- vision,","cited_arxiv_id":null,"evidence_quote":"Supplies CLIP, the general-purpose encoder whose gap to DinoBloom quantifies the value of specialized pretraining."},{"cited_title":"All-idb: The acute lymphoblastic leukemia image database for image processing,","cited_arxiv_id":null,"evidence_quote":"Supplies ALL-IDB2, the held-out dataset used for Stage 1 evaluation and for the ALL half of the Stage 2 test set."},{"cited_title":"Retrieval Augmented Classification for Long-Tail Visual Recognition","cited_arxiv_id":"2202.11233","evidence_quote":"Supplies the retrieval-augmented classification method (top-k neighbor voting) that the paper adapts and benchmarks."},{"cited_title":"All challenge dataset of ISBI 2019 (C-NMC 2019),","cited_arxiv_id":null,"evidence_quote":"Supplies C-NMC 2019, the only training source of ALL cells, whose synthetic background replacement is central to the shortcut explanation for Stage 2 collapse."},{"cited_title":"A single-cell morphological dataset of leukocytes from aml patients and non-malignant controls,","cited_arxiv_id":null,"evidence_quote":"Supplies the AML-Cytomorphology dataset, used as the AML source for Stage 2 training and as the AML half of the held-out test set."},{"cited_title":"Do foundation models truly outperform domain-specific models? evidence from digital pathology,","cited_arxiv_id":null,"evidence_quote":"The closest prior benchmark quantifying pretraining specialization; the paper extends it by harmonizing leukemia labels and adding a held-out protocol for hematology."}],"review_version":1}