{"id":"c3989543-9201-4758-b010-da636b2343ba","arxiv_id":"2501.15724","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic literature review of computational pathology foundation models that catalogs datasets, SSL adaptation strategies, and evaluation tasks into taxonomies.","lead":"This paper surveys foundation models for computational pathology, covering pre-training datasets, adaptation strategies, and evaluation tasks. It organizes 28 models into uni-modal and multi-modal groups and proposes a six-way taxonomy of evaluation tasks.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's availability column contains demonstrably wrong entries (Phikon, Phikon-v2, RudolfV), undermining the reliability of the survey's central structured map.","rationale":"The reader's CONDITIONAL verdict identifies the lack of search methodology and inclusion criteria as the weakest assumption, making the survey's completeness claim unverifiable. My stress-test finds a more immediate and more concrete problem: the central Table 2 contains factually wrong availability information for at least three well-known public models. This is an accuracy failure, not merely a methodological gap, and it directly attacks the strongest claim that the paper 'provides a structured map of 28 foundation models.' The reader's internal-summaries-are-consistent observation holds for the prose descriptions, but the table that operationalizes the survey's contribution is unreliable in a column that matters for practical use. The error is fixable, so it does not justify rejecting the survey; it does justify keeping the CONDITIONAL verdict and adding a concrete correction requirement. I also agree with the reader that a systematic literature search test would settle the completeness question, but the availability audit is the one check that most decisively tests the paper's central factual claim because it can be falsified by direct inspection of public repositories. This is why my verdict recommendation remains UNCHANGED rather than escalating to REJECT: the concerns are addressable and do not invalidate the survey's overall organizational contribution.","tokens_in":14603,"tokens_out":6815,"duration_ms":65846,"concrete_test":"Independently verify the 'Availability' column of Table 2 by visiting the official repositories/HuggingFace pages for all 28 rows. Specifically, check the model pages for Phikon (owkin/phikon), Phikon-v2 (owkin/phikon-v2), and RudolfV (Aignostics/rudolfv). If all three are confirmed publicly downloadable while Table 2 marks them '✗', then the table is falsified on 3/28 entries. Then audit the remaining 25 rows the same way. If three or more entries are wrong, the survey's central structured map must be corrected before the comprehensiveness claim can be evaluated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central claim is to provide a comprehensive and accurate structured review of 28 CPathFMs. The most immediately checkable part of that claim is Table 2, whose 'Availability' column is wrong for at least three prominent models: Phikon (Filiot et al., 2023) has public weights on HuggingFace, Phikon-v2 (Filiot et al., 2024) is explicitly released, and RudolfV (Dippel et al., 2024) is distributed under a public model license, yet all three are marked '✗' in Table 2. These are not peripheral entries; Phikon and Phikon-v2 are widely used public baselines, and RudolfV is a notable open-weights pathology model. A survey whose main contribution is the structured comparison in Tables 1–2 cannot support its 'valuable resource' claim if the availability column is wrong on roughly 10% of rows. This factual unreliability is independent of the reader's coverage concern: even if the model list were exhaustive, the map would still mislead practitioners. The absence of explicit inclusion criteria (Sections 4–5) compounds the problem by making the table's correctness and completeness un-auditable, but the availability errors are sufficient on their own to require revision.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews computational pathology foundation models (CPathFMs), covering 28 models, their pre-training datasets, adaptation strategies, and evaluation tasks. It organizes models into uni-modal and multi-modal categories, tabulates pre-training datasets in Table 1, summarizes architectures and adaptation strategies in Table 2, and introduces a six-category taxonomy of evaluation tasks in Figure 3. The paper also discusses data, adaptation, and evaluation challenges and proposes future directions such as trustworthy CPathFMs, MxIF imaging, and standardized benchmarking.","tokens_in":14874,"tokens_out":6130,"duration_ms":54344,"significance":"A well-curated survey of CPathFMs would be valuable to a broad community. The paper compiles a substantial amount of information in compact form, and the proposed taxonomy of evaluation tasks is a useful organizing device. However, because the survey's main deliverable is the structured map in Table 2, factual errors in the availability column—at least Phikon, Phikon-v2, and RudolfV are incorrectly marked as unavailable—directly undermine its reliability. The absence of a stated methodology also weakens the comprehensiveness claim. With corrections and a methodology section, the paper could serve as a useful reference.","major_comments":[{"comment":"The Availability column lists Phikon (Filiot et al., 2023), Phikon-v2 (Filiot et al., 2024), and RudolfV (Dippel et al., 2024) as unavailable, but all three models have public weights: Phikon and Phikon-v2 are distributed on HuggingFace, and RudolfV is released under a public model license. These are not peripheral entries; Phikon and Phikon-v2 are widely used public baselines. Because Table 2 is the paper's central structured comparison, these errors mislead practitioners and must be corrected, and every other row should be re-verified against primary sources.","section":"Table 2"},{"comment":"The paper claims to provide a 'comprehensive' review of 28 existing and up-to-date models and to summarize evaluation tasks 'for the first time,' but it gives no inclusion/exclusion criteria, search strategy, or date cutoff. This makes the completeness claim unauditable: a reader cannot determine why these 28 models were chosen or whether the six evaluation categories in Figure 3 are exhaustive. A methodology subsection describing how models and tasks were selected, categorized, and cross-checked is needed.","section":"Sections 4–5 and Introduction"},{"comment":"The evaluation taxonomy mixes task type, granularity, and experimental setting in a single hierarchy. For example, 'OOD generalization' and 'imbalanced' are listed as subcategories under classification alongside tile-level and WSI-level granularity, while supervised/zero-shot/few-shot settings appear as parallel dimensions. This conflation makes the claimed 'six main perspectives' less well-defined and the model assignments harder to interpret. Clarify the dimensions of the taxonomy or present them as separate axes.","section":"Section 5 and Figure 3"}],"minor_comments":[{"comment":"The header contains 'A vailability' with an erroneous space; it should read 'Availability'.","section":"Table 2"},{"comment":"The sentence 'In addition to qualitative analysis, some CPathFMs have undergone qualitative analysis' appears to use 'qualitative' twice; one instance should likely be 'quantitative'.","section":"Section 5"},{"comment":"The Phikon-v2 row lists 'Proprietary×4' as a data source, which is unclear; specify the four proprietary sources or explain the notation in a footnote.","section":"Table 1"},{"comment":"The figure is densely packed; consider increasing font size and using labels or patterns in addition to color to distinguish uni-modal and multi-modal models for accessibility.","section":"Figure 3"},{"comment":"Reference [Zhou et al., 2024a] contains the malformed author string 'Lifeng others Wang'; this should be 'Lifeng Wang et al.'","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The availability errors in Table 2 are concrete and checkable. If the authors correct those entries, verify the rest of the table, and add a methodology section, the paper would be publishable as a survey. I would not reject solely for lacking a formal systematic-review protocol, but the factual errors must be fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things. First, this is a genuinely useful survey: it covers 28 CPathFMs including recent ones like Atlas and MUSK, organizes adaptation strategies into uni-modal and multi-modal, and proposes a six-category evaluation taxonomy. Second, the central comparison table has at least three wrong entries: Phikon, Phikon-v2, and RudolfV are all marked as unavailable, but all three have public weights. That is a factual error in the most checked part of the paper, and it needs fixing before this can serve as a reliable reference.\n\nWhat's new: the coverage is more current than prior surveys (through early 2025), and the explicit breakdown of adaptation strategies—domain-specific tuning, from-scratch, frozen—plus the evaluation taxonomy in Figure 3, are useful organizational contributions. The dataset table is a good compilation, and the descriptive summaries of individual models read accurately to me.\n\nMain weaknesses: the reliability of Table 2 is the biggest issue. Beyond the three availability errors, the 'Pre-training Strategy' column uses shorthand (S/D/F) that is not fully defined for every model. The paper also claims to be 'comprehensive' but gives no search strategy, inclusion criteria, or date cutoff, so the 28-model list is un-auditable. The evaluation taxonomy is asserted without comparing it to prior frameworks, making 'for the first time' a strong claim. These are addressable: fix the table, add a methodology paragraph, and soften the claims.\n\nWho is this for: students and practitioners wanting a map of the field up to early 2025. It is not a benchmark paper and produces no new measurements, so it will not change research directions. But as a reference it has clear value once corrected.\n\nRecommendation: yes, send to peer review, but conditional on fixing the availability column and adding inclusion criteria. A desk reject would be wrong because the paper fills a real coverage gap; an accept without revision would also be wrong because the factual errors are checkable and misleading.","headline":"Useful survey with up-to-date coverage, but Table 2 has verifiable availability errors and the 'comprehensive' claim needs explicit inclusion criteria.","tokens_in":15334,"tokens_out":1726,"would_cite":false,"duration_ms":15186,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey maps 28 computational pathology foundation models through three lenses: pre-training datasets, adaptation strategies, and evaluation tasks.","keywords":["computational pathology","foundation models","self-supervised learning","whole-slide images","vision-language models","evaluation taxonomy","histopathology datasets","adaptation strategies"],"falsifier":"A reader could check Table 2 against the full list of publicly released computational pathology foundation models as of the survey's final revision, or check Figure 3's six categories against every evaluation task reported in the cited model papers; finding a released model with a distinct self-supervised backbone that is absent from the table, or an evaluation task that fits none of the six categories, would falsify the survey's coverage claim.","tokens_in":14446,"feed_emoji":"🔬","tokens_out":5083,"duration_ms":45195,"temperature":0.7,"pith_summary":"This survey tries to give the field of computational pathology foundation models a usable map: which datasets are used to pre-train them, how the models adapt general self-supervised learning frameworks to pathology, and how the research community evaluates them. It claims that although the models differ in scale and modality, they fall into two paradigms—uni-modal models trained only on images and multi-modal models that align images with text—and that their evaluations can be summarized into six task categories. A sympathetic reader would care because the field lacks standardized benchmarks; a structured comparison is a step toward measuring which of these models actually generalizes to clinical use.","feed_headline":"Survey maps 28 pathology AI models and how they are tested","feed_subtitle":"It groups pre-training data, adaptation strategies, and six evaluation tasks into one comparison map.","key_machinery":"The organizing device is a two-axis classification: uni-modal versus multi-modal pre-training, crossed with a six-category taxonomy of evaluation tasks. The survey constructs an architecture-and-adaptation matrix that records each model's self-supervised backbone, parameter count, input type, and pre-training strategy, and a task taxonomy that lists which models were evaluated on each task. This matrix-plus-taxonomy structure is what carries the review's claims of comprehensiveness and comparability.","core_discovery":"The central claim is organizational: 28 current computational pathology foundation models can be systematically reviewed through three lenses—pre-training datasets, adaptation strategies, and evaluation tasks—and doing so reveals that the field has converged on a small number of self-supervised learning frameworks. Uni-modal models mostly adapt DINO, DINOv2, or masked image modeling frameworks, while multi-modal models mostly adapt CLIP or CoCa; their evaluations cluster into classification, retrieval, generation, segmentation, prediction, and visual question answering. The paper further argues that no standardized benchmark exists and that this lack is the main obstacle to comparing models across institutions and tasks.","pith_inferences":["The paper's own tables suggest that private, large-scale slide collections dominate the largest models, so the public datasets it lists may understate the field's real data advantage; that reading goes beyond what the paper states explicitly.","A testable extension is to convert the evaluation-task taxonomy into a checklist and use it to score future computational pathology foundation model papers for evaluation coverage.","The six task categories may overlap—report generation and visual question answering both test cross-modal understanding—so a future taxonomy might merge them or add a separate cross-modal reasoning axis.","Because the survey excludes generative pathology assistants and non-histopathology multimodal models, its map is deliberately narrower than all medical foundation models; readers should not generalize beyond the histopathology scope the paper defines."],"forward_implications":["Researchers can use the six-category taxonomy to position a new model's evaluation and see which task categories are under-tested.","The dominance of DINOv2 for uni-modal models and CLIP/CoCa for multi-modal models suggests that future models will likely build on these backbones rather than invent entirely new pre-training frameworks.","The lack of standardized benchmarks implies that results across papers remain not directly comparable until a common evaluation suite is adopted.","Coverage gaps, such as few models trained on multiplex immunofluorescence images, point to specific data types where foundation models are still missing.","The survey's future directions imply that clinical deployment will require work on fairness, explainability, security, and transparency before these models can be trusted in practice."],"supporting_citations":[{"why":"Supplies the large public cancer slide cohort that most surveyed uni-modal and multi-modal models use as a pre-training source.","marker":"[Weinstein et al., 2013]"},{"why":"Supplies the genotype-tissue expression slide collection used by several surveyed models to broaden pre-training diversity.","marker":"[Consortium et al., 2015]"},{"why":"Introduces masked autoencoders, the masked image modeling framework that the survey tracks as one of the main self-supervised adaptation routes.","marker":"[He et al., 2022]"},{"why":"Presents DINOv2, the dominant backbone for the surveyed uni-modal computational pathology foundation models.","marker":"[Oquab et al., 2023]"},{"why":"Presents UNI, a reference DINOv2-based uni-modal model whose scale, data, and evaluation the survey repeatedly uses as a comparison point.","marker":"[Chen et al., 2024]"},{"why":"Presents Virchow, an example of a very large-scale uni-modal model trained on private slide data, illustrating the private-data trend in the survey's dataset analysis.","marker":"[Vorontsov et al., 2023]"},{"why":"Presents PLIP, a CLIP-based multi-modal model that the survey categorizes under its vision-language adaptation strategy.","marker":"[Huang et al., 2023]"},{"why":"Presents CONCH, a CoCa-based multi-modal model that the survey uses as an exemplar of the CoCa adaptation route.","marker":"[Lu et al., 2024]"},{"why":"Presents GigaPath, a whole-slide-level model that the survey uses to illustrate slide-level evaluation and long-context adaptation.","marker":"[Xu et al., 2024]"},{"why":"Presents TITAN, a multi-modal whole-slide foundation model whose three-stage pre-training the survey describes as an advanced adaptation strategy.","marker":"[Ding et al., 2024]"}],"fun_headline_variants":["28 pathology AI models, zero standard benchmarks","Pathology foundation models lack a common test","Survey: 28 CPathFMs, no benchmark to compare them","Pathology AI: 28 models, one missing piece—benchmarks","How 28 pathology models train and test—survey"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's map is only as complete as the authors' informal selection of 28 models and six task categories; no systematic search strategy or inclusion criteria is given, so the claim of comprehensiveness rests on an unstated judgment about which models and tasks count.","fun_headline_variants_meta":{"raw":{"variants":["28 pathology AI models, zero standard benchmarks","Pathology foundation models lack a common test","Survey: 28 CPathFMs, no benchmark to compare them","Pathology AI: 28 models, one missing piece—benchmarks","How 28 pathology models train and test—survey"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000338,"raw_usage":{"total_tokens":1818,"prompt_tokens":848,"completion_tokens":970,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":890}},"tokens_in":464,"tokens_out":970,"duration_ms":9457,"temperature":1.0,"reasoning_tokens":890,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:59:06.271468+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could check Table 2 against the full list of publicly released computational pathology foundation models as of the survey's final revision, or check Figure 3's six categories against every evaluation task reported in the cited model papers; finding a released model with a distinct self-supervised backbone that is absent from the table, or an evaluation task that fits none of the six categories, would falsify the survey's coverage claim.","supporting_citations":[{"cited_title":"The cancer genome atlas pan-cancer analysis project","cited_arxiv_id":null,"evidence_quote":"Supplies the large public cancer slide cohort that most surveyed uni-modal and multi-modal models use as a pre-training source."},{"cited_title":"The genotype-tissue expression (gtex) pilot analysis: multitis- sue gene regulation in humans","cited_arxiv_id":null,"evidence_quote":"Supplies the genotype-tissue expression slide collection used by several surveyed models to broaden pre-training diversity."},{"cited_title":"Towards a general-purpose foundation model for computational pathology","cited_arxiv_id":null,"evidence_quote":"Presents UNI, a reference DINOv2-based uni-modal model whose scale, data, and evaluation the survey repeatedly uses as a comparison point."},{"cited_title":"A visual–language foundation model for pathology image analysis using medical twitter","cited_arxiv_id":null,"evidence_quote":"Presents PLIP, a CLIP-based multi-modal model that the survey categorizes under its vision-language adaptation strategy."}],"review_version":1}