Pith. sign in

REVIEW 2 cited by

CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.12077 v1 pith:ATIVX2HW submitted 2024-12-16 cs.CV

classification cs.CV
keywords cpath-omnimodelspathologymodeltasksimagevisualachieves
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The emergence of large multimodal models (LMMs) has brought significant advancements to pathology. Previous research has primarily focused on separately training patch-level and whole-slide image (WSI)-level models, limiting the integration of learned knowledge across patches and WSIs, and resulting in redundant models. In this work, we introduce CPath-Omni, the first 15-billion-parameter LMM designed to unify both patch and WSI level image analysis, consolidating a variety of tasks at both levels, including classification, visual question answering, captioning, and visual referring prompting. Extensive experiments demonstrate that CPath-Omni achieves state-of-the-art (SOTA) performance across seven diverse tasks on 39 out of 42 datasets, outperforming or matching task-specific models trained for individual tasks. Additionally, we develop a specialized pathology CLIP-based visual processor for CPath-Omni, CPath-CLIP, which, for the first time, integrates different vision models and incorporates a large language model as a text encoder to build a more powerful CLIP model, which achieves SOTA performance on nine zero-shot and four few-shot datasets. Our findings highlight CPath-Omni's ability to unify diverse pathology tasks, demonstrating its potential to streamline and advance the field of foundation model in pathology.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A bilateral RL framework with a GRPO-trained reasoning branch and an RL-trained token allocator improves pathology VQA, subtyping, and detection while cutting tokens per image.

  2. GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning

    cs.CV 2025-07 reject novelty 4.0 of 10

    GNN-ViTCap combines deep embedded clustering, graph-based aggregation, and large language models to classify and caption microscopic whole slide images, reporting high F1, AUC, BLEU, and METEOR scores on BreakHis and ...

Pith tools