Pith. sign in

REVIEW 2 cited by

PathInsight: Instruction Tuning of Multimodal Datasets and Models for Intelligence Assisted Diagnosis in Histopathology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.07037 v1 pith:GA4SUQEA submitted 2024-08-13 cs.CV cs.AI

classification cs.CVcs.AI
keywords modelsmultimodaldatasetdatasetsfine-tunedmodeladdressingclassification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pathological diagnosis remains the definitive standard for identifying tumors. The rise of multimodal large models has simplified the process of integrating image analysis with textual descriptions. Despite this advancement, the substantial costs associated with training and deploying these complex multimodal models, together with a scarcity of high-quality training datasets, create a significant divide between cutting-edge technology and its application in the clinical setting. We had meticulously compiled a dataset of approximately 45,000 cases, covering over 6 different tasks, including the classification of organ tissues, generating pathology report descriptions, and addressing pathology-related questions and answers. We have fine-tuned multimodal large models, specifically LLaVA, Qwen-VL, InternLM, with this dataset to enhance instruction-based performance. We conducted a qualitative assessment of the capabilities of the base model and the fine-tuned model in performing image captioning and classification tasks on the specific dataset. The evaluation results demonstrate that the fine-tuned model exhibits proficiency in addressing typical pathological questions. We hope that by making both our models and datasets publicly available, they can be valuable to the medical and research communities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient and Comprehensive Feature Extraction in Large Vision-Language Model for Pathology Analysis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A pathology-specialized LVLM with mixed-task feature enhancement and attention-guided detail completion reports higher accuracy than existing LVLMs across many diagnostic tasks.

  2. WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image

    cs.CV 2024-12 reject novelty 6.0 of 10

    WSI-LLaVA, trained on a large AI-generated WSI question-answer benchmark, is reported to outperform prior models on morphology and diagnosis, though the evaluation loop is largely closed through GPT-4o.

Pith tools