Pith. sign in

REVIEW 1 cited by

MedICaT: A Dataset of Medical Images, Captions, and Textual References

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.06000 v1 pith:7AI2RBVW submitted 2020-10-12 cs.CV cs.CL

classification cs.CVcs.CL
keywords figuresmedicatimagesdatasetmedicalreferencestextunderstanding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Understanding the relationship between figures and text is key to scientific document understanding. Medical figures in particular are quite complex, often consisting of several subfigures (75% of figures in our dataset), with detailed text describing their content. Previous work studying figures in scientific papers focused on classifying figure content rather than understanding how images relate to the text. To address challenges in figure retrieval and figure-to-text alignment, we introduce MedICaT, a dataset of medical images in context. MedICaT consists of 217K images from 131K open access biomedical papers, and includes captions, inline references for 74% of figures, and manually annotated subfigures and subcaptions for a subset of figures. Using MedICaT, we introduce the task of subfigure to subcaption alignment in compound figures and demonstrate the utility of inline references in image-text matching. Our data and code can be accessed at https://github.com/allenai/medicat.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedVision: Benchmarking Quantitative Medical Image Analysis

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A 30.8M-pair medical imaging benchmark shows pretrained vision-language models are poor at detection, tumor-size, and angle/distance measurement, and that fine-tuning on the benchmark substantially improves them.

Pith tools