Pith. sign in

REVIEW 1 cited by

Anatomical Structure-Guided Medical Vision-Language Pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.09294 v1 pith:FM2USLOC submitted 2024-03-14 cs.CV cs.CL

classification cs.CVcs.CL
keywords anatomicallearningalignmentexistencefindingframeworkimageimage-report
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning medical visual representations through vision-language pre-training has reached remarkable progress. Despite the promising performance, it still faces challenges, i.e., local alignment lacks interpretability and clinical relevance, and the insufficient internal and external representation learning of image-report pairs. To address these issues, we propose an Anatomical Structure-Guided (ASG) framework. Specifically, we parse raw reports into triplets <anatomical region, finding, existence>, and fully utilize each element as supervision to enhance representation learning. For anatomical region, we design an automatic anatomical region-sentence alignment paradigm in collaboration with radiologists, considering them as the minimum semantic units to explore fine-grained local alignment. For finding and existence, we regard them as image tags, applying an image-tag recognition decoder to associate image features with their respective tags within each sample and constructing soft labels for contrastive learning to improve the semantic association of different image-report pairs. We evaluate the proposed ASG framework on two downstream tasks, including five public benchmarks. Experimental results demonstrate that our method outperforms the state-of-the-art methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training

    eess.IV 2025-08 conditional novelty 6.0 of 10

    On abdominal CT, a vision-language pre-training method using organ-level normal/abnormal contrastive learning and a VQ-VAE normality model achieves 84.9% average zero-shot AUC, beating prior methods by 3.6%.

Pith tools