REVIEW 3 cited by
Advancing Radiograph Representation Learning with Masked Record Modeling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Advancing Radiograph Representation Learning with Masked Record Modeling
read the original abstract
Modern studies in radiograph representation learning rely on either self-supervision to encode invariant semantics or associated radiology reports to incorporate medical expertise, while the complementarity between them is barely noticed. To explore this, we formulate the self- and report-completion as two complementary objectives and present a unified framework based on masked record modeling (MRM). In practice, MRM reconstructs masked image patches and masked report tokens following a multi-task scheme to learn knowledge-enhanced semantic representations. With MRM pre-training, we obtain pre-trained models that can be well transferred to various radiography tasks. Specifically, we find that MRM offers superior performance in label-efficient fine-tuning. For instance, MRM achieves 88.5% mean AUC on CheXpert using 1% labeled data, outperforming previous R$^2$L methods with 100% labels. On NIH ChestX-ray, MRM outperforms the best performing counterpart by about 3% under small labeling ratios. Besides, MRM surpasses self- and report-supervised pre-training in identifying the pneumonia type and the pneumothorax area, sometimes by large margins.
Forward citations
Cited by 3 Pith papers
-
Super-Generalist: Towards Comprehensive and Accurate Medical Image Understanding via Generalist-Specialist Synergy
Injecting multi-expert anatomy/lesion segmentation priors into vision–language alignment and calibrating text attention with lesion masks yields broad CT diagnosis plus specialist-level tumor performance and lesion grounding.
-
RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture
RadJEPA learns chest X-ray encoders from unlabeled images via latent prediction in a joint embedding architecture, exceeding prior state-of-the-art on classification, segmentation, and report generation.
-
RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture
A JEPA-style encoder pretrained on 839k unlabeled chest X-rays matches or exceeds vision-language and DINO baselines on CXR classification, segmentation, and report generation.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.