Pith. sign in

REVIEW 5 cited by

Advancing human-centric AI for robust X-ray analysis through holistic self-supervised learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.01469 v1 pith:2RMXAUH6 submitted 2024-05-02 cs.CV cs.AI

Advancing human-centric AI for robust X-ray analysis through holistic self-supervised learning

classification cs.CV cs.AI
keywords modelsfoundationraydinoanalysisbiasesmedicalradiologyself-supervision
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

AI Foundation models are gaining traction in various applications, including medical fields like radiology. However, medical foundation models are often tested on limited tasks, leaving their generalisability and biases unexplored. We present RayDINO, a large visual encoder trained by self-supervision on 873k chest X-rays. We compare RayDINO to previous state-of-the-art models across nine radiology tasks, from classification and dense segmentation to text generation, and provide an in depth analysis of population, age and sex biases of our model. Our findings suggest that self-supervision allows patient-centric AI proving useful in clinical workflows and interpreting X-rays holistically. With RayDINO and small task-specific adapters, we reach state-of-the-art results and improve generalization to unseen populations while mitigating bias, illustrating the true promise of foundation models: versatility and robustness.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Capability $\neq$ Interpretability: Human Interpretability of Vision Foundation Models

    cs.CV 2026-05 conditional novelty 7.0

    Foundation models yield less human-interpretable features than supervised vision transformers, with interpretability tied to activation locality and coarse semantic alignment rather than task performance.

  2. Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift

    cs.CV 2026-07 conditional novelty 6.0

    Mammography-specific VLMs lead mean OOD linear-probe performance across 15 datasets, but robustness depends on pretraining objective and is highly dataset-heterogeneous, not on mammography exposure alone.

  3. Mirror-Fusion Attention for Reflection-Aware Self-Supervised Representation Learning

    cs.CV 2026-07 unverdicted novelty 6.0

    MFASSL adds mirror-paired views, a lightweight Mirror-Fusion Attention module, and reflection-consistency losses to improve SSL on bilateral data with ~2.7% extra parameters.

  4. Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have

    cs.CV 2026-06 unverdicted novelty 6.0

    FINO adapts vision foundation models to scientific domains via metadata-guided self-supervised learning and outperforms both unsupervised domain adaptation and fully supervised methods without using task labels for th...

  5. Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation

    cs.CV 2026-02 unverdicted novelty 6.0

    MARL-Rad trains region-specific and global agents with reinforcement learning on clinical rewards to produce more accurate radiology reports than prior methods on MIMIC-CXR and IU X-ray datasets.