{"work":{"id":"91ab4cd1-413f-4e31-be3f-78e309d19006","openalex_id":"https://openalex.org/W4391159111","doi":"10.48550/arxiv.2401.12208","arxiv_id":"2401.12208","raw_key":null,"title":"A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation","authors":null,"authors_text":"Zhihong Chen, Maya Varma, Jean-Benoit Del- brouck, Magdalini Paschali, Louis Blanke- meier, Dave Van Veen, Jeya Maria Jose Vala- narasu, Alaa Youssef, Joseph Paul Cohen, Ed- uardo Pontes Reis, Emily B","year":2024,"venue":"cs.CV","abstract":"Over 1.4 billion chest X-rays (CXRs) are performed annually due to their cost-effectiveness as an initial diagnostic test. This scale of radiological studies provides a significant opportunity to streamline CXR interpretation and documentation. While foundation models are a promising solution, the lack of publicly available large-scale datasets and benchmarks inhibits their iterative development and real-world evaluation. To overcome these challenges, we constructed a large-scale dataset (CheXinstruct), which we utilized to train a vision-language foundation model (CheXagent). We systematically demonstrated competitive performance across eight distinct task types on our novel evaluation benchmark (CheXbench). Beyond technical validation, we assessed the real-world utility of CheXagent in directly drafting radiology reports. Our clinical assessment with eight radiologists revealed a 36% time saving for residents using CheXagent-drafted reports, while attending radiologists showed no significant time difference editing resident-drafted or CheXagent-drafted reports. The CheXagent-drafted reports improved the writing efficiency of both radiology residents and attending radiologists in 81% and 61% of cases, respectively, without loss of quality. Overall, we demonstrate that CheXagent can effectively perform a variety of CXR interpretation tasks and holds potential to assist radiologists in routine clinical workflows.","external_url":"https://arxiv.org/abs/2401.12208","cited_by_count":24,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2401.12208","created_at":"2026-05-10T05:00:37.127194+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"A vision- language foundation model to enhance efficiency of chest x-ray interpretation","render_title":"A vision- language foundation model to enhance efficiency of chest x-ray interpretation"},"hub":{"state":{"work_id":"91ab4cd1-413f-4e31-be3f-78e309d19006","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":23,"external_cited_by_count":24,"distinct_field_count":5,"first_pith_cited_at":"2023-05-17T17:50:16+00:00","last_pith_cited_at":"2026-07-07T06:23:08+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-06T08:00:34.153502+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":4}],"polarity_counts":[{"context_polarity":"background","n":4}],"runs":{},"summary":{},"graph":{},"authors":[]}}