{"work":{"id":"dd32b8a1-4ad4-4155-b031-c317b565c6e7","openalex_id":"https://openalex.org/W4400141942","doi":"10.48550/arxiv.2406.19280","arxiv_id":"2406.19280","raw_key":null,"title":"HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale","authors":null,"authors_text":"Chen, J","year":2024,"venue":"cs.CV","abstract":"The rapid development of multimodal large language models (MLLMs), such as GPT-4V, has led to significant advancements. However, these models still face challenges in medical multimodal capabilities due to limitations in the quantity and quality of medical vision-text data, stemming from data privacy concerns and high annotation costs. While pioneering approaches utilize PubMed's large-scale, de-identified medical image-text pairs to address these limitations, they still fall short due to inherent data noise. To tackle this, we refined medical image-text pairs from PubMed and employed MLLMs (GPT-4V) in an 'unblinded' capacity to denoise and reformat the data, resulting in the creation of the PubMedVision dataset with 1.3 million medical VQA samples. Our validation demonstrates that: (1) PubMedVision can significantly enhance the medical multimodal capabilities of current MLLMs, showing significant improvement in benchmarks including the MMMU Health & Medicine track; (2) manual checks by medical experts and empirical results validate the superior data quality of our dataset compared to other data construction methods. Using PubMedVision, we train a 34B medical MLLM HuatuoGPT-Vision, which shows superior performance in medical multimodal scenarios among open-source MLLMs.","external_url":"https://arxiv.org/abs/2406.19280","cited_by_count":8,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2406.19280","created_at":"2026-05-10T01:04:49.685428+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"arXiv preprint arXiv:2406.19280 , year=","render_title":"arXiv preprint arXiv:2406.19280 , year="},"hub":{"state":{"work_id":"dd32b8a1-4ad4-4155-b031-c317b565c6e7","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":26,"external_cited_by_count":8,"distinct_field_count":5,"first_pith_cited_at":"2024-12-25T15:12:34+00:00","last_pith_cited_at":"2026-07-06T16:49:10+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T08:19:28.797249+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":1}],"polarity_counts":[{"context_polarity":"background","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}