{"work":{"id":"3d3f25c0-31e8-4859-bb3d-0d719b47a63d","openalex_id":"https://openalex.org/W4414687708","doi":"10.48550/arxiv.2507.05201","arxiv_id":"2507.05201","raw_key":null,"title":"MedGemma Technical Report","authors":null,"authors_text":"Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri, Atilla P. Kiraly, Madeleine Traverse, Timo Kohlberger, Shawn Xu, F Guimaraes Silvio Jamil, Cían Hughes, Charles Lau","year":2025,"venue":"cs.AI","abstract":"Artificial intelligence (AI) has significant potential in healthcare applications, but its training and deployment faces challenges due to healthcare's diverse data, complex tasks, and the need to preserve privacy. Foundation models that perform well on medical tasks and require less task-specific tuning data are critical to accelerate the development of healthcare AI applications. We introduce MedGemma, a collection of medical vision-language foundation models based on Gemma 3 4B and 27B. MedGemma demonstrates advanced medical understanding and reasoning on images and text, significantly exceeding the performance of similar-sized generative models and approaching the performance of task-specific models, while maintaining the general capabilities of the Gemma 3 base models. For out-of-distribution tasks, MedGemma achieves 2.6-10% improvement on medical multimodal question answering, 15.5-18.1% improvement on chest X-ray finding classification, and 10.8% improvement on agentic evaluations compared to the base models. Fine-tuning MedGemma further improves performance in subdomains, reducing errors in electronic health record information retrieval by 50% and reaching comparable performance to existing specialized state-of-the-art methods for pneumothorax classification and histopathology patch classification. We additionally introduce MedSigLIP, a medically-tuned vision encoder derived from SigLIP. MedSigLIP powers the visual understanding capabilities of MedGemma and as an encoder achieves comparable or better performance than specialized medical image encoders. Taken together, the MedGemma collection provides a strong foundation of medical image and text capabilities, with potential to significantly accelerate medical research and development of downstream applications. The MedGemma collection, including tutorials and model weights, can be found at https://goo.gle/medgemma.","external_url":"https://arxiv.org/abs/2507.05201","cited_by_count":27,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2507.05201","created_at":"2026-05-08T22:14:19.993353+00:00","updated_at":"2026-08-05T02:49:54.815029+00:00","title_quality_ok":false,"display_title":"MedGemma Technical Report","render_title":"MedGemma Technical Report"},"hub":{"state":{"work_id":"3d3f25c0-31e8-4859-bb3d-0d719b47a63d","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":183,"external_cited_by_count":27,"distinct_field_count":15,"first_pith_cited_at":"2025-05-27T19:37:51+00:00","last_pith_cited_at":"2026-07-09T09:03:42+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T03:19:39.043416+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":9},{"context_role":"method","n":4},{"context_role":"baseline","n":2}],"polarity_counts":[{"context_polarity":"background","n":9},{"context_polarity":"use_method","n":4},{"context_polarity":"baseline","n":2}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"MedGemma Technical Report","claims":[{"claim_text":"Artificial intelligence (AI) has significant potential in healthcare applications, but its training and deployment faces challenges due to healthcare's diverse data, complex tasks, and the need to preserve privacy. Foundation models that perform well on medical tasks and require less task-specific tuning data are critical to accelerate the development of healthcare AI applications. We introduce MedGemma, a collection of medical vision-language foundation models based on Gemma 3 4B and 27B. MedGemma demonstrates advanced medical understanding and reasoning on images and text, significantly exce","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"datasets, macro-averaged performance is reported within each dataset's own label set; an 8-label subset shared by all three datasets was used for direct cross-dataset comparison. Reference Labels MERLIN labels were used as published. For the private institutional dataset, study-level labels were derived from radiology reports by two independent LLM-assisted NLP pipelines based on MedGemma [25] and Qwen[26] using the RATE framework released with Pillar-0 [7]: each pipeline extracted findings from","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"86 1.51 40.78 27.32×1.0 - Autoregressive Medical Models LLaVA-Med [25] 6.92 0.65 10.11 11.39 7.36 0.72 10.01 17.04 6.69 0.65 10.07 9.22×1.0 36.33 Lingshu-7B [52] 21.95 1.12 31.34 27.54 23.63 1.25 20.26 34.82 22.24 1.16 32.79 26.94×1.0 53.70 Hulu-Med-7B [21] 22.38 1.15 26.25 23.43 21.66 1.14 22.77 23.56 22.26 1.18 22.89 22.30×1.0 38.78 MedGemma-27B [39] 20.30 0.79 38.83 31.79 23.41 0.91 26.00 37.74 20.80 0.82 38.48 30.24×1.0 16.78 Lingshu-32B [52] 20.40 1.06 30.92 24.32 22.10 1.20 20.33 31.49 20.","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"We use 1000 samples each for DDxPlus and PMC-Patients-v2, and all (201) samples for AgentClinic-MedQA. Baseline models.We compare MedExAgent against a total of seven baseline models. This include four general-purpose LLMs: Qwen3-8B, Qwen3-14B, Qwen3-32B [42], Llama-3-8B-Instruct [12], and three medical LLMs: Aloe-Beta-70B [ 11], MedGemma-27B-text-it [36] and HuatuoGPT-o1- 70B [7]. Note that the base model Meditron-3-8B is not included since it's not an instruction-tuned 6 Table 1: Diagnostic per","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"the deterministic input-to-output mapping that audit-grade reconstruction relies upon, while learned compressors and dense-embedding summarisers do not. Empirical evidence in regulated-domain benchmarks supports the architectural preference for context- preserving over context-rewriting techniques. Biomedical-domain robustness studies report substantial accuracy degradation under adversarial or perturbed context [24, 25]; the FinanceBench benchmark over SEC 10-K/10-Q/8-K filings reports retrieva","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Conceição and Paulo R. C. Lopes, LASIGE, Departamento de Informática, Faculdade de Ciências, Universidade de Lisboa The lasigeBioTM team(37) developed a RAG-based system using the Mistral-7B-Instruct- v0.3 model (49), augmented with external knowledge from Wikipedia and the Mondo ontology (50). They compared this configuration with the performance of MedGemma-27B (51). Their experiments showed that although Mistral-7B-Instruct benefited from the addition of external knowledge, it did not match t","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"net), highlighting the need for con- tinuous data enrichment and the development of interoperable data integration tools, such as MCP-FHIR [21]. In addition, the system pri- marily operates on structured tabular data and does not yet support multimodal inputs, including clinical notes, medical imaging, and physiologi- cal signals. Incorporating the medical foundation models [36] into the system would enable more comprehensive and realistic clinical research work- flows. Lastly, while major compo","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks MedGemma Technical Report because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":3,"context_role":"method"},{"n":2,"context_role":"baseline"}]},"error":null,"updated_at":"2026-05-25T05:05:37.593527+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[]},"error":null,"updated_at":"2026-05-25T05:05:37.587648+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T13:00:59.178901+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":10},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":10},{"title":"P., Li, L., Aljunied, M., Yuan, R., Wang, J., Xiao, C., Chen, G., Liu, C., Li, Z., et al","work_id":"63908fd8-1967-4f5a-ae33-9d555390043d","shared_citers":8},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":8},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":7},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":7},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":7},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":7},{"title":"arXiv preprint arXiv:2305.10415 , year=","work_id":"c8726e94-2bdf-4008-9942-e5228fc1fa64","shared_citers":6},{"title":"arXiv preprint arXiv:2501.18362 (2025)","work_id":"17099117-68ce-46a6-a057-ac254bc140ac","shared_citers":6},{"title":"BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs","work_id":"6fcf8750-00b8-4f3e-9f0b-965a879a5dff","shared_citers":5},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":5},{"title":"Gemma 3 Technical Report","work_id":"f93e08bf-9e96-409b-8ac6-b8385fd17fd7","shared_citers":5},{"title":"Pathvqa: 30000+ questions for medical visual question answering","work_id":"4e35c15f-5a72-4a89-a773-9d4036871506","shared_citers":5},{"title":"arXiv preprint arXiv:2504.00993 (2025)","work_id":"bc0e894e-d353-47d9-8714-8b2f7ced1749","shared_citers":4},{"title":"arXiv preprint arXiv:2510.08668 , year=","work_id":"7a8fdc22-20a2-4741-865e-cb4dfc77234d","shared_citers":4},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":4},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":4},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":4},{"title":"LLaVA-Video: Video Instruction Tuning With Synthetic Data","work_id":"e598f516-d992-449a-ab6d-6c788b3a1d7b","shared_citers":4},{"title":"Medframeqa: A multi-image medical vqa benchmark for clinical reasoning","work_id":"288f899f-01a2-48e4-809e-7b2281450a1b","shared_citers":4},{"title":"URLhttps://www.nature.com/articles/s41597-019-0322-0","work_id":"b5ff10f9-da4b-488f-8068-51b18358da75","shared_citers":4},{"title":"Acosta, Josh Miller, Ouwen Huang, and Pranav Rajpurkar","work_id":"f61477dc-52c2-43a1-8e38-06993fbb2da9","shared_citers":3},{"title":"Ca- pabilities of GPT-4 on medical challenge problems","work_id":"b8eb14ee-3d13-45e1-a013-3ddefdd77f6f","shared_citers":3}],"time_series":[{"n":51,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T13:00:55.393990+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T13:01:03.075235+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"MedGemma Technical Report","claims":[{"claim_text":"Artificial intelligence (AI) has significant potential in healthcare applications, but its training and deployment faces challenges due to healthcare's diverse data, complex tasks, and the need to preserve privacy. Foundation models that perform well on medical tasks and require less task-specific tuning data are critical to accelerate the development of healthcare AI applications. We introduce MedGemma, a collection of medical vision-language foundation models based on Gemma 3 4B and 27B. MedGemma demonstrates advanced medical understanding and reasoning on images and text, significantly exce","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"datasets, macro-averaged performance is reported within each dataset's own label set; an 8-label subset shared by all three datasets was used for direct cross-dataset comparison. Reference Labels MERLIN labels were used as published. For the private institutional dataset, study-level labels were derived from radiology reports by two independent LLM-assisted NLP pipelines based on MedGemma [25] and Qwen[26] using the RATE framework released with Pillar-0 [7]: each pipeline extracted findings from","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"86 1.51 40.78 27.32×1.0 - Autoregressive Medical Models LLaVA-Med [25] 6.92 0.65 10.11 11.39 7.36 0.72 10.01 17.04 6.69 0.65 10.07 9.22×1.0 36.33 Lingshu-7B [52] 21.95 1.12 31.34 27.54 23.63 1.25 20.26 34.82 22.24 1.16 32.79 26.94×1.0 53.70 Hulu-Med-7B [21] 22.38 1.15 26.25 23.43 21.66 1.14 22.77 23.56 22.26 1.18 22.89 22.30×1.0 38.78 MedGemma-27B [39] 20.30 0.79 38.83 31.79 23.41 0.91 26.00 37.74 20.80 0.82 38.48 30.24×1.0 16.78 Lingshu-32B [52] 20.40 1.06 30.92 24.32 22.10 1.20 20.33 31.49 20.","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"We use 1000 samples each for DDxPlus and PMC-Patients-v2, and all (201) samples for AgentClinic-MedQA. Baseline models.We compare MedExAgent against a total of seven baseline models. This include four general-purpose LLMs: Qwen3-8B, Qwen3-14B, Qwen3-32B [42], Llama-3-8B-Instruct [12], and three medical LLMs: Aloe-Beta-70B [ 11], MedGemma-27B-text-it [36] and HuatuoGPT-o1- 70B [7]. Note that the base model Meditron-3-8B is not included since it's not an instruction-tuned 6 Table 1: Diagnostic per","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"the deterministic input-to-output mapping that audit-grade reconstruction relies upon, while learned compressors and dense-embedding summarisers do not. Empirical evidence in regulated-domain benchmarks supports the architectural preference for context- preserving over context-rewriting techniques. Biomedical-domain robustness studies report substantial accuracy degradation under adversarial or perturbed context [24, 25]; the FinanceBench benchmark over SEC 10-K/10-Q/8-K filings reports retrieva","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Conceição and Paulo R. C. Lopes, LASIGE, Departamento de Informática, Faculdade de Ciências, Universidade de Lisboa The lasigeBioTM team(37) developed a RAG-based system using the Mistral-7B-Instruct- v0.3 model (49), augmented with external knowledge from Wikipedia and the Mondo ontology (50). They compared this configuration with the performance of MedGemma-27B (51). Their experiments showed that although Mistral-7B-Instruct benefited from the addition of external knowledge, it did not match t","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"net), highlighting the need for con- tinuous data enrichment and the development of interoperable data integration tools, such as MCP-FHIR [21]. In addition, the system pri- marily operates on structured tabular data and does not yet support multimodal inputs, including clinical notes, medical imaging, and physiologi- cal signals. Incorporating the medical foundation models [36] into the system would enable more comprehensive and realistic clinical research work- flows. Lastly, while major compo","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks MedGemma Technical Report because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":3,"context_role":"method"},{"n":2,"context_role":"baseline"}]},"error":null,"updated_at":"2026-05-25T05:05:38.646569+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"MedGemma Technical Report","claims":[{"claim_text":"Artificial intelligence (AI) has significant potential in healthcare applications, but its training and deployment faces challenges due to healthcare's diverse data, complex tasks, and the need to preserve privacy. Foundation models that perform well on medical tasks and require less task-specific tuning data are critical to accelerate the development of healthcare AI applications. We introduce MedGemma, a collection of medical vision-language foundation models based on Gemma 3 4B and 27B. MedGemma demonstrates advanced medical understanding and reasoning on images and text, significantly exce","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks MedGemma Technical Report because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T13:00:47.523999+00:00"}},"summary":{"title":"MedGemma Technical Report","claims":[{"claim_text":"Artificial intelligence (AI) has significant potential in healthcare applications, but its training and deployment faces challenges due to healthcare's diverse data, complex tasks, and the need to preserve privacy. Foundation models that perform well on medical tasks and require less task-specific tuning data are critical to accelerate the development of healthcare AI applications. We introduce MedGemma, a collection of medical vision-language foundation models based on Gemma 3 4B and 27B. MedGemma demonstrates advanced medical understanding and reasoning on images and text, significantly exce","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks MedGemma Technical Report because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":10},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":10},{"title":"P., Li, L., Aljunied, M., Yuan, R., Wang, J., Xiao, C., Chen, G., Liu, C., Li, Z., et al","work_id":"63908fd8-1967-4f5a-ae33-9d555390043d","shared_citers":8},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":8},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":7},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":7},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":7},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":7},{"title":"arXiv preprint arXiv:2305.10415 , year=","work_id":"c8726e94-2bdf-4008-9942-e5228fc1fa64","shared_citers":6},{"title":"arXiv preprint arXiv:2501.18362 (2025)","work_id":"17099117-68ce-46a6-a057-ac254bc140ac","shared_citers":6},{"title":"BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs","work_id":"6fcf8750-00b8-4f3e-9f0b-965a879a5dff","shared_citers":5},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":5},{"title":"Gemma 3 Technical Report","work_id":"f93e08bf-9e96-409b-8ac6-b8385fd17fd7","shared_citers":5},{"title":"Pathvqa: 30000+ questions for medical visual question answering","work_id":"4e35c15f-5a72-4a89-a773-9d4036871506","shared_citers":5},{"title":"arXiv preprint arXiv:2504.00993 (2025)","work_id":"bc0e894e-d353-47d9-8714-8b2f7ced1749","shared_citers":4},{"title":"arXiv preprint arXiv:2510.08668 , year=","work_id":"7a8fdc22-20a2-4741-865e-cb4dfc77234d","shared_citers":4},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":4},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":4},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":4},{"title":"LLaVA-Video: Video Instruction Tuning With Synthetic Data","work_id":"e598f516-d992-449a-ab6d-6c788b3a1d7b","shared_citers":4},{"title":"Medframeqa: A multi-image medical vqa benchmark for clinical reasoning","work_id":"288f899f-01a2-48e4-809e-7b2281450a1b","shared_citers":4},{"title":"URLhttps://www.nature.com/articles/s41597-019-0322-0","work_id":"b5ff10f9-da4b-488f-8068-51b18358da75","shared_citers":4},{"title":"Acosta, Josh Miller, Ouwen Huang, and Pranav Rajpurkar","work_id":"f61477dc-52c2-43a1-8e38-06993fbb2da9","shared_citers":3},{"title":"Ca- pabilities of GPT-4 on medical challenge problems","work_id":"b8eb14ee-3d13-45e1-a013-3ddefdd77f6f","shared_citers":3}],"time_series":[{"n":51,"year":2026}],"dependency_candidates":[]},"authors":[]}}