MedFlowBench evaluates VLM agents on full radiology and pathology studies by requiring both task answers and verifiable evidence like key slices and regions of interest, revealing that answer-only scores overestimate performance.
Med-glip: Advancing medical language-image pre-training with large-scale grounded dataset
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 3years
2026 3representative citing papers
BioMedVR proposes a visual reprogramming framework using confusion minimization via LLM attributes, suppression loss, and mixture-of-prompt experts to adapt VLMs to biomedical imaging.
Context alignment in medical VLMs raises AUC from 0.918 to 0.925, cuts hallucinated keywords from 1.14 to 0.25, shortens explanations to 15.3 words, and maintains calibrated uncertainty without raising model confidence.
citing papers explorer
-
MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows
MedFlowBench evaluates VLM agents on full radiology and pathology studies by requiring both task answers and verifiable evidence like key slices and regions of interest, revealing that answer-only scores overestimate performance.
-
BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming
BioMedVR proposes a visual reprogramming framework using confusion minimization via LLM attributes, suppression loss, and mixture-of-prompt experts to adapt VLMs to biomedical imaging.
-
Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models
Context alignment in medical VLMs raises AUC from 0.918 to 0.925, cuts hallucinated keywords from 1.14 to 0.25, shortens explanations to 15.3 words, and maintains calibrated uncertainty without raising model confidence.