MedFlowBench evaluates VLM agents on full radiology and pathology studies by requiring both task answers and verifiable evidence like key slices and regions of interest, revealing that answer-only scores overestimate performance.
arXiv preprint arXiv:2505.24173 , year=
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
Med-R2 Bench is a new adversarial benchmark revealing that medical VLMs show sequential performance drops along clinical workflow stages and depend more on correct prompts than on visual grounding.
SaaS-Bench benchmark shows LLM-based agents achieve under 4% end-to-end success on 106 realistic professional tasks spanning 23 deployable SaaS platforms.
citing papers explorer
-
MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows
MedFlowBench evaluates VLM agents on full radiology and pathology studies by requiring both task answers and verifiable evidence like key slices and regions of interest, revealing that answer-only scores overestimate performance.
-
Med-R2: An Adversarial Benchmark for Evidence-Grounded Reasoning in Medical VLMs
Med-R2 Bench is a new adversarial benchmark revealing that medical VLMs show sequential performance drops along clinical workflow stages and depend more on correct prompts than on visual grounding.
-
SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?
SaaS-Bench benchmark shows LLM-based agents achieve under 4% end-to-end success on 106 realistic professional tasks spanning 23 deployable SaaS platforms.