REVIEW 4 major objections 4 minor 61 references
MedVision: Benchmarking Quantitative Medical Image Analysis
T0 review · 4 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read Current vision-language models fail at quantitative medical measurement, but supervised fine-tuning on the new MedVision dataset—30.8 million image-annotation pairs across 22 public datasets—cuts tumor-size error from above 50% to about 30%
desk verdict A genuinely useful new benchmark for quantitative medical VLM evaluation, but the detection results need a y-flip sanity check before the headline claim is safe. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the pairing of a large-scale quantitative annotation pipeline with pixel-spacing-aware prompting: MedVision derives bounding boxes, ellipse-fitted bidirectional tumor diameters, and landmark-based angle/distance ground truths from 22 public segmentation/landmark datasets, then feeds each VLM a text prompt containing the physical pixel size adjusted to that model's own resize/crop pipeline. This lets a text-generating model perform regression on visual quantities. Supervised fine-tuning with LoRA on task-specific splits supplies the missing mapping from visual evidence to numeric outputs.
What would settle it
Hold out a fixed-resize model and deliberately perturb the pixel size in the prompt by ±20% on a sample of T/L and A/D cases; if the model's numeric outputs shift by roughly the same percentage, the spacing prompt is doing the work and the reported accuracy depends on it. Conversely, if outputs are unchanged, the model is ignoring the pixel size and the benchmark's calibration assumption is not load-bearing.
Extended reading notes
Core claim
MedVision establishes that quantitative medical image analysis is a distinct capability that off-the-shelf VLMs lack. On detection, baseline models achieve IoU below 15% for anatomical structures and below 10% for tumors/lesions despite success rates above 90% in producing well-formed coordinates; after LoRA fine-tuning on 1M MedVision samples, Qwen2.5-VL reaches IoU of 71.6% (7B) and 74.6% (32B) on anatomy and 41.2%/44.3% on tumors/lesions. For T/L size, baseline mean relative error ranges 50.1–116.9%; fine-tuning on 5K samples reduces it to ~30% with MAE ~13mm. For angle/distance, baseline angle error is 29.9–66.1° and distance error up to 81.6mm; fine-tuning brings angle error below 4° an
Load-bearing premise
The measured T/L and A/D errors assume that the adjusted pixel size placed in each prompt exactly matches the physical spacing of the image the model actually sees after its internal resize/crop; if that calibration is off, part of the reported error is a prompt artifact rather than genuine measurement failure.
Editorial extensions
If this is right
- If the result holds, clinical measurement tasks such as tumor staging and scoliosis angle assessment become plausible targets for fine-tuned VLMs, not just for dedicated segmentation/regression networks.
- Existing medical VLM benchmarks that only test categorical or qualitative answers underestimate or miss a capability that is both trainable and clinically central.
- SFT with only 5K samples suffices for substantial T/L and A/D improvement, suggesting measurement ability is learnable from modest, well-annotated data.
- Small structures and small angles remain a hard floor for current models, so clinical deployment would need guardrails for small targets.
- Out-of-distribution transfer is partial: the learned quantitative skills generalize across planes and some unseen datasets, but diversity and annotation consistency remain critical.
Reading between the lines
- The paper does not ablate the pixel-spacing prompt, but the design implies a calibration sensitivity: if any deployed model's preprocessing differs from the one assumed at benchmark time, measured performance could drop, so production systems should expose or fix the resize pipeline.
- The paper does not test RECIST-style uni-dimensional measurements, but extending the benchmark to those or to 3D volumetric estimates would test whether the learned skill transfers to other clinical measurement conventions.
- A direct extension of the reported size-accuracy trend: oversampling small targets (below 5% relative size) or using higher-resolution crops should lift the documented bottleneck, since the paper shows a monotone relationship between target size and detection accuracy.
- The paper does not separate perceptual error from output-calibration error, but comparing the same encoder with an object-detection head could reveal whether the measured gap is a vision failure or a language-output failure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MedVision, a large-scale benchmark and dataset for quantitative medical image analysis, aggregating 22 public datasets into 30.8 million image-annotation pairs covering three tasks: detection of anatomical structures and abnormalities, tumor/lesion (T/L) size estimation, and angle/distance (A/D) measurement. It evaluates 15 off-the-shelf VLMs and reports that they perform poorly on all three tasks; after LoRA-based supervised fine-tuning on MedVision, Qwen2.5-VL variants show large improvements in detection IoU, T/L size MRE, and A/D measurement error. The paper includes out-of-distribution evaluations and failure-mode analyses. Code, data, and model checkpoints are released. The central empirical claim is that off-the-shelf VLMs lack quantitative medical-image reasoning and that SFT on MedVision substantially closes the gap.
Significance. If the central claim holds, MedVision fills a real gap: existing medical VLM benchmarks are categorical/qualitative, while quantitative measurement is clinically important. The dataset scale, multi-modality coverage, open release, and OOD evaluation are clear strengths. The paper also separates instruction-following success rate from numeric/localization accuracy, which is a useful evaluation design. However, the headline conclusions rest on two unvalidated assumptions: (i) that baseline detection failures are not partly caused by a coordinate-origin mismatch between prompt and pretraining, and (ii) that the 'adjusted pixel size' supplied in prompts correctly reflects the physical spacing of the image after each VLM's internal preprocessing. Additionally, T/L ground truth is an ellipse fit rather than a clinical measurement, and SFT results are single-run without variance estimates. The resource is significant, but these load-bearing points need to be addressed before the claims are fully established.
major comments (4)
- [Appendix 4.4 / Table 2] The detection prompt requests relative coordinates with the origin at the lower-left corner of the image, but most evaluated VLMs were pretrained to emit bounding-box coordinates with a top-left origin. If a baseline follows its pretrained convention, its predicted y-coordinates are vertically mirrored relative to the benchmark's ground truth, and IoU can be near zero even for a perfectly localized box. The reported SR does not verify coordinate-system compliance. Please add a sanity check: flip predicted y-coordinates (y -> 1-y) for all baselines and recompute detection metrics. If IoU rises substantially, the 'off-the-shelf VLMs fail at detection' claim is partly an artifact of coordinate convention; if not, the concern is resolved. This is load-bearing for the detection half of the central claim.
- [Section 2.2 / Appendix 4.5 / Tables 3-4] T/L size and A/D measurement prompts include an 'adjusted pixel size' derived from each model's image resizing pipeline. The paper does not validate that this value matches the physical spacing of the actual image content after the model's internal cropping, resizing, or padding. If the pixel size is wrong, all predicted lengths and angles are systematically biased, and the large SFT improvements (e.g., MRE dropping from >50% to ~30%, angle error to <4 degrees) could partly reflect correction of a prompt-calibration artifact. Please include an explicit calibration check, such as measuring objects of known physical size across models, or testing whether predictions scale linearly with the stated pixel size.
- [Section 2.1(ii)] The T/L size ground truth is defined as the major and minor axis lengths of an ellipse fitted to the segmentation mask. This is an operationalization, but it is not identical to RECIST's longest diameter or to routine clinical bidirectional measurements. Since the paper motivates T/L size estimation as clinically relevant, please justify that ellipse axes are a clinically meaningful proxy, for example by comparing against manual measurements on a subset, or by explicitly recasting the claim as 'ellipse-based tumor size estimation' rather than 'tumor/lesion size' generally.
- [Tables 2-4] All SFT results are reported from a single run, with no error bars, confidence intervals, or multi-seed variance. The central claim is that SFT dramatically improves performance over baselines; without variance estimates, one cannot exclude the possibility that part of the reported gain is due to training stochasticity, especially since LoRA hyperparameters are fixed ad hoc (r=16, alpha=16). Please report at least 3 seeds for the main SFT models, or, if compute is prohibitive, provide explicit bootstrap confidence intervals or a clear caveat that single-run numbers may not be stable.
minor comments (4)
- [Table 2] The MedDr row appears garbled: '4.174.94.4' is not a valid parsing of the numeric columns. Please fix the table formatting.
- [Table 1] The total row is ambiguous: '22 / 9.2M 23 / 9.6K 5.6 / 2.4K' does not make clear how these counts relate to the advertised 30.8M image-annotation pairs. Please define whether these are train/test totals or per-task counts and ensure the numbers are consistent.
- [Table 1] The dataset name 'autoPEI-III' appears to be a typo for 'autoPET-III'. Please correct.
- [Figures 6 and 8] In Figure 8, labels starting with 'A' and 'L' are not defined in the caption; please provide a legend or table of angle/distance names. Similarly, Figure 6 would benefit from a clear definition of the T/L labels.
Circularity Check
No significant circularity: the benchmark's ground truths come from external public datasets, evaluation is on held-out patient-level test splits, and SFT gains are measured against those independent annotations.
full rationale
The paper's central claims are (1) off-the-shelf VLMs are poor at quantitative medical image analysis and (2) supervised fine-tuning on MedVision substantially improves them. Both claims are evaluated against ground-truth bounding boxes, ellipse-fitted major/minor axes, and landmark-derived angles/distances that are generated from external public datasets (AbdomenAtlas, BraTS24, Ceph-Bio-400, FeTA24, KiTS23, etc.), not from the VLMs' own outputs. The dataset is split at the patient level into 70% training and 30% test, and all SFT results are reported on the held-out test split. There is no parameter fitted to the test target and then renamed as a prediction; the MRE/MAE/IoU metrics are computed by direct comparison of parsed model outputs to independently defined ground truths. The only self-citations are data-source references (OAIZIB-CM from the authors' prior work, and a HuggingFace hosting link for Ceph-Bio-400 under the first author's name); these are provenance details, not load-bearing mathematical or empirical support for the benchmark conclusions. The reviewer's coordinate-origin concern (lower-left vs top-left prompt convention for detection) and the pixel-spacing calibration concern are potential threats to the validity or interpretation of specific measurements, but they are not circularity: they do not make the predictions equal to the ground truth by construction, nor do they fit the target result. Under the stated hard rules, correctness risks that are not circular reductions do not raise the circularity score. Overall, the derivation chain is self-contained against externally sourced annotations, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (3)
- Minimum detection target size =
10 pixels
- Ellipse buffer zone for T/L annotations =
10% shrunk/enlarged bounding box
- LoRA SFT hyperparameters =
r=16, alpha=16, dropout=0.05
assumptions (5)
- domain assumption Pixel spacing in image headers is correct and sufficient for physical measurement.
- domain assumption Public segmentation masks and landmark labels are accurate enough to serve as measurement ground truth.
- ad hoc to paper Fitted-ellipse major/minor axes are a valid operationalization of tumor/lesion size.
- domain assumption The adjusted pixel size written into the prompt matches the actual image content after each model's resize pipeline.
- standard math Standard coordinate geometry and ellipse fitting are valid.
Cite this review
Pith. "Pith review of MedVision: Benchmarking Quantitative Medical Image Analysis." pith.science (2026). https://pith.science/paper/GPVB54JJ
@misc{pith2026251118676,
author = {Pith},
title = {Pith review of: MedVision: Benchmarking Quantitative Medical Image Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/GPVB54JJ}},
note = {Machine review of arXiv:2511.18676}
}
read the original abstract
Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., "Is this normal or abnormal?") or qualitative descriptive tasks. However, clinical decision-making often relies on quantitative assessments, such as measuring the size of a tumor or the angle of a joint, from which physicians draw their own diagnostic conclusions. This quantitative reasoning capability remains underexplored and poorly supported in existing VLMs. In this work, we introduce MedVision, a large-scale dataset and benchmark specifically designed to evaluate and improve VLMs on quantitative medical image analysis. MedVision spans 22 public datasets covering diverse anatomies and modalities, with 30.8 million image-annotation pairs. We focus on three representative quantitative tasks: (1) detection of anatomical structures and abnormalities, (2) tumor/lesion (T/L) size estimation, and (3) angle/distance (A/D) measurement. We show that current off-the-shelf VLMs perform poorly on these tasks. However, supervised and reinforcement fine-tuning on MedVision significantly enhances performance across detection, T/L estimation, and A/D measurement. MedVision provides a foundation for developing VLMs with robust quantitative reasoning capabilities in medical imaging.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Felix Ambellan, Alexander Tack, Moritz Ehlke, and Stefan Zachow. Automated segmentation of knee bone and car- tilage combining statistical shape knowledge and convolu- tional neural networks: Data from the osteoarthritis initia- tive.Medical image analysis, 52:109–118, 2019. 2
2019
-
[2]
A” and “L
Michela Antonelli, Annika Reinke, Spyridon Bakas, Key- van Farahani, Annette Kopp-Schneider, Bennett A Landman, Geert Litjens, Bjoern Menze, Olaf Ronneberger, Ronald M Summers, et al. The medical segmentation decathlon.Nature Figure 8. SFT improves VLM performance in A/D size measure- ment tasks. Labels start with “A” and “L” are angle and distance measur...
2022
-
[3]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities.arXiv preprint arXiv:2308.12966, 1(2):3,
-
[4]
Qwen2.5-vl technical report.arXiv preprint arXiv:2502.13923, 2025
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhao- hai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Jun- yang Lin. Qwen2.5-vl technical repor...
arXiv 2025
-
[5]
Pedro R.A.S. Bassi, Mehmet Can Yavuz, Ibrahim Ethem Hamamci, Sezgin Er, Xiaoxi Chen, Wenxuan Li, Bjoern Menze, Sergio Decherchi, Andrea Cavalli, Kang Wang, Yang Yang, Alan Yuille, and Zongwei Zhou. Radgpt: Con- 8 structing 3d image-text tumor datasets. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 23720–23730, 2025. 1
2025
-
[6]
Vqa-med: Overview of the medical visual question answering task at imageclef 2019
Asma Ben Abacha, Sadid A Hasan, Vivek V Datla, Dina Demner-Fushman, and Henning M ¨uller. Vqa-med: Overview of the medical visual question answering task at imageclef 2019. InProceedings of CLEF (Conference and Labs of the Evaluation Forum) 2019 Working Notes. 9-12 September 2019, 2019. 1
2019
-
[7]
Overview of the vqa-med task at imageclef 2021: Visual question answer- ing and generation in the medical domain
Asma Ben Abacha, Mourad Sarrouti, Dina Demner- Fushman, Sadid A Hasan, and Henning M¨uller. Overview of the vqa-med task at imageclef 2021: Visual question answer- ing and generation in the medical domain. InProceedings of the CLEF 2021 Conference and Labs of the Evaluation Forum-working notes. 21-24 September 2021, 2021. 1
2021
-
[8]
Deep learning techniques for automatic mri cardiac multi-structures segmentation and diagnosis: is the problem solved?IEEE transactions on medical imaging, 37 (11):2514–2525, 2018
Olivier Bernard, Alain Lalande, Clement Zotti, Freder- ick Cervenansky, Xin Yang, Pheng-Ann Heng, Irem Cetin, Karim Lekadir, Oscar Camara, Miguel Angel Gonzalez Ballester, et al. Deep learning techniques for automatic mri cardiac multi-structures segmentation and diagnosis: is the problem solved?IEEE transactions on medical imaging, 37 (11):2514–2525, 2018. 2
2018
Show all 61 references
-
[9]
Making the most of text semantics to improve biomedical vision–language processing
Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C Castro, Anton Schwaighofer, Stephanie Hyland, Maria Wetscherek, Tristan Naumann, Aditya Nori, Javier Alvarez-Valle, et al. Making the most of text semantics to improve biomedical vision–language processing. InEuro- pea...
-
[10]
Segmenting the inferior alveolar canal in cbcts volumes: the toothfairy challenge.IEEE Transactions on Medical Imaging, 2024
Federico Bolelli, Luca Lumetti, Shankeeth Vinayahalingam, Mattia Di Bartolomeo, Arrigo Pellacani, Kevin Marchesini, Niels Van Nistelrooij, Pieter Van Lierop, Tong Xi, Yusheng Liu, et al. Segmenting the inferior alveolar canal in cbcts volumes: the toothfairy challenge.IEEE Tra...
2024
-
[11]
Segmenting max- illofacial structures in cbct volumes
Federico Bolelli, Kevin Marchesini, Niels van Nistelrooij, Luca Lumetti, Vittorio Pipoli, Elisa Ficarra, Shankeeth Vinayahalingam, and Costantino Grana. Segmenting max- illofacial structures in cbct volumes. InProceedings of the Computer Vision and Pattern Recognition Conferen...
2025
-
[12]
Chexpert plus: Augmenting a large chest x-ray dataset with text ra- diology reports, patient demographics and additional image formats.arXiv preprint arXiv:2405.19538, 2024
Pierre Chambon, Jean-Benoit Delbrouck, Thomas Sounack, Shih-Cheng Huang, Zhihong Chen, Maya Varma, Steven QH Truong, Chu The Chuong, and Curtis P Langlotz. Chexpert plus: Augmenting a large chest x-ray dataset with text ra- diology reports, patient demographics and additional ...
2024 arXiv
-
[13]
Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale, 2024
Junying Chen, Chi Gui, Ruyi Ouyang, Anningzhe Gao, Shu- nian Chen, Guiming Hardy Chen, Xidong Wang, Ruifei Zhang, Zhenyang Cai, Ke Ji, Guangjun Yu, Xiang Wan, and Benyou Wang. Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale, 2024. 4, 1
2024
-
[14]
Adolescent idiopathic scoliosis.Nature re- views disease primers, 1(1):1–21, 2015
Jack C Cheng, Ren ´e M Castelein, Winnie C Chu, Aina J Danielsson, Matthew B Dobbs, Theodoros B Grivas, Christina A Gurnett, Keith D Luk, Alain Moreau, Peter O Newton, et al. Adolescent idiopathic scoliosis.Nature re- views disease primers, 1(1):1–21, 2015. 1
2015
-
[15]
Moawad, Yury Velichko, Benedikt Wiestler, Talissa Altes, Patil Basavasagar, Martin Bendszus, Gianluca Brugnara, Jaeyoung Cho, Yaseen Dhemesh, Brandon K
Maria Correia de Verdier, Rachit Saluja, Louis Gagnon, Do- minic LaBella, Ujjwall Baid, Nourel Hoda Tahon, Martha Foltyn-Dumitru, Jikai Zhang, Maram Alafif, Saif Baig, Ken Chang, Gennaro D’Anna, Lisa Deptula, Diviya Gupta, Muhammad Ammar Haider, Ali Hussain, Michael Iv, Mari- ...
2024
-
[16]
SKM-TEA: A dataset for accelerated MRI reconstruction with dense image labels for quantitative clinical evaluation
Arjun D Desai, Andrew M Schmidt, Elka B Rubin, Christo- pher Michael Sandino, Marianne Susan Black, Valentina Mazzoli, Kathryn J Stevens, Robert Boutin, Christopher Re, Garry E Gold, Brian Hargreaves, and Akshay Chaudhari. SKM-TEA: A dataset for accelerated MRI reconstruction ...
2021
-
[17]
The proposed ninth edition tnm classification of lung cancer
Frank C Detterbeck, Gavitt A Woodard, Anna S Bader, Sanja Dacic, Michael J Grant, Henry S Park, and Lynn T Tanoue. The proposed ninth edition tnm classification of lung cancer. Chest, 166(4):882–895, 2024. 1
2024
-
[18]
Dawant, Hexin Dong, Sergio Escalera, Yubo Fan, Lasse Hansen, Mattias P
Reuben Dorent, Aaron Kujawa, Marina Ivory, Spyridon Bakas, Nicola Rieke, Samuel Joutard, Ben Glocker, Jorge Cardoso, Marc Modat, Kayhan Batmanghelich, Arseniy Belkov, Maria Baldeon Calisto, Jae Won Choi, Benoit M. Dawant, Hexin Dong, Sergio Escalera, Yubo Fan, Lasse Hansen, Ma...
2021
-
[19]
New response evaluation criteria in solid tumours: re- vised recist guideline (version 1.1).European journal of can- cer, 45(2):228–247, 2009
Elizabeth A Eisenhauer, Patrick Therasse, Jan Bogaerts, Lawrence H Schwartz, Danielle Sargent, Robert Ford, Janet Dancey, Stephen Arbuck, Steve Gwyther, Margaret Mooney, 9 et al. New response evaluation criteria in solid tumours: re- vised recist guideline (version 1.1).Europe...
2009
-
[20]
Gsco: Towards generalizable ai in medicine via generalist-specialist collaboration, 2024
Sunan He, Yuxiang Nie, Hongmei Wang, Shu Yang, Yihui Wang, Zhiyuan Cai, Zhixuan Chen, Yingxue Xu, Luyang Luo, Huiling Xiang, Xi Lin, Mingxiang Wu, Yifan Peng, George Shih, Ziyang Xu, Xian Wu, Qiong Wang, Ronald Cheong Kin Chan, Varut Vardhanabhuti, Winnie Chiu Wing Chu, Yefeng...
2024
-
[21]
Pathvqa: 30000+ questions for medical visual question answering, 2020
Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. Pathvqa: 30000+ questions for medical visual question answering, 2020. 1
2020
-
[22]
Dense biased networks with deep priori anatomy and hard region adaptation: Semi- supervised learning for fine renal artery segmentation.Med- ical image analysis, 63:101722, 2020
Yuting He, Guanyu Yang, Jian Yang, Yang Chen, Youyong Kong, Jiasong Wu, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger, Pengfei Shao, et al. Dense biased networks with deep priori anatomy and hard region adaptation: Semi- supervised learning for fine renal artery segmentation...
2020
-
[23]
Meta grayscale adaptive network for 3d integrated renal structures segmen- tation.Medical image analysis, 71:102055, 2021
Yuting He, Guanyu Yang, Jian Yang, Rongjun Ge, Youy- ong Kong, Xiaomei Zhu, Shaobo Zhang, Pengfei Shao, Huazhong Shu, Jean-Louis Dillenseger, et al. Meta grayscale adaptive network for 3d integrated renal structures segmen- tation.Medical image analysis, 71:102055, 2021. 2
2021
-
[24]
The kits21 chal- lenge: Automatic segmentation of kidneys, renal tumors, and renal cysts in corticomedullary-phase ct, 2023
Nicholas Heller, Fabian Isensee, Dasha Trofimova, Re- sha Tejpaul, Zhongchen Zhao, Huai Chen, Lisheng Wang, Alex Golts, Daniel Khapun, Daniel Shats, Yoel Shoshan, Flora Gilboa-Solomon, Yasmeen George, Xi Yang, Jian- peng Zhang, Jing Zhang, Yong Xia, Mengran Wu, Zhiyang Liu, Ed...
2023
-
[25]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models, 2021. 4
2021
-
[26]
Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm
Yutao Hu, Tianbin Li, Quanfeng Lu, Wenqi Shao, Junjun He, Yu Qiao, and Ping Luo. Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pages 22170–22183,
-
[27]
Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison
Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Sil- viana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, et al. Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. InProceedings of the AAAI...
2019
-
[28]
Amos: A large-scale abdominal multi-organ benchmark for versatile medical image segmentation.Ad- vances in neural information processing systems, 35:36722– 36732, 2022
Yuanfeng Ji, Haotian Bai, Chongjian Ge, Jie Yang, Ye Zhu, Ruimao Zhang, Zhen Li, Lingyan Zhanng, Wanling Ma, Xi- ang Wan, et al. Amos: A large-scale abdominal multi-organ benchmark for versatile medical image segmentation.Ad- vances in neural information processing systems, 35...
2022
-
[29]
Mimic-cxr, a de- identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019
Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng. Mimic-cxr, a de- identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019. 1
2019
-
[30]
Explaining chest x-ray pathologies in natural language
Maxime Kayser, Cornelius Emde, Oana-Maria Camburu, Guy Parsons, Bartlomiej Papiez, and Thomas Lukasiewicz. Explaining chest x-ray pathologies in natural language. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 701–713. Springer,
-
[31]
Anahita Fathi Kazerooni, Nastaran Khalili, Xinyang Liu, Deep Gandhi, Zhifan Jiang, Syed Muhammed Anwar, Jake Albrecht, Maruf Adewole, Udunna Anazodo, Hannah An- derson, Ujjwal Baid, Timothy Bergquist, Austin J. Borja, Evan Calabrese, Verena Chung, Gian-Marco Conte, Farouk Dako...
2024
-
[32]
Cleveland, Raman- deep Kang, Uma M
Dominic LaBella, Valeriia Abramova, Mehdi Astaraki, An- dre Ferreira, Zhifan Jiang, Mason C. Cleveland, Raman- deep Kang, Uma M. Lal-Trehan Estrada, Cansu Yalcin, Rachika E. Hamadache, Clara Lisazo, Adri `a Casamitjana, Joaquim Salvi, Arnau Oliver, Xavier Llad ´o, Iuliana Toma...
2024
-
[33]
A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018
Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018. 1
2018
-
[34]
Deep learning for segmentation using an open large-scale dataset in 2d echocardiography.IEEE transac- tions on medical imaging, 38(9):2198–2210, 2019
Sarah Leclerc, Erik Smistad, Joao Pedrosa, Andreas Østvik, Frederic Cervenansky, Florian Espinosa, Torvald Espeland, Erik Andreas Rye Berg, Pierre-Marc Jodoin, Thomas Gre- nier, et al. Deep learning for segmentation using an open large-scale dataset in 2d echocardiography.IEEE...
2019
-
[35]
Llava-onevision: Easy visual task transfer, 2024
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Zi- wei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer, 2024. 1, 4
2024
-
[36]
Llava-med: Training a large language- and-vision assistant for biomedicine in one day.arXiv preprint arXiv:2306.00890, 2023
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. Llava-med: Training a large language- and-vision assistant for biomedicine in one day.arXiv preprint arXiv:2306.00890, 2023. 4, 1
2023 arXiv
-
[37]
Wenxuan Li, Chongyu Qu, Xiaoxi Chen, Pedro RAS Bassi, Yijia Shi, Yuxiang Lai, Qian Yu, Huimin Xue, Yixiong Chen, Xiaorui Lin, et al. Abdomenatlas: A large-scale, detailed- annotated, & multi-center dataset for efficient transfer learn- ing and open algorithmic benchmarking.Med...
2024
-
[38]
Healthgpt: A medical large vision- language model for unifying comprehension and generation via heterogeneous knowledge adaptation, 2025
Tianwei Lin, Wenqiao Zhang, Sijing Li, Yuqian Yuan, Binhe Yu, Haoyuan Li, Wanggui He, Hao Jiang, Mengze Li, Xiao- hui Song, Siliang Tang, Jun Xiao, Hui Lin, Yueting Zhuang, and Beng Chin Ooi. Healthgpt: A medical large vision- language model for unifying comprehension and gene...
2025
-
[39]
Pmc-clip: Con- trastive language-image pre-training using biomedical docu- ments, 2023
Weixiong Lin, Ziheng Zhao, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. Pmc-clip: Con- trastive language-image pre-training using biomedical docu- ments, 2023. 1
2023
-
[40]
Fully automatic system for accurate localisation and analysis of cephalometric landmarks in lateral cephalograms.Scien- tific reports, 6(1):33581, 2016
Claudia Lindner, Ching-Wei Wang, Cheng-Ta Huang, Chung-Hsing Li, Sheng-Wei Chang, and Tim F Cootes. Fully automatic system for accurate localisation and analysis of cephalometric landmarks in lateral cephalograms.Scien- tific reports, 6(1):33581, 2016. 2, 3
2016
-
[41]
Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering,
Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering,
-
[42]
Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023. 1
2023
-
[43]
Enhancing patch-based learning for the segmentation of the mandibular canal.IEEE Access, 12: 79014–79024, 2024
Luca Lumetti, Vittorio Pipoli, Federico Bolelli, Elisa Ficarra, and Costantino Grana. Enhancing patch-based learning for the segmentation of the mandibular canal.IEEE Access, 12: 79014–79024, 2024. 2
2024
-
[44]
Abdomenct-1k: Is abdominal organ segmentation a solved problem?IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):6695–6714, 2022
Jun Ma, Yao Zhang, Song Gu, Cheng Zhu, Cheng Ge, Yichi Zhang, Xingle An, Congcong Wang, Qiyuan Wang, Xin Liu, Shucheng Cao, Qi Zhang, Shangqing Liu, Yunpeng Wang, Yuhui Li, Jian He, and Xiaoping Yang. Abdomenct-1k: Is abdominal organ segmentation a solved problem?IEEE Transact...
2022
-
[45]
Jun Ma, Yao Zhang, Song Gu, Cheng Ge, Shihao Mae, Adamo Young, Cheng Zhu, Xin Yang, Kangkang Meng, Ziyan Huang, Fan Zhang, Yuanke Pan, Shoujin Huang, Ji- acheng Wang, Mingze Sun, Rongguo Zhang, Dengqiang Jia, Jae Won Choi, Nat ´alia Alves, Bram de Wilde, Gre- gor Koehler, Haor...
2024
-
[46]
Fetal brain tissue annotation and segmentation challenge re- sults.Medical image analysis, 88:102833, 2023
Kelly Payette, Hongwei Bran Li, Priscille De Dumast, Rox- ane Licandro, Hui Ji, Md Mahfuzur Rahman Siddiquee, Daguang Xu, Andriy Myronenko, Hao Liu, Yuchen Pei, et al. Fetal brain tissue annotation and segmentation challenge re- sults.Medical image analysis, 88:102833, 2023. 2
2023
-
[47]
Steiner, Can Kir- 11 mizibayrak, Rory Pilgrim, Daniel Golden, and Lin Yang
Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri, Atilla Kiraly, Madeleine Traverse, Timo Kohlberger, Shawn Xu, Fayaz Jamil, C ´ıan Hughes, Charles Lau, Justin Chen, Fereshteh Mahvar, Liron Yatziv, Tiffany Chen, Bram Ster- ling, Stefanie Anna Baby, Susanna Maria Baby, Jerem...
2025
-
[48]
Laparoscopic partial nephrectomy with segmental renal artery clamping: technique and clinical outcomes.European urology, 59(5):849–855, 2011
Pengfei Shao, Chao Qin, Changjun Yin, Xiaoxin Meng, Xi- aobing Ju, Jie Li, Qiang Lv, Wei Zhang, and Zhengquan Xu. Laparoscopic partial nephrectomy with segmental renal artery clamping: technique and clinical outcomes.European urology, 59(5):849–855, 2011. 2
2011
-
[49]
Pengfei Shao, Lijun Tang, Pu Li, Yi Xu, Chao Qin, Qiang Cao, Xiaobing Ju, Xiaoxin Meng, Qiang Lv, Jie Li, et al. Precise segmental renal artery clamping under the guidance of dual-source computed tomography angiography during la- paroscopic partial nephrectomy.European urology...
2012
-
[50]
Medicat: A dataset of medical images, captions, and textual references
Sanjay Subramanian, Lucy Lu Wang, Sachin Mehta, Ben Bogin, Madeleine Van Zuylen, Sravanthi Parasa, Sameer Singh, Matt Gardner, and Hannaneh Hajishirzi. Medicat: A dataset of medical images, captions, and textual references. arXiv preprint arXiv:2010.06000, 2020. 1
2010 arXiv
-
[51]
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Ta- tiana Matejovicova, Alexandre Ram ´e, Morgane Rivi `ere, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Ca...
2025
-
[52]
Lingshu: A generalist foun- dation model for unified multimodal medical understanding and reasoning, 2025
LASA Team, Weiwen Xu, Hou Pong Chan, Long Li, Ma- hani Aljunied, Ruifeng Yuan, Jianyu Wang, Chenghao Xiao, Guizhen Chen, Chaoqun Liu, Zhaodonghui Li, Yu Sun, Ju- nao Shen, Chaojun Wang, Jie Tan, Deli Zhao, Tingyang Xu, Hao Zhang, and Yu Rong. Lingshu: A generalist foun- dation...
2025
-
[53]
To- talsegmentator: robust segmentation of 104 anatomic struc- tures in ct images.Radiology: Artificial Intelligence, 5(5): e230024, 2023
Jakob Wasserthal, Hanns-Christian Breit, Manfred T Meyer, Maurice Pradella, Daniel Hinck, Alexander W Sauter, Tobias Heye, Daniel T Boll, Joshy Cyriac, Shan Yang, et al. To- talsegmentator: robust segmentation of 104 anatomic struc- tures in ct images.Radiology: Artificial Int...
2023
-
[54]
Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data
Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Hui Hui, Yanfeng Wang, and Weidi Xie. Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data. Nature Communications, 16(1):7866, 2025. 1
2025
-
[55]
Kaiyuan Yang, Fabio Musio, Yihui Ma, Norman Juch- ler, Johannes C. Paetzold, Rami Al-Maskari, Luciano H¨oher, Hongwei Bran Li, Ibrahim Ethem Hamamci, Anjany Sekuboyina, Suprosanna Shit, Houjing Huang, Chinmay Prabhakar, Ezequiel de la Rosa, Bastian Wittmann, Diana Waldmannstet...
2025
-
[56]
Quantifying knee car- tilage shape and lesion: From image to metrics
Yongcheng Yao and Weitian Chen. Quantifying knee car- tilage shape and lesion: From image to metrics. InAppli- cations of Medical Artificial Intelligence, pages 162–172, Cham, 2025. Springer Nature Switzerland. 2
2025
-
[57]
Cartimorph: A framework for au- tomated knee articular cartilage morphometrics.Medical Im- age Analysis, 91:103035, 2024
Yongcheng Yao, Junru Zhong, Liping Zhang, Sheheryar Khan, and Weitian Chen. Cartimorph: A framework for au- tomated knee articular cartilage morphometrics.Medical Im- age Analysis, 91:103035, 2024. 2
2024
-
[58]
Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.Advances in Neural Information Processing Systems, 37: 94327–94427, 2024
Jin Ye, Guoan Wang, Yanjun Li, Zhongying Deng, Wei Li, Tianbin Li, Haodong Duan, Ziyan Huang, Yanzhou Su, Benyou Wang, et al. Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.Advances in Neural Information Processing Systems, 37: 94327–9...
2024
-
[59]
Drvd-bench: Do vision- language models reason like human doctors in medical im- age diagnosis?, 2025
Tianhong Zhou, Yin Xu, Yingtao Zhu, Chuxi Xiao, Haiyang Bian, Lei Wei, and Xuegong Zhang. Drvd-bench: Do vision- language models reason like human doctors in medical im- age diagnosis?, 2025. 1
2025
-
[60]
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shen- glong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Xuehui Wang, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, Hao Li, Jiahao Wang, Nianchen Deng, Songze L...
2025
-
[61]
Hospedales
Yongshuo Zong, Oisin Mac Aodha, and Timothy M. Hospedales. Self-supervised multimodal learning: A sur- vey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. 1 13 MedVision: Dataset and Benchmark for Quantitative Medical Image Analysis Supplementary Material...
2024
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.