REVIEW 4 major objections 5 minor 128 references
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Multimodal language models generalize to unseen medical images by recombining modality, anatomy, and task factors; this compositional mechanism accounts for most multi-task training gains.
desk verdict A substantial new benchmark and a clear empirical pattern, but the causal claim about compositional generalization overreaches the Related-vs-Unrelated design. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The MAT-Triplet decomposition: every medical image is assigned one Modality, one Anatomical area, and one medical Task, and datasets sharing the same triplet are merged into subsets. The argument treats these three labels as independent, recombinable factors, so that training on, say, CT-Lung-Cancer and X-ray-Brain-Cancer should support the unseen combination CT-Brain-Cancer. The central controlled comparison is Related versus Baseline+: equal-size training sets that either do or do not share at least one MAT element with the target, isolating the compositional contribution from raw data volume. A scaling experiment then disrupts CG by dropping one element at a time from the related pool, showing that the shared element is what carries the transfer.
What would settle it
Match two training pools for low-level image statistics and data-source overlap while differing only in whether their manually assigned MAT labels share an element with the target; if target accuracy is equal across the pools, the compositional interpretation is wrong. A second check is to reassign MAT tags randomly across a fixed image pool and split into related and unrelated by the reassigned tags, then see whether target accuracy follows the tags or the image content.
Extended reading notes
Core claim
The paper's central claim is that MLLMs exhibit compositional generalization over the MAT-Triplet: a model that has seen one factor in some contexts can recombine it with other learned factors to handle a target combination it never saw during training. Concretely, training on datasets that share at least one MAT-Triplet element with the target improves target accuracy, while training on equally sized unrelated data leaves accuracy near random. This transfer appears across three MLLM backbones, across classification and detection tasks, and when all three elements of the target come from three different datasets. Deliberately removing a shared element from otherwise related training data causes a large accuracy drop, and training on all related data matches the performance of training on all available data, which the paper reads as evidence that compositional generalization is a principal mechanism behind multi-task generalization.
Load-bearing premise
The load-bearing premise is that Modality, Anatomical area, and Task are the right independent, transferable factors, so that sharing any one of them is what causes the target gain; if the improvements actually come from low-level image similarity, overlapping data sources, or generic multi-task regularization that merely correlates with the tags, the compositional interpretation does not follow.
Editorial extensions
If this is right
- Related medical data can substitute for target data in low-resource settings: adding related combinations alongside a small amount of target data reaches peak accuracy faster than target data alone.
- A newly emerging condition, such as a novel disease, could be handled without any dedicated target training examples if the model is trained on data sharing at least one MAT element with the new imaging task.
- Multi-task training benefits in medical MLLMs come largely from overlapping MAT factors, so constructing multi-task curricula around such overlap should be more effective than mixing arbitrary medical data.
- The effect generalizes across different MLLM architectures and across classification, detection, and segmentation, so the mechanism is not an artifact of one model family.
- Because disrupting CG reduces but does not eliminate generalization, other generalization mechanisms also contribute alongside compositional generalization.
Reading between the lines
- If the MAT decomposition is the true causal structure, the same selection rule should transfer to other descriptive axes; the appendix's population-group and finer-disease results hint that the triplet could be extended, though the paper does not claim that as a main result.
- A practical selection heuristic follows: before collecting target labels for a new medical imaging task, fine-tune on any public data sharing at least one MAT element with the target, and use that as a cheap baseline.
- The compositional account predicts a gradient: transfer should scale with the number of shared elements and with the diversity within each element; this gradient is only partially tested and could be measured directly.
- The same decompositional logic should apply outside medicine to any domain with orthogonal, recombinable descriptive factors, but the paper only provides evidence for medical images.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Med-MAT, a VQA dataset built from 106 public medical image datasets, reorganized into 53 subsets tagged with a MAT-Triplet (Modality, Anatomical area, medical Task). Using LLaVA-v1.5-7B and two additional backbones, the authors compare generalization to a target subset after training on data that shares at least one MAT element (Related) versus data sharing none (Baseline+/All Unrelated). They report that Related training substantially outperforms Unrelated training, that removing shared MAT elements from multi-task training degrades performance, and that related data help in low-data and cross-task (classification/detection) settings. They conclude that compositional generalization is a main driver of multi-task generalization in medical MLLMs.
Significance. The Med-MAT resource is a substantial contribution: it provides a large, publicly released benchmark with explicit factor annotations, and the study spans multiple model families, task types, and includes reproducibility details such as code links and repeated-run statistics in Appendix A.3. If the causal interpretation were established, the work would offer practical guidance for data selection in medical MLLM training and a framework for studying compositional generalization in a high-stakes domain. However, the current evidence is weakened by confounds between the MAT-tag definition and low-level image/format similarity, by a selection procedure that highlights only strong cases in the scaling experiments, and by the absence of error bars in the main tables. These issues do not invalidate the dataset contribution, but they do prevent the paper from supporting its strongest causal claim as written.
major comments (4)
- [Section 3.1, Table 3; Appendix A.2] The central Related-versus-Unrelated contrast is confounded with low-level image distribution and task format. Related data share Modality, Anatomical area, or Task, but these MAT tags also correlate with image acquisition style, windowing, question templates, and label vocabularies (e.g., CT-Lung subsets use similar prompts and answer spaces, while Fundus/Microscopy subsets do not). Baseline+ is sampled from the latter, so the observed gain could be ordinary feature/format transfer rather than recombination of abstract factors. Appendix A.2 confirms leakage for finer disease attributes (COVID vs. pneumonia, gain of only 1.34 points over unrelated data in Table 8), showing that the MAT factor boundaries are not clean. The paper needs a control that matches source distribution or task format while varying only the MAT composition, or an analysis showing MAT overlap predicts gain beyond low-level image similarity.
- [Section 4.1, Figure 4; Take-aways 5 and 6] The causal claim that disrupting CG reduces generalization rests on only two target subsets (Subset 03 and Subset 28), and the 'w/o Modality/Area/Task' manipulations remove entire groups of datasets, changing the number of datasets, task diversity, and possibly the difficulty of the remaining training mixture while holding only total sample count fixed. The observed drops (e.g., 73 to 62/58/48 for Subset 03) are interpreted as evidence that CG drives multi-task gains, but they are equally consistent with a generic multi-task diversity effect. No error bars are reported for these two target subsets, and the Limitations section itself concedes that some generalization remains after disruption. Additional target subsets, statistical repetition, and an analysis that varies dataset composition independently of MAT overlap are needed to support Take-away 6.
- [Section 5.1, Figure 5] The generalization-without-target-data analysis is built on a selection rule that keeps only combinations where Trained already exceeds both Baseline and Baseline+ by at least 10 accuracy points. This explicitly cherry-picks the strongest cases before measuring the scaling curve, so Figure 5 cannot provide unbiased evidence that related data are generally useful in the absence of target data. The authors should report results for all combinations from Table 3, or at least for a pre-registered random sample, alongside the selected subset, so that the reader can assess the strength of Take-away 7.
- [Tables 1, 3, and 4; Appendix A.3] The main quantitative evidence is presented without variance estimates. Table 1, Table 3, and Table 4 report single runs, whereas Appendix A.3 provides mean and standard deviation for only eight selected combinations. Given that Table 3 contains multiple failures and that several reported gains are small (e.g., +1 or +2 points), single-run results are insufficient to establish that the observed related-data advantage is robust. I recommend reporting repeated-run mean and standard deviation for all main tables or, at minimum, for the target subsets used in Figure 4 and the scaling experiments.
minor comments (5)
- [Figure 5 caption and text] The text in Section 5.1 refers to a purple line for Unrelated data, but the figure legend and caption show green and red lines; the colors should be made consistent.
- [Appendix B.2] The prose says 'following the template in Table 8,' but the template is presented as Figure 8; the cross-reference should be corrected.
- [Table 1 caption] The caption says 'different models,' but Section 2.2 describes experiments with a single model (LLaVA-v1.5-7B-Vicuna). The caption should be clarified to avoid implying multi-model results in Table 1.
- [Table 3] The table would be easier to interpret if it included an explicit 'Direction Type' column for each row, because several rows share the same text strings but differ by which element is fixed; the footnote is easy to miss.
- [Section 2.1] The use of ImageWikiQA as a non-medical training set to balance option counts is mentioned but never justified or ablated; a brief explanation of its role would help readers assess potential format confounds.
Circularity Check
The CG claim is partly built into the MAT-tag operationalization, though the underlying related-vs-unrelated contrasts are genuine experiments.
-
self definitional
[Section 3.1, Table 3; Take-away 2]
"To ensure that our conclusions are not influenced by the amount of training data, we randomly sampled an equal number of data from the Unrelated subsets, and this configuration is referred to as Baseline+. ... In the Baseline+ setting, we removed all datasets sharing any MAT-Triplet element with the Target data. Consequently, Baseline+ models perform at near-random levels on the test set, indicating they failed to acquire target-relevant knowledge. This suggests that only datasets related through the MAT-Triplet can help the model learn and generalize to new target tasks."
The 'Related' treatment is defined as sharing at least one manually assigned MAT-Triplet element, and the outcome 'CG Helps' is scored as Trained (Related) beating Baseline+ (Unrelated). The conclusion that MLLMs generalize by recombining MAT elements is therefore read directly off the selection rule that created the comparison. The experiment measures the effect of MAT-tag overlap; it does not measure a compositional recombination mechanism separate from those tags. Any gain from MAT-sharing data is automatically labeled CG, so the existence claim is partly a restatement of the operationalization rather than an independent test of the mechanism.
-
renaming known result
[Section 4.1, Figure 4; Take-away 5]
"w/o Modality/Area/Task are trained on All Related datasets but omit those sharing the same element as the Target Data, to intentionally disrupt CG. ... Take-away 5: Disrupting CG leads to a significant decline in generalization ability. (Q2)."
The manipulation called 'disrupting CG' consists of removing from the training set exactly the data that share a MAT-Triplet element with the target, i.e., removing the property used to define 'Related'. Reporting that this deletion lowers target accuracy is equivalent to saying the defined treatment has an effect; it does not independently establish that the underlying mechanism is compositional recombination. The drop is tied to the deletion operation by construction, and the label 'CG' is applied to the removal after the fact.
full rationale
The paper's empirical backbone is real: equal-sized Related versus Unrelated training sets are compared, the scaling experiment varies the amount and composition of training data, and the disruption experiment manipulates MAT overlap. These results could in principle have gone the other way, so the contrasts are not vacuous. However, the central mechanistic claim that compositional generalization is the driver is partly self-definitional. 'Related' is defined by the MAT-tag overlap whose effect is then measured, and 'disrupting CG' is defined as removing that same overlap, so the interpretive language adds a causal mechanism that the experiments do not independently isolate. The design also does not control for low-level image similarity, source overlap, or shared question/answer formats correlated with the MAT tags; Appendix A.2 even reports leakage between fine-grained attributes because images look similar. No load-bearing self-citation chain was found, and the dataset construction is independent, so the circularity is partial rather than total. The paper establishes a useful empirical correlation between MAT-tag relatedness and transfer gains, but the claim that this proves a compositional recombination mechanism is a renaming of the operational definition.
Assumptions & free parameters
assumptions (4)
- ad hoc to paper Every medical image can be usefully and causally decomposed into Modality, Anatomical area, and Task, and these are the elements the model recombines.
- domain assumption LLaVA-v1.5-7B pretraining contains minimal medical image knowledge, so CG gains are not prior exposure.
- domain assumption Unrelated datasets matched for sample count are an adequate control isolating element overlap.
- domain assumption No target test images leak into related training sets through shared sources or duplicates.
invented entities (1)
-
MAT-Triplet (Modality, Anatomical area, Task) as the compositional decomposition of medical images
Cite this review
Pith. "Pith review of Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging." pith.science (2026). https://pith.science/paper/VMOMLD3Y
@misc{pith2026241220070,
author = {Pith},
title = {Pith review of: Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging},
year = {2026},
howpublished = {\url{https://pith.science/paper/VMOMLD3Y}},
note = {Machine review of arXiv:2412.20070}
}
read the original abstract
Medical imaging provides essential visual insights for diagnosis, and multimodal large language models (MLLMs) are increasingly utilized for its analysis due to their strong generalization capabilities; however, the underlying factors driving this generalization remain unclear. Current research suggests that multi-task training outperforms single-task as different tasks can benefit each other, but they often overlook the internal relationships within these tasks. To analyze this phenomenon, we attempted to employ compositional generalization (CG), which refers to the models' ability to understand novel combinations by recombining learned elements, as a guiding framework. Since medical images can be precisely defined by Modality, Anatomical area, and Task, naturally providing an environment for exploring CG, we assembled 106 medical datasets to create Med-MAT for comprehensive experiments. The experiments confirmed that MLLMs can use CG to understand unseen medical images and identified CG as one of the main drivers of the generalization observed in multi-task training. Additionally, further studies demonstrated that CG effectively supports datasets with limited data and confirmed that MLLMs can achieve CG across classification and detection tasks, underscoring its broader generalization potential. Med-MAT is available at https://github.com/FreedomIntelligence/Med-MAT.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Andrea Acevedo, Anna Merino, Santiago Alf \'e rez, \'A ngel Molina, Laura Bold \'u , and Jos \'e Rodellar. 2020. A dataset of microscopic peripheral blood cell images for development of automatic recognition systems. Data in brief, 30:105474
2020
-
[2]
Walid Al-Dhabyani, Mohammed Gomaa, Hussien Khaled, and Aly Fahmy. 2020. Dataset of breast ultrasound images. Data in brief, 28:104863
2020
-
[3]
Tazuddin Ahmed, Joydip Paul, Tasnim Jahan, S
Shams Nafisa Ali, Md. Tazuddin Ahmed, Joydip Paul, Tasnim Jahan, S. M. Sakeef Sani, Nawshaba Noor, and Taufiq Hasan. 2022. Monkeypox skin lesion detection using deep learning models: A preliminary feasibility study. arXiv preprint arXiv:2207.03342
arXiv 2022
-
[4]
Sharib Ali, Barbara Braden, Dominique Lamarque, Stefano Realdon, Adam Bailey, Renato Cannizzaro, Noha Ghatwary, Jens Rittscher, Christian Daul, and James East. 2020. https://doi.org/10.21227/f8xg-wb80 Endoscopy disease detection and segmentation (edd2020)
-
[5]
MD Anouk Stein, Carol Wu, Chris Carr, George Shih, Jamie Dulkowski, kalpathy, Leon Chen, Luciano Prevedello, MD Marc Kohli, Mark McDonald, Peter, Phil Culliton, Safwan Halabi MD, and Tian Xia. 2018. Rsna pneumonia detection challenge. https://kaggle.com/competitions/rsna-pneumonia-detection-challenge. Kaggle
2018
-
[6]
Will Arevalo. 2020. Chexpert v1.0 small. https://www.kaggle.com/datasets/willarevalo/chexpert-v10-small. Kaggle
2020
-
[7]
A Asraf and Z Islam. 2021. Covid19, pneumonia and normal chest x-ray pa dataset. mendeley data v1 (2021)
2021
-
[8]
Francisco José Fumero Batista, Tinguaro Diaz-Aleman, Jose Sigut, Silvia Alayon, Rafael Arnay, and Denisse Angel-Pereira. 2020. https://doi.org/10.5566/ias.2346 Rim-one dl: A unified retinal image database for assessing glaucoma using deep learning . Image Analysis & Stereology, 39(3):161--167
Show all 128 references
-
[9]
Dev Batra. 2024. Fracture detection using x-ray images. https://www.kaggle.com/datasets/devbatrax/fracture-detection-using-x-ray-images. Kaggle
2024
-
[10]
Veronica Elisa Castillo Ben \' tez, Ingrid Castro Matto, Julio C \'e sar Mello Rom \'a n, Jos \'e Luis V \'a zquez Noguera, Miguel Garc \' a-Torres, Jordan Ayala, Diego P Pinto-Roa, Pedro E Gardel-Sotomayor, Jacques Facon, and Sebastian Alberto Grillo. 2021. Dataset from fundu...
2021
-
[11]
BenO, jljones, Kumar H, Meg Risdal, MRao, Vadim Sherman, Vipul, Wendy Kan, and Yau Ben-Or. 2017. Intel & mobileodt cervical cancer screening. https://kaggle.com/competitions/intel-mobileodt-cervical-cancer-screening. Kaggle
2017
-
[12]
Jorge Bernal, F Javier S \'a nchez, Gloria Fern \'a ndez-Esparrach, Debora Gil, Cristina Rodr \' guez, and Fernando Vilari \ n o. 2015. Wm-dova maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians. Computerized medical imaging and g...
2015
-
[13]
Bukun. 2019. Breast cancer histopathological database (breakhis). https://www.kaggle.com/datasets/ambarish/breakhis. Kaggle
2019
-
[14]
Grillo, Cynthia Villalba, Jacques Facon, Veronica Elisa Castillo Benítez , Ingrid Castro Matto , and Diego Aquino-Brítez
Olivia Cardozo, Verena Ojeda, Rodrigo Parra, Julio César Mello-Román, José Luis Vázquez Noguera , Miguel García-Torres, Federico Divina, Sebastian A. Grillo, Cynthia Villalba, Jacques Facon, Veronica Elisa Castillo Benítez , Ingrid Castro Matto , and Diego Aquino-Brítez. 2023....
2023
-
[15]
Ling-Ping Cen, Jie Ji, Jian-Wei Lin, Si-Tong Ju, Hong-Jie Lin, Tai-Ping Li, Yun Wang, Jian-Feng Yang, Yu-Fen Liu, Shaoying Tan, et al. 2021. Automatic detection of 39 fundus diseases and conditions in retinal photographs using deep neural networks. Nature communications, 12(1):4828
2021
-
[16]
Delong Chen, Jianfeng Liu, Wenliang Dai, and Baoyuan Wang. 2024 a . Visual instruction tuning with polite flamingo. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 17745--17753
2024
-
[17]
Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny. 2023 a . Minigpt-v2: large language model as a unified interface for vision-language multi-task learning. arXiv preprint a...
2023 arXiv
-
[18]
Junying Chen, Ruyi Ouyang, Anningzhe Gao, Shunian Chen, Guiming Hardy Chen, Xidong Wang, Ruifei Zhang, Zhenyang Cai, Ke Ji, Guangjun Yu, et al. 2024 b . Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale. arXiv preprint arXiv:2406.19280
2024 arXiv
-
[19]
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. 2023 b . Shikra: Unleashing multimodal llm's referential dialogue magic. arXiv preprint arXiv:2306.15195
2023 arXiv
-
[20]
Pingjun Chen. 2018. Knee osteoarthritis severity grading dataset. Mendeley Data, 1(10.17632)
2018
-
[21]
Muhammad E. H. Chowdhury, Tawsifur Rahman, Amith Khandakar, Rashid Mazhar, Muhammad Abdul Kadir, Zaid Bin Mahbub, Khandakar Reajul Islam, Muhammad Salman Khan, Atif Iqbal, Nasser Al Emadi, Mamun Bin Ibne Reaz, and Mohammad Tariqul Islam. 2020. https://doi.org/10.1109/ACCESS.20...
2020
-
[22]
Noel Codella, Veronica Rotemberg, Philipp Tschandl, M Emre Celebi, Stephen Dusza, David Gutman, Brian Helba, Aadi Kalloo, Konstantinos Liopyris, Michael Marchetti, et al. 2019. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin im...
2019 arXiv
-
[23]
Noel CF Codella, David Gutman, M Emre Celebi, Brian Helba, Michael A Marchetti, Stephen W Dusza, Aadi Kalloo, Konstantinos Liopyris, Nabin Mishra, Harald Kittler, et al. 2018. Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on bi...
2018
-
[24]
Marc Combalia, Noel CF Codella, Veronica Rotemberg, Brian Helba, Veronica Vilaplana, Ofer Reiter, Cristina Carrera, Alicia Barreiro, Allan C Halpern, Susana Puig, et al. 2019. Bcn20000: Dermoscopic lesions in the wild. arXiv preprint arXiv:1908.02288
2019 arXiv
-
[25]
Will Cukierski. 2018. Histopathologic cancer detection. https://kaggle.com/competitions/histopathologic-cancer-detection. Kaggle
2018
-
[26]
Training Data. 2023. Computed tomography of the brain. https://www.kaggle.com/datasets/trainingdatapro/computed-tomography-ct-of-the-brain. Kaggle
2023
-
[27]
Coen de Vente, Koenraad A. Vermeer, Nicolas Jaccard, He Wang, Hongyi Sun, Firas Khader, Daniel Truhn, Temirgali Aimyshev, Yerkebulan Zhanibekuly, Tien-Dung Le, Adrian Galdran, Miguel Ángel González Ballester, Gustavo Carneiro, Devika R G, Hrishikesh P S, Densen Puthussery, Hon...
2023 arXiv
-
[28]
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. 2023. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378
2023 arXiv
-
[29]
Fernando Feltrin. 2022. Brain tumor mri images 17 classes. https://www.kaggle.com/datasets/fernando2rad/brain-tumor-mri-images-17-classes. Kaggle
2022
-
[30]
Mohammad Fraiwan, Ziad Audat, Luay Fraiwan, and Tarek Manasreh. 2022. Using deep transfer learning to detect scoliosis and spondylolisthesis from x-ray images. Plos one, 17(5):e0267851
2022
-
[31]
Huazhu Fu, Fei Li, José Ignacio Orlando, Hrvoje Bogunović, Xu Sun, Jingan Liao, Yanwu Xu, Shaochong Zhang, and Xiulan Zhang. 2019. https://doi.org/10.21227/55pk-8z03 Palm: Pathologic myopia challenge
2019 doi
-
[32]
Ioannis Giotis, Nynke Molders, Sander Land, Michael Biehl, Marcel F Jonkman, and Nicolai Petkov. 2015. Med-node: A computer-assisted melanoma diagnosis system using non-dermoscopic images. Expert systems with applications, 42(19):6578--6585
2015
-
[33]
Haifan Gong, Guanqi Chen, Ranran Wang, Xiang Xie, Mingzhi Mao, Yizhou Yu, Fei Chen, and Guanbin Li. 2021. Multi-task learning for thyroid nodule segmentation with thyroid region prior. In 2021 IEEE 18th international symposium on biomedical imaging (ISBI), pages 257--261. IEEE
2021
-
[34]
Haifan Gong, Jiaxin Chen, Guanqi Chen, Haofeng Li, Fei Chen, and Guanbin Li. 2022. Thyroid region prior guided attention for ultrasound segmentation of thyroid nodules. Computers in Biology and Medicine, 106389:1--12
2022
-
[35]
Shivanand Gornale and Pooja Patravali. 2020. https://doi.org/10.17632/t9ndx37v5h.1 Digital knee x-ray images
2020 doi
-
[36]
Matthew Groh, Caleb Harris, Luis Soenksen, Felix Lau, Rachel Han, Aerin Kim, Arash Koochek, and Omar Badri. 2021. Evaluating deep neural networks trained on clinical images in dermatology with the fitzpatrick 17k dataset. In Proceedings of the IEEE/CVF Conference on Computer V...
2021
-
[37]
David Gutman, Noel CF Codella, Emre Celebi, Brian Helba, Michael Marchetti, Nabin Mishra, and Allan Halpern. 2016. Skin lesion analysis toward melanoma detection: A challenge at the international symposium on biomedical imaging (isbi) 2016, hosted by the international skin ima...
2016 arXiv
-
[38]
Saba Hesaraki. 2022. Breast ultrasound images dataset (busi). https://www.kaggle.com/datasets/sabahesaraki/breast-ultrasound-images-dataset. Kaggle
2022
-
[39]
Md Nazmul Islam, Mehedi Hasan, Md Kabir Hossain, Md Golam Rabiul Alam, Md Zia Uddin, and Ahmet Soylu. 2022 a . Vision transformer and explainable transfer learning models for auto detection of kidney cyst, stone and tumor from ct-radiography. Scientific Reports, 12(1):1--14
2022
-
[40]
Towhidul Islam, Mohammad Arafat Hussain, Forhad Uddin Hasan Chowdhury, and B M Riazul Islam. 2022 b . https://doi.org/10.1101/2022.08.01.502199 A web-scrapped skin image database of monkeypox, chickenpox, smallpox, cowpox, and measles . bioRxiv 2022.08.01.502199
2022 doi
-
[41]
Stefan Jaeger, Sema Candemir, Sameer Antani, Y \` -Xi \'a ng J W \'a ng, Pu-Xuan Lu, and George Thoma. 2014. Two public chest x-ray datasets for computer-aided screening of pulmonary diseases. Quantitative imaging in medicine and surgery, 4(6):475
2014
-
[42]
Debesh Jha, Pia H Smedsrud, Michael A Riegler, P l Halvorsen, Thomas de Lange, Dag Johansen, and H vard D Johansen. 2020. Kvasir-seg: A segmented polyp dataset. In MultiMedia Modeling: 26th International Conference, MMM 2020, Daejeon, South Korea, January 5--8, 2020, Proceedin...
2020
-
[43]
Kai Jin, Xingru Huang, Jingxing Zhou, Yunxiang Li, Yan Yan, Yibao Sun, Qianni Zhang, Yaqi Wang, and Juan Ye. 2022. Fives: A fundus image dataset for artificial intelligence based vessel segmentation. Scientific data, 9(1):475
2022
-
[44]
JR2NGB. 2019. Cataract dataset. https://www.kaggle.com/datasets/jr2ngb/cataractdataset. Kaggle
2019
-
[45]
Nur Karaca. 2022. Nlm montgomery cxr set. https://www.kaggle.com/datasets/nurkaraca/nlm-montgomerycxrset. Kaggle
2022
-
[46]
Karthik, Maggie, and Sohier Dane. 2019. Aptos 2019 blindness detection. https://kaggle.com/competitions/aptos2019-blindness-detection. Kaggle
2019
-
[47]
Andrey Katanskiy. 2019. Skin cancer isic. https://www.kaggle.com/datasets/nodoubttome/skin-cancer9-classesisic. Kaggle
2019
-
[48]
Jakob Nikolas Kather, Niels Halama, and Alexander Marx. 2018. https://doi.org/10.5281/zenodo.1214456 100,000 histological images of human colorectal cancer and healthy tissue
2018 doi
-
[49]
Daniel Kermany. 2018. Labeled optical coherence tomography (oct) and chest x-ray images for classification. Mendeley data
2018
-
[50]
Felipe Campos Kitamura. 2018. https://doi.org/10.34740/KAGGLE/DSV/152137 Head ct - hemorrhage
2018 doi
-
[51]
Jorge F Lazo, Benoit Rosa, Michele Catellani, Matteo Fontana, Francesco A Mistretta, Gennaro Musi, Ottavio de Cobelli, Michel de Mathelin, and Elena De Momi. 2023. Semi-supervised bladder tissue classification in multi-domain endoscopic images. IEEE Transactions on Biomedical ...
2023
-
[52]
Trang Le, Casper F Winsnes, Ulrika Axelsson, Hao Xu, Jayasankar Mohanakrishnan Kaimal, Diana Mahdessian, Shubin Dai, Ilya S Makarov, Vladislav Ostankovich, Yang Xu, et al. 2022. Analysis of the human protein atlas weakly supervised single-cell classification competition. Natur...
2022
-
[53]
Phuc H Le-Khac, Graham Healy, and Alan F Smeaton. 2020. Contrastive representation learning: A framework and review. Ieee Access, 8:193907--193934
2020
-
[54]
Rebecca Sawyer Lee, Francisco Gimenez, Assaf Hoogi, Kanae Kawai Miyake, Mia Gorovoy, and Daniel L Rubin. 2017. A curated mammography data set for use in computer-aided detection and diagnosis research. Scientific data, 4(1):1--9
2017
-
[55]
Sangjune L Lee, Poonam Yadav, Yin Li, Jason J Meudt, Jessica Strang, Dustin Hebel, Alyx Alfson, Stephanie J Olson, Tera R Kruser, Jennifer B Smilowitz, et al. 2024. Dataset for gastrointestinal tract segmentation on serial mris for abdominal tumor radiotherapy. Data in Brief, ...
2024
-
[56]
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2024. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems, 36
2024
-
[57]
Yuanpeng Li, Liang Zhao, Jianyu Wang, and Joel Hestness. 2019. Compositional generalization for primitive substitutions. arXiv preprint arXiv:1910.02612
2019 arXiv
-
[58]
Yuexiang Li, Nanjun He, and Yawen Huang. 2022. Single domain generalization via spontaneous amplitude spectrum diversification. In MICCAI Workshop on Resource-Efficient Medical Image Analysis, pages 32--41. Springer
2022
-
[59]
Jie Lian, Jingyu Liu, Shu Zhang, Kai Gao, Xiaoqing Liu, Dingwen Zhang, and Yizhou Yu. 2021. A structure-aware relation network for thoracic diseases detection and segmentation. IEEE Transactions on Medical Imaging, 40(8):2042--2052
2021
-
[60]
Xiao Liang. 2021. Adam dataset. https://www.kaggle.com/datasets/xiaoliang2121/adamdataset. Kaggle
2021
-
[61]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning
2023
-
[62]
Jacob A Macdonald, Zhe Zhu, Brandon Konkel, and Mazurowski. 2020. Siim-acr pneumothorax segmentation. https://doi.org/10.5281/zenodo.7774566. Zenodo
2020 doi
-
[63]
K Scott Mader. 2017. Mias mammography. https://www.kaggle.com/datasets/kmader/mias-mammography. Kaggle
2017
-
[64]
Salman Maqbool, Aqsa Riaz, Hasan Sajid, and Osman Hasan. 2020. m2caiseg: Semantic segmentation of laparoscopic images using convolutional neural networks. arXiv preprint arXiv:2008.10134
2020 arXiv
-
[65]
Christian Matek, Sebastian Krappe, Christian M \"u nzenmayer, Torsten Haferlach, and Carsten Marr. 2021. An expert-annotated dataset of bone marrow cytology in hematologic malignancies. The Cancer Imaging Archive
2021
-
[66]
Sarah Matta, Mathieu Lamard, Philippe Zhang, Alexandre Le Guilcher, Laurent Borderie, B \'e atrice Cochener, and Gwenol \'e Quellec. 2024. A systematic review of generalization research in medical image classification. arXiv preprint arXiv:2403.12167
2024 arXiv
-
[67]
Teresa Mendonca, M Celebi, T Mendonca, and J Marques. 2015. Ph2: A public database for the analysis of dermoscopic images. Dermoscopy image analysis
2015
-
[68]
Meta AI . 2024. Llama 3.2: Revolutionizing edge ai and vision with open, customizable models. https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge\\-mobile-devices/
2024
-
[69]
Shentong Mo and Paul Pu Liang. 2024. Multimed: Massively multimodal and multitask medical understanding. arXiv preprint arXiv:2408.12682
2024 arXiv
-
[70]
Paul Mooney. 2017. Blood cell images. https://www.kaggle.com/datasets/paultimothymooney/blood-cells. Kaggle
2017
-
[71]
Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Eduardo Pontes Reis, and Pranav Rajpurkar. 2023. Med-flamingo: a multimodal medical few-shot learner. In Machine Learning for Health (ML4H), pages 353--367. PMLR
2023
-
[72]
Loris Nanni, Michelangelo Paci, Florentino Luciano Caetano dos Santos, Heli Skottman, Kati Juuti-Uusitalo, and Jari Hyttinen. 2016. Texture descriptors ensembles enable image-based classification of maturation of human stem cell-derived retinal pigmented epithelium. PLoS One, ...
2016
-
[73]
Hieu T Nguyen, Ha Q Nguyen, Hieu H Pham, Khanh Lam, Linh T Le, Minh Dao, and Van Vu. 2023. Vindr-mammo: A large-scale benchmark dataset for computer-aided diagnosis in full-field digital mammography. Scientific Data, 10(1):277
2023
-
[74]
Masoud Nickparvar. 2021 a . Brain tumor mri dataset. https://www.kaggle.com/datasets/masoudnickparvar/brain-tumor-mri-dataset. Kaggle
2021
-
[75]
Msoud Nickparvar. 2021 b . https://doi.org/10.34740/KAGGLE/DSV/2645886 Brain tumor mri dataset
2021
-
[77]
Nikita Orlov, Wayne Chen, David Eckley, Tomasz Macura, Lior Shamir, Elaine Jaffe, and Ilya Goldberg. 2010 b . https://doi.org/10.1109/TITB.2010.2050695 Automatic classification of lymphoma images with transform-based global features . IEEE transactions on information technolog...
2010
-
[78]
Silvia Ovreiu, Elena-Anca Paraschiv, and Elena Ovreiu. 2021. Deep learning & digital fundus images: Glaucoma detection using densenet. In 2021 13th international conference on electronics, computers and artificial intelligence (ECAI), pages 1--4. IEEE
2021
-
[79]
Andre GC Pacheco, Gustavo R Lima, Amanda S Salomao, Breno Krohling, Igor P Biral, Gabriel G de Angelo, F \'a bio CR Alves Jr, Jos \'e GM Esgario, Alana C Simora, Pedro BC Castro, et al. 2020. Pad-ufes-20: A skin lesion dataset composed of patient data and clinical images colle...
2020
-
[80]
Sachin Panchal, Ankita Naik, Manesh Kokare, Samiksha Pachade, Rushikesh Naigaonkar, Prerana Phadnis, and Archana Bhange. 2023. Retinal fundus multi-disease image dataset (rfmid) 2.0: a dataset of frequently and rarely identified diseases. Data, 8(2):29
2023
-
[81]
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824
2023 arXiv
-
[82]
H Hieu Pham, T Thanh Tran, and Ha Quy Nguyen. 2022. Vindr-pcxr: An open, large-scale pediatric chest x-ray dataset for interpretation of common thoracic diseases. PhysioNet (version 1.0. 0), 10:2
2022
-
[83]
Hieu Huy Pham, H Nguyen Trung, and Ha Quy Nguyen. 2021. Vindr-spinexr: A large annotated medical image dataset for spinal lesions detection and classification from radiographs. PhysioNet
2021
-
[85]
Konstantin Pogorelov, Kristin Ranheim Randel, Carsten Griwodz, Sigrun Losada Eskeland, Thomas de Lange, Dag Johansen, Concetto Spampinato, Duc-Tien Dang-Nguyen, Mathias Lux, Peter Thelin Schmidt, Michael Riegler, and P l Halvorsen. 2017 b . https://doi.org/10.1145/3083187.3083...
2017
-
[86]
Praveen. 2019. Coronahack chest x-ray dataset. https://www.kaggle.com/datasets/praveengovi/coronahack-chest-xraydataset. Kaggle
2019
-
[87]
Pavle Prentasic, Sven Loncaric, Zoran Vatavuk, Goran Bencic, Marko Subasic, Tomislav Petković, Lana Dujmovic, Maja Malenica Ravlic, Nikolina Budimlija, and Rašeljka Tadić. 2013. https://doi.org/10.1109/ISPA.2013.6703830 Diabetic retinopathy image database(dridb): A new databas...
2013
-
[88]
Xianbiao Qi, Guoying Zhao, Jie Chen, and Matti Pietik \"a inen. 2016. Hep-2 cell classification: The role of gaussian scale space theory as a pre-processing approach. Pattern Recognition Letters, 82:36--43
2016
-
[89]
Raddar. 2019. Chest x-rays (indiana university). https://www.kaggle.com/datasets/raddar/chest-xrays-indiana-university?select=indiana_reports.csv. Kaggle
2019
-
[90]
Tawsifur Rahman, Amith Khandakar, Muhammad Abdul Kadir, Khandaker Rejaul Islam, Khandakar F Islam, Rashid Mazhar, Tahir Hamid, Mohammad Tariqul Islam, Saad Kashem, Zaid Bin Mahbub, et al. 2020. Reliable tuberculosis detection using chest x-ray with deep learning, segmentation ...
2020
-
[91]
MOHD ZAID RASHID. 2024. Oral cancer dataset. https://www.kaggle.com/datasets/zaidpy/oral-cancer-dataset. Kaggle
2024
-
[92]
Sucheng Ren, Xiaoke Huang, Xianhang Li, Junfei Xiao, Jieru Mei, Zeyu Wang, Alan Yuille, and Yuyin Zhou. 2024. Medical vision generalist: Unifying medical imaging tasks in context. arXiv preprint arXiv:2406.05565
2024 arXiv
-
[93]
Manuel Alejandro Rodr \' guez, Hasan AlMarzouqi, and Panos Liatsis. 2022. Multi-label retinal disease classification using transformers. IEEE Journal of Biomedical and Health Informatics
2022
-
[94]
Veronica Rotemberg, Nicholas Kurtansky, Brigid Betz-Stablein, Liam Caffery, Emmanouil Chousakos, Noel Codella, Marc Combalia, Stephen Dusza, Pascale Guitera, David Gutman, et al. 2021. A patient-centric dataset of images and metadata for identifying melanomas using clinical co...
2021
-
[95]
Khaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn, Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, et al. 2024. Capabilities of gemini models in medicine. arXiv preprint arXiv:2404.18416
2024 arXiv
-
[96]
Salman Sajid. 2024. Oral diseases. https://www.kaggle.com/datasets/salmansajid05/oral-diseases/data. Kaggle
2024
-
[97]
F Shaker. 2018. Human sperm head morphology dataset (hushem). Mendeley Data, 3
2018
-
[98]
Julio Silva-Rodr \' guez, Adri \'a n Colomer, Mar \' a A Sales, Rafael Molina, and Valery Naranjo. 2020. Going deeper through the gleason scoring scale: An automatic end-to-end system for histology prostate grading and cribriform pattern detection. Computer Methods and Program...
2020
-
[99]
Eduardo Soares, Plamen Angelov, Sarah Biaso, Michele Higa Froes, and Daniel Kanda Abe. 2020. Sars-cov-2 ct-scan dataset:a large dataset of real patients ct scans for sars-cov-2 identification. Cold Spring Harbor Laboratory Press
2020
-
[100]
Malliga Subramanian, Kogilavani Shanmugavadivel, Obuli Sai Naren, K Premkumar, and K Rankish. 2022. https://doi.org/10.1109/ICCCI54379.2022.9740985 Classification of retinal oct images using deep learning . In 2022 International Conference on Computer Communication and Informa...
2022
-
[101]
Summers and Ronald. 2020. Chestxray nihcc. https://nihcc.app.box.com/v/ChestXray-NIHCC/folder/36938765345. NIH
2020
-
[102]
SunneYi. 2021. https://tianchi.aliyun.com/dataset/93929 Chest CT-Scan images Dataset
2021
-
[103]
Siham Tabik, Anabel G \'o mez-R \' os, Jos \'e Luis Mart \' n-Rodr \' guez, Iv \'a n Sevillano-Garc \' a, Manuel Rey-Area, David Charte, Emilio Guirado, Juan-Luis Su \'a rez, Juli \'a n Luengo, MA Valero-Gonz \'a lez, et al. 2020. Covidgr dataset and covid-sdnet methodology fo...
2020
-
[104]
Yihong Tang, Ao Qu, Zhaokai Wang, Dingyi Zhuang, Zhaofeng Wu, Wei Ma, Shenhao Wang, Yunhan Zheng, Zhan Zhao, and Jinhua Zhao. 2024. Sparkle: Mastering basic spatial capabilities in vision language models elicits generalization to composite spatial reasoning. arXiv preprint arX...
2024
-
[105]
Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann, Mohamed Amin, Pi-Chuan Chang, Andrew Carroll, Charles Lau, Ryutaro Tanno, Ira Ktena, et al. 2024. Towards generalist biomedical ai. NEJM AI, 1(3):AIoa2300138
2024
-
[106]
Peking University. 2019. Odir-2019 dataset. https://odir2019.grand-challenge.org/introduction/. Grand Challenge
2019
-
[107]
Preet Viradiya. 2020. Brain tumor dataset. https://www.kaggle.com/datasets/preetviradiya/brain-tumor-dataset. Kaggle
2020
-
[108]
Haiyang Wang, Hao Tang, Li Jiang, Shaoshuai Shi, Muhammad Ferjad Naeem, Hongsheng Li, Bernt Schiele, and Liwei Wang. 2025. Git: Towards generalist vision transformer through universal language interface. In European Conference on Computer Vision, pages 55--73. Springer
2025
-
[109]
Linda Wang, Zhong Qiu Lin, and Alexander Wong. 2020. https://doi.org/10.1038/s41598-020-76550-z Covid-net: a tailored deep convolutional neural network design for detection of covid-19 cases from chest x-ray images . Scientific Reports, 10(1):19549
2020 doi
-
[111]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024 b . https://arxiv.org/abs/2409.12191 Qwen...
2024 arXiv
-
[112]
Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al. 2024 c . Visionllm: Large language model is also an open-ended decoder for vision-centric tasks. Advances in Neural Information Processing Systems, 36
2024
-
[113]
Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and Ronald M Summers. 2017. Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Proceedings of the IEEE conference...
2017
-
[114]
wjXiaochuangw. 2019. https://tianchi.aliyun.com/dataset/93666 Covid-19-ct scan images
2019
-
[115]
Zhenlin Xu, Marc Niethammer, and Colin A Raffel. 2022. Compositional generalization in unsupervised compositional representation learning: A study on disentanglement and emergent language. Advances in Neural Information Processing Systems, 35:25074--25087
2022
-
[116]
Anna Zawacki, Carol Wu, George Shih, Julia Elliott, Mikhail Fomitchev, Mohannad Hussain, Paras Lakhani, Phil Culliton, and Shunxing Bao. 2019. Siim-acr pneumothorax segmentation. https://kaggle.com/competitions/siim-acr-pneumothorax-segmentation. Kaggle
2019
-
[117]
Yaya Zha. 2021. Rus-chn. https://aistudio.baidu.com/datasetdetail/69582/0. AI Studio
2021
-
[118]
Ao Zhang, Liming Zhao, Chen-Wei Xie, Yun Zheng, Wei Ji, and Tat-Seng Chua. 2023 a . Next-chat: An lmm for chat, detection and segmentation. arXiv preprint arXiv:2311.04498
2023 arXiv
-
[119]
Edward Zhang and Sauman Das. 2022. Glaucoma detection. https://www.kaggle.com/datasets/sshikamaru/glaucoma-detection. Kaggle
2022
-
[120]
Ruipeng Zhang, Qinwei Xu, Chaoqin Huang, Ya Zhang, and Yanfeng Wang. 2022. Semi-supervised domain generalization for medical image analysis. In 2022 IEEE 19th International Symposium on Biomedical Imaging (ISBI), pages 1--5. IEEE
2022
-
[121]
Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Yu Liu, Kai Chen, and Ping Luo. 2023 b . Gpt4roi: Instruction tuning large language model on region-of-interest. arXiv preprint arXiv:2307.03601
2023 arXiv
-
[122]
Tao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan, Shengqiong Wu, Shunping Ji, Change Loy Chen, and Shuicheng Yan. 2024 a . Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding. In NeurIPS
2024
-
[123]
Xiaoman Zhang, Chaoyi Wu, Ziheng Zhao, Weixiong Lin, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023 c . Pmc-vqa: Visual instruction tuning for medical visual question answering. arXiv preprint arXiv:2305.10415
2023 arXiv
-
[124]
Yuhui Zhang, Alyssa Unell, Xiaohan Wang, Dhruba Ghosh, Yuchang Su, Ludwig Schmidt, and Serena Yeung-Levy. 2024 b . Why are visually-grounded language models bad at image classification? arXiv preprint arXiv:2405.18415
2024 arXiv
-
[125]
Jinyu Zhao, Yichen Zhang, Xuehai He, and Pengtao Xie. 2020. Covid-ct-dataset: a ct scan dataset about covid-19. arXiv preprint arXiv:2003.13865
2020 arXiv
-
[126]
Chuang Zhu, Wenkai Chen, Ting Peng, Ying Wang, and Mulan Jin. 2021 a . Hard sample aware noise robust learning for histopathology image classification. IEEE transactions on medical imaging, 41(4):881--894
2021
-
[127]
Chuang Zhu, Wenkai Chen, Ting Peng, Ying Wang, and Mulan Jin. 2021 b . Hard sample aware noise robust learning for histopathology image classification. IEEE transactions on medical imaging, 41(4):881--894
2021
-
[128]
Xile Zhu. 2022. Lc25000. https://www.kaggle.com/datasets/xilezhu/lc25000. Kaggle
2022
-
[129]
Абеуов Нурмұхаммед Бақтыбекұлы. 2021. Augemnted ocular diseases. https://www.kaggle.com/datasets/nurmukhammed7/augemnted-ocular-diseases. Kaggle
2021
-
[130]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[131]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.