REVIEW 3 major objections 7 minor 70 references
UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment
T0 review · 3 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Requiring each modality to predict alone removes the text shortcut and improves multi-modal diagnosis.
desk verdict UniMod is a solid empirical paper with honest diagnostics, but its central story (independent supervision removes shortcuts) is not properly isolated: the missing IFE ablation and a contradictory stray number keep it from being fully convincing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the modality-separated attention mask combined with independent unimodal classification heads. The mask allows image tokens to attend only to image tokens and text tokens only to text tokens, so the mean-pooled image and text embeddings are genuinely unimodal; the final prediction token is the only place cross-modal interaction occurs. Three classification losses—on the image-only, text-only, and multi-modal paths—then force each modality to extract diagnostic features on its own, while an MSE cross-modal alignment loss transfers knowledge between modalities and a supervised contrastive loss structures each modality's representation space by diagnosis. The framework is assembled from established pieces, but the paper's claim is that the independent supervision, not gradient modulation, is what removes the shortcut.
What would settle it
Take the cleaned Harvard-Glaucoma and CheXpert Plus notes and run an independent audit for diagnostic cues, then train the text-only branch on examples that contain no phrase a keyword classifier would flag; if text-only AUC collapses or if a held-out keyword classifier can predict the label from the cleaned text at high accuracy, the shortcut-learning story and the comparison to gradient balancing would need re-quantification.
Extended reading notes
Core claim
The paper claims that supervising image-only, text-only, and multi-modal predictions at the same time removes the lowest-loss route to shortcut learning: with the fused loss alone, a model can satisfy the objective by riding the easier text branch while the image encoder stays nearly non-diagnostic; with independent unimodal losses, neither branch can defer to the other. UniMod operationalizes this with a custom attention mask that keeps image and text token streams separate except at a final prediction token, mean-pooled unimodal embeddings, cross-modality alignment that pulls same-patient image and text representations together, and supervised contrastive alignment that clusters same-diagnosis patients within each modality. The reported result is 0.850 AUC on Harvard-Glaucoma and 0.966 AUC on CheXpert Plus, outperforming OGM-GE and Gradient Blending by 1.6-1.8% and over 5% respectively, and a 5-class multi-label extension improves mean AUC by 0.097 over CGGM without architectural change.
Load-bearing premise
The load-bearing premise is that the text-cleaning step removes all diagnostic label leakage from the clinical notes, so the text-only supervision is learning from genuine clinical reasoning rather than from leaked keywords; if leakage remains, the shortcut is not actually closed.
Editorial extensions
If this is right
- If the central claim is right, gradient-balancing methods such as OGM-GE and Gradient Blending address the symptom rather than the cause; changing the training objective to include unimodal supervision is the effective intervention.
- A model trained with UniMod retains much of its accuracy when one modality is absent at test time, which matters clinically because notes may be incomplete or imaging-only diagnosis may be required.
- The image branch becomes genuinely diagnostic: on CheXpert Plus the image-text AUC gap shrinks from 0.58 under standard multimodal LoRA to 0.06 under UniMod.
- The method transfers to multi-label diagnosis with no architectural change, suggesting the objective, not task-specific design, drives the gain.
Reading between the lines
- Beyond the paper: any vision-language model with shared or separate encoders should benefit from replacing the fused-only objective with per-modality supervision whenever one modality is cheaper to exploit than the other.
- Beyond the paper: the attention mask changes accuracy little, but the paper argues it is a correctness precondition; an external reader could verify by checking whether the text branch's learned features change when the mask is removed.
- Beyond the paper: a direct extension would audit the cleaned notes for residual leaked diagnostic cues; if leakage survives cleaning, the text-only baseline's high AUC could be inflated, which would change how the shortcut is quantified.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes UniMod, a multi-modal medical diagnosis framework that combines fundus images and clinical notes (Harvard-Glaucoma) or chest X-rays and radiology reports (CheXpert Plus). The method adds three classification losses on image-only, text-only, and multi-modal predictions, enforces a modality-separated attention mask, and adds cross-modality MSE alignment plus within-modality supervised contrastive learning, with GradNorm weighting and LoRA fine-tuning of an InternVL2.5-8B backbone. The authors report AUC improvements over OGM-GE and Gradient Blending (0.850 vs 0.835/0.837 on Harvard-Glaucoma; 0.966 vs 0.918/0.919 on CheXpert Plus), missing-modality robustness, and a 5-class multi-label extension that improves mean AUC over CGGM by 0.097. The paper also provides loss-level diagnostics intended to show that standard multi-modal training satisfies the fused objective while leaving the image branch under-optimized.
Significance. If the central causal claim holds, the paper makes a useful and clean contribution: it identifies that gradient-level balancing does not alter the training objective, and that directly supervising each modality's independent prediction is a simple mechanism for discouraging shortcut learning in vision-language medical models. The modality-separated attention design is justified as a correctness precondition rather than an accuracy device, and the missing-modality robustness results in Table 4 and Figures 5-6 are compelling evidence that UniMod produces more balanced modality reliance. The extension to multi-label diagnosis without architectural change is also a positive feature. However, the empirical case for the core mechanism is weakened by the absence of a direct ablation of independent feature extraction, by an internal inconsistency between the reported full-model AUC values, and by incomplete documentation of the label-leakage cleaning step. The paper's significance therefore depends on whether these points can be resolved in revision.
major comments (3)
- [Section 5.4, Table 3] The central mechanism of the paper, Independent Feature Extraction (IFE), is never directly ablated. Table 3 removes only Within and Cross, and both rows keep IFE enabled. No row removes L_img_cls and L_txt_cls while retaining L_cross and L_within, even though the contribution list claims this is the decisive component. The only quantitative evidence for IFE is the statement that 'replacing it under the same alignment losses reaches 0.849 AUC against 0.857 for the full model'; this result has no experimental detail, no variance, and conflicts with Table 1, where UniMod achieves 0.850 AUC on Harvard-Glaucoma. Please add a w/o IFE ablation row with the same alignment losses and report the full model's AUC consistently, together with seed-level variance.
- [Section 5.1, Addressing potential label leakage] The text-cleaning description is not sufficient to establish that clinical notes are free of label leakage. The claim that '0% of samples contain direct label leakage' after cleaning needs the complete cleaning pattern list and an audit procedure; otherwise the text-only baseline and UniMod's text branch could still exploit diagnostic keywords, which would undermine the shortcut-learning narrative and the comparison with gradient-balancing methods. Please provide the full list of removed patterns, representative cleaned and uncleaned examples, and a quantitative check such as text-only AUC on raw versus cleaned notes.
- [Section 5.2, Table 1] The main results in Table 1 are reported on a single split without error bars, even though Table 5 reports means over three seeds. On Harvard-Glaucoma the reported improvement over Gradient Blending is 0.850 versus 0.837 AUC, a 0.013 difference that could plausibly lie within seed-to-seed variation for this setup. Please report mean and standard deviation over at least three seeds for the main results and state whether the differences against OGM-GE and G-Blend are statistically significant.
minor comments (7)
- [Abstract] There is a missing space in 'We proposeUniMod'; similar spacing issues with 'UniMod' appear in the body text.
- [Section 5.1, Table 2] Table references alternate between 'Table' and 'Tbl.'; please use a consistent style.
- [Table 1, Zero-shot row] The recall of 1.000 with AUC 0.472 on Harvard-Glaucoma suggests the zero-shot model is effectively predicting all samples as positive; please clarify the thresholding procedure and discuss the below-chance AUC.
- [Section 5.4, Table 3] The text mentions GradNorm and modality-separation ablations with specific AUC deltas, but Table 3 does not include these rows. Please either add them to the ablation table or clearly state that they are reported only in the text.
- [Section 4.1, 'Why separate the streams'] The discussion of the mask ablation in Section 4.1 would fit more naturally in Section 5.4, and the 'seed-to-seed standard deviation' is mentioned without reporting the actual standard deviation values.
- [Figure 3] The 'greener is better optimized' convention is not accessible to color-blind readers; please add numeric loss values or use a non-color visual cue.
- [Section 5.1, default cleaned text] Please state explicitly whether the text-cleaning is also applied to the text-only baseline and to the case-study reports shown in Figure 4.
Circularity Check
No significant circularity: UniMod's unimodal-supervision mechanism is an intervention tested on external benchmarks; self-citations are peripheral.
full rationale
The paper's central claim is that adding image-only and text-only classification losses to the fused objective removes the lowest-loss shortcut and forces each modality to become diagnostic. That claim is an architectural intervention with the resulting AUCs evaluated on held-out test splits against externally published baselines (OGM-GE, Gradient Blending, CGGM), not a quantity derived from its own fitted constants. The unimodal losses L_img_cls and L_txt_cls are defined on separate representations and are not, by construction, equal to the multi-modal loss or to the reported AUC improvements. The paper's supporting mechanism evidence (CE_mm, CE_img, CE_txt and missing-modality robustness) is observational and could be debated, but that is an experimental-support or correctness concern, not circularity. The self-citations to Kan et al., Zheng et al., and MuteBench ([28], [29], [70]) appear only in related-work context and are not load-bearing: no uniqueness theorem, ansatz, or fitted result is imported from prior work by the same authors. The skeptical concern about the missing IFE-only ablation (Table 3 removes only Within and Cross) is a gap in isolating the independent feature extraction effect, not a reduction of the prediction to the input. Because the method is evaluated against external benchmarks with no fitted parameter renamed as a prediction and no self-citation chain supporting the main mechanism, the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (1)
- SupCon temperature tau =
0.07
assumptions (3)
- domain assumption InternVL2.5-8B's pre-trained weights place image and text tokens in a common embedding space where direct MSE alignment (Eq. 7) is a useful knowledge-transfer signal without a learned projection.
- domain assumption The text cleaning in Section 5.1 removes every direct diagnostic cue, leaving 0% of samples with direct label leakage.
- standard math GradNorm provides a stable, near-optimal adaptive weighting of the classification and alignment losses.
Cite this review
Pith. "Pith review of UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment." pith.science (2026). https://pith.science/paper/CNNIKPCS
@misc{pith2026260810316,
author = {Pith},
title = {Pith review of: UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/CNNIKPCS}},
note = {Machine review of arXiv:2608.10316}
}
read the original abstract
Multi-modal learning combining medical images and clinical text is promising for disease diagnosis. However, standard multi-modal training leads to shortcut learning: models exploit the easier modality (e.g., diagnostic cues in text) while neglecting harder-to-learn features (e.g., subtle visual patterns). We propose UniMod, a framework that mitigates shortcut learning by requiring each modality to predict the diagnosis on its own. It supervises image-only, text-only, and multi-modal classification simultaneously, so each modality must extract diagnostic features. We add cross-modality alignment for knowledge transfer and within-modality supervised contrastive alignment over same-diagnosis patients. On Harvard-Glaucoma, UniMod reaches 0.850 AUC, outperforming OGM-GE and Gradient Blending by 1.6-1.8%; on CheXpert Plus, it reaches 0.966 AUC, surpassing them by over 5%. UniMod also extends to 5-class multi-label diagnosis without architectural change, improving mean AUC by 0.097 over CGGM.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski,...
work page 2022
-
[2]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2. 5-vl technical re...
arXiv 2025
-
[3]
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2018. Multi- modal Machine Learning: A Survey and Taxonomy.IEEE Transactions on Pattern Analysis and Machine Intelligence(2018)
work page 2018
-
[4]
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arber, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models.arXiv preprint arXiv:2108.07258(2021)
arXiv 2021
-
[5]
Pierre Chambon, Jean-Benoit Delbrouck, Thomas Sounack, Shih-Cheng Huang, Zhihong Chen, Maya Varma, Steven QH Truong, Chu The Chuong, and Curtis P. Langlotz. 2024. CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image Formats. arXiv preprint arXiv:2405.19538(2024)
arXiv 2024
-
[6]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. InInterna- tional conference on machine learning. PmLR, 1597–1607
2020
-
[7]
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Universal image-text representation learning. InEuropean Conference on Computer Vision. 104–120
work page 2020
-
[8]
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich. 2018. Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks. InInternational conference on machine learning. PMLR, 794–803
2018
Show all 70 references
-
[9]
Zhihong Chen, Yuhao Du, Jinpeng Hu, Yang Liu, Guanbin Li, Xiang Wan, and Tsung-Hui Chang. 2022. Multi-modal masked autoencoders for medical vision- and-language pre-training. InInternational Conference on Medical Image Comput- ing and Computer-Assisted Intervention. 679–689
2022
-
[10]
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al . 2024. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling.arXiv preprint arXiv:2412.05271(2024)
2024 arXiv
-
[11]
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multi- modality, long context, and next generation agentic...
2025 arXiv
-
[12]
Alex J DeGrave, Joseph D Janizek, and Su-In Lee. 2021. AI for radiographic COVID-19 detection selects shortcuts over signal.Nature Machine Intelligence3, 7 (2021), 610–619
2021
-
[13]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient Finetuning of Quantized LLMs.Advances in Neural Information Processing Systems36 (2023), 10088–10115
2023
-
[14]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, and Sylvain Gelly. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. InInternational...
2020
-
[15]
Tom Fawcett. 2006. An introduction to ROC analysis.Pattern Recognition Letters 27, 8 (2006), 861–874
2006
-
[16]
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020. Shortcut learn- ing in deep neural networks.Nature Machine Intelligence2, 11 (2020), 665–673
2020
-
[17]
Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhao- han Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhao- han Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. 2020. Bootstrap your own...
2020
-
[18]
Zirun Guo, Tao Jin, and Zhou Zhao. 2024. Classifier-Guided Gradient Modulation for Enhanced Multimodal Learning. InAdvances in Neural Information Processing Systems
2024
-
[19]
Haibo He and Edwardo A Garcia. 2009. Learning from imbalanced data.IEEE Transactions on Knowledge and Data Engineering21, 9 (2009), 1263–1284
2009
-
[20]
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Mo- mentum Contrast for Unsupervised Visual Representation Learning. InIEEE/CVF Conference on Computer Vision and Pattern Recognition
2020
-
[21]
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. InInternational Conference on Machine Learning. 2790–2799
2019
-
[22]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. InInternational Conference on Learning Representations
2022
-
[23]
Shih-Cheng Huang, Liyue Shen, Matthew P Lungren, and Serena Yeung. 2021. Gloria: A multimodal global-local representation learning framework for label- efficient medical image recognition. InProceedings of the IEEE/CVF International Conference on Computer Vision. 3942–3951
2021
-
[24]
Fushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang, and Song Guo. 2024. C2KD: Bridging the Modality Gap for Cross-Modal Knowledge Distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion
2024
-
[25]
Mong, Safwan S
Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Silviana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, Jayne Seekins, David A. Mong, Safwan S. Halabi, Jesse K. Sandberg, Ricky Jones, David B. Larson, Curtis P. Langlotz, Bhavik ...
2019
-
[26]
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision- language representation learning with noisy text supervision. InInternational Conference on Machine Learning. 4904–4916
2021
-
[27]
Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih ying Deng, Roger G Mark, and Steven Horng. 2019. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports.Scientific Data6, 1 (2019), 317
2019
-
[28]
Ziwen Kan, Yishuo Chen, Kecheng Li, Andrew Wen, Xiaomeng Wang, Liwei Wang, Jihao Duan, Song Wang, Hongfang Liu, and Tianlong Chen. 2026. TRACE: A Temporal Conditional Estimation for Multimodal Time Series Foundation Models.arXiv preprint arXiv:2606.06285(2026)
2026 arXiv
-
[29]
Ziwen Kan, Wugeng Zheng, Tianlong Chen, and Song Wang. 2026. PAMF: Prior-Aware Multimodal Fusion for Incomplete Time Series Data.arXiv preprint arXiv:2606.06328(2026)
2026 arXiv
-
[30]
Alex Kendall, Yarin Gal, and Roberto Cipolla. 2018. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7482–7491
2018
-
[31]
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised Contrastive Learning. InAdvances in Neural Information Processing Systems
2020
-
[32]
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language trans- former without convolution or region supervision. InInternational Conference on Machine Learning. 5583–5594
2021
-
[33]
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2023. Llava-med: Train- ing a large language-and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems36 (20...
2023
-
[34]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.arXiv preprint arXiv:2301.12597(2023)
2023 arXiv
-
[35]
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language repre- sentation learning with momentum distillation.Advances in Neural Information Processing Systems34 (2021), 9694–9705
2021
-
[36]
Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. InProceedings of the 59th Annual Meeting of the Associa- tion for Computational Linguistics. 4582–4597. MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Gu et al
2021
-
[37]
Victor Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Y Zou. 2022. Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning.Advances in Neural Information Processing Systems(2022)
2022
-
[38]
Weixiong Lin, Ziheng Zhao, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023. Pmc-clip: Contrastive language-image pre-training using biomedical documents.arXiv preprint arXiv:2303.07240(2023)
2023 arXiv
-
[39]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruc- tion tuning. InAdvances in Neural Information Processing Systems
2023
-
[40]
Xiaoxuan Liu, Livia Faes, Aditya U Kale, Siegfried K Wagner, Dun Jack Fu, Alice Bruynseels, Thushika Mahendiran, Gabriella Moraes, Mohith Shamdas, and Christoph Kern. 2019. A comparison of deep learning performance against health- care professionals in detecting diseases from ...
2019
-
[41]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. InProceedings of the IEEE/CVF International Conference on Computer Vision. 10012–10022
2021
-
[42]
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretrain- ing task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in Neural Information Processing Systems32 (2019)
2019
-
[43]
Yan Luo, Min Shi, Yu Tian, Tobias Elze, and Mengyu Wang. 2023. Har- vard glaucoma detection and progression: A multimodal multitask dataset and generalization-reinforced semi-supervised learning. InProceedings of the IEEE/CVF International Conference on Computer Vision. 20471–20482
2023
-
[44]
Natalia Neverova, Christian Wolf, Graham Taylor, and Florian Nebout. 2016. ModDrop: Adaptive multi-modal gesture recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence38, 8 (2016), 1692–1706
2016
-
[45]
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding.arXiv preprint arXiv:1807.03748(2018)
2018 arXiv
-
[46]
OpenAI. 2023. GPT-4V(ision) System Card. https://openai.com/research/gpt-4v- system-card
2023
-
[47]
Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu. 2022. Balanced multimodal learning via on-the-fly gradient modulation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8238–8247
2022
-
[48]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual mod- els from natural language supervision. InInternatio...
2021
-
[49]
Lungren, and Andrew Y
Pranav Rajpurkar, Jeremy Irvin, Kaylie Zhu, Brandon Yang, Hershel Mehta, Tony Duan, Daisy Ding, Aarti Bagul, Curtis Langlotz, Katie Shpanskaya, Matthew P. Lungren, and Andrew Y. Ng. 2017. CheXNet: Radiologist-level pneumonia detec- tion on chest x-rays with deep learning.arXiv...
2017 arXiv
-
[50]
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-net: Convolutional networks for biomedical image segmentation. InInternational Conference on Medical image computing and computer-assisted intervention. Springer, 234–241
2015
-
[51]
Sebastian Ruder. 2017. An overview of multi-task learning in deep neural net- works.arXiv preprint arXiv:1706.05098(2017)
2017 arXiv
-
[52]
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. 2020. Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization. InInternational Conference on Learning Representations
2020
-
[53]
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Nathaneal Scharli, Aakanksha Chowdhery, Philip Mansfield, Blaise Ague...
2023
-
[54]
Yih-Chung Tham, Xiang Li, Tien Y Wong, Harry A Quigley, Tin Aung, and Ching- Yu Cheng. 2014. Global prevalence of glaucoma and projections of glaucoma burden through 2040: a systematic review and meta-analysis.Ophthalmology 121, 11 (2014)
2014
-
[55]
Ekin Tiu, Ellie Talius, Pujan Patel, Curtis P Langlotz, Andrew Y Ng, and Pranav Rajpurkar. 2022. Expert-level detection of pathologies from unannotated chest X-ray images via self-supervised learning.Nature Biomedical Engineering6, 12 (2022), 1399–1406
2022
-
[56]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lam- ple. 2023. Llama: Open and efficient foundation ...
2023 arXiv
-
[57]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. InAdvances in Neural Information Processing Systems. 5998–6008
2017
-
[58]
Tongzhou Wang and Phillip Isola. 2020. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. InInternational Conference on Machine Learning
2020
-
[59]
Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al. 2025. Internvl3. 5: Ad- vancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265(2025)
2025 arXiv
-
[60]
Weiyao Wang, Du Tran, and Matt Feiszli. 2020. What makes training multi- modal classification networks hard?. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12695–12705
2020
-
[61]
Zifeng Wang, Zhenbang Wu, Dinesh Agarwal, and Jimeng Sun. 2022. Medclip: Contrastive learning from unpaired medical images and text.arXiv preprint arXiv:2210.10163(2022)
2022 arXiv
-
[62]
Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023. Medklip: Medical knowledge enhanced language-image pre-training.medRxiv (2023), 2023–01
2023
-
[63]
Peng Xu, Xiatian Zhu, and David A Clifton. 2023. Multimodal learning with transformers: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence45, 10 (2023), 12113–12132
2023
-
[64]
Jinhui Yi, Huan Yan, Haotian Wang, Jian Yuan, and Yong Li. 2023. Deepsta: A spatial-temporal attention network for logistics delivery timely rate prediction in anomaly conditions. InProceedings of the 32nd ACM International Conference on Information and Knowledge Management. 4916–4922
2023
-
[65]
Jinhui Yi, Huan Yan, Haotian Wang, Jian Yuan, and Yong Li. 2024. Learning to Estimate Package Delivery Time in Mixed Imbalanced Delivery and Pickup Logistics Services. InProceedings of the 32nd ACM International Conference on Advances in Geographic Information Systems. 432–443
2024
-
[66]
Kihyun You, Jawook Gu, Jiyeon Ham, Beomhee Park, Jiho Kim, Eun K Hong, Woonhyuk Baek, and Byungseok Roh. 2023. Cxr-clip: Toward large scale chest x-ray language-image pre-training. InInternational Conference on Medical Image Computing and Computer-Assisted Intervention. 101–111
2023
-
[67]
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. 2021. Bar- low twins: Self-supervised learning via redundancy reduction. InInternational Conference on Machine Learning. 12310–12320
2021
-
[68]
Lungren, Tristan Naumann, Sheng Wang, and Hoifung Poon
Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, Cliff Wong, Andrea Tupini, Yu Wang, Matt Mazzola, Swadheen Shukla, Lars Liden, Jianfeng Gao, Angela Crabtree, Brian Piening, Carlo Bifulco, Matthew P....
2023 arXiv
-
[69]
Ying Zhang, Tao Xiang, Timothy M Hospedales, and Huchuan Lu. 2018. Deep mutual learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4320–4328
2018
-
[70]
Wugeng Zheng, Ziwen Kan, Tianlong Chen, Chen Chen, and Song Wang. 2026. MuteBench: Modality Unavailability Tolerance Evaluation for Incomplete Multi- modal Fusion.arXiv preprint arXiv:2605.15235(2026)
2026 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.