REVIEW 2 major objections 2 minor 51 references
Attention Consistent Longitudinal Medical Visual Question Answering Guided by Vision Foundation Models
T0 review · 2 major / 2 minor · reviewed 2026-06-28 · grok-4.3
Pith's one-line read Saliency-conditioned generation after mild affine registration forms a framework for longitudinal medical VQA on chest X-rays.
desk verdict The paper gives a usable training recipe that combines mild affine registration with frozen DINO masks and auxiliary losses for longitudinal chest X-ray VQA, but provides no direct test that the registration preserves real change signals. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The shared saliency mask produced by a frozen DINO-based mask generator combined with a trainable adaptive mask generator and applied after mild affine registration, which conditions the multimodal decoder on consistent attention regions across image pairs.
What would settle it
An ablation experiment in which the registration module is removed or the auxiliary consistency losses are dropped and the model nevertheless matches or exceeds the reported BLEU, ROUGE-L, CIDEr, and METEOR scores on Medical-Diff-VQA would falsify the necessity of those components.
Extended reading notes
Core claim
The central claim is that an attention-guided encoder-decoder that first applies lightweight affine registration to reduce nuisance motion, then generates shared saliency masks from a frozen DINO-based generator plus a trainable adaptive generator, and finally conditions a multimodal transformer decoder on the masked image pairs, produces higher automatic metric scores on the Medical-Diff-VQA benchmark while supplying intrinsic interpretability through the masks; auxiliary losses for mask rebuilding, Gram-style consistency, and uniformity further stabilize learning and clarify change signals.
Load-bearing premise
The lightweight affine registration module with a small regularizer meaningfully reduces nuisance motion without discarding diagnostically relevant change signals.
Editorial extensions
If this is right
- Strong automatic metric performance on the Medical-Diff-VQA benchmark.
- Intrinsic interpretability supplied by the shared saliency mask.
- Demonstration that vision foundation models can be utilized in biomedicine by jointly optimizing supervised and unsupervised objectives.
Reading between the lines
- The same registration-plus-shared-mask pattern could be tested on longitudinal CT or MRI pairs to check whether the framework transfers beyond chest X-rays.
- The auxiliary uniformity and consistency losses might reduce the amount of paired longitudinal annotations needed for training.
- If the generated masks reliably highlight anatomical change, they could serve as an additional visualization aid for clinicians reviewing model outputs.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an attention-guided encoder-decoder for longitudinal medical VQA on chest X-rays. It adds a lightweight affine registration module (with small regularizer) to co-register current and reference images, applies saliency masks from a frozen DINO-based generator plus a trainable adaptive generator, feeds the masked pairs into an image encoder, and uses a multimodal transformer decoder. Auxiliary losses (mask rebuilding, pairwise Gram-style consistency, KoLeo uniformity) inspired by DINO-v3 are included for stabilization. On Medical-Diff-VQA the model is reported to achieve strong BLEU/ROUGE-L/CIDEr/METEOR scores with intrinsic interpretability via the shared masks; the authors conclude that saliency-conditioned generation plus mild pre-alignment constitutes a principled framework for longitudinal medical VQA.
Significance. If the performance claims are supported by proper baselines, ablations, and verification that registration preserves change signals, the work would illustrate a practical way to combine vision foundation models with mild alignment and attention consistency for temporal reasoning tasks in medical imaging. The simultaneous use of supervised and unsupervised objectives is a constructive element that could be relevant beyond this specific benchmark.
major comments (2)
- [Abstract and method description] Abstract (registration module description): the assumption that the lightweight affine registration module with a small regularizer reduces nuisance motion without discarding diagnostically relevant longitudinal change signals is load-bearing for the central claim yet unsupported. No ablation that removes the registration step, no quantitative measure of preserved change (e.g., Dice on annotated lesions pre-/post-registration), and no difference-map visualizations on Medical-Diff-VQA cases with evolving pathology are described. Without these, performance gains cannot be confidently attributed to the proposed saliency-conditioned framework.
- [Abstract] Abstract (evaluation claims): the statement that the model 'delivers strong BLEU, ROUGE-L, CIDEr, and METEOR scores' is presented without any baseline comparisons, statistical significance tests, ablation results, or error analysis. This absence prevents assessment of whether the reported numbers constitute an advance over prior longitudinal VQA methods and therefore undermines the claim that the results support the framework as principled.
minor comments (2)
- The title uses the phrase 'Attention Consistent' but the abstract does not define or operationalize the term; a concise definition or reference to the relevant loss or masking mechanism would improve clarity.
- [Abstract] The auxiliary objectives are listed by name only; a short equation or pseudocode block showing how the mask rebuilding loss, Gram-style consistency loss, and KoLeo uniformity loss are formulated and weighted would aid reproducibility.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on the registration module and the evaluation claims in the abstract. We address each major comment below and indicate planned revisions.
read point-by-point responses
-
Referee: [Abstract and method description] Abstract (registration module description): the assumption that the lightweight affine registration module with a small regularizer reduces nuisance motion without discarding diagnostically relevant longitudinal change signals is load-bearing for the central claim yet unsupported. No ablation that removes the registration step, no quantitative measure of preserved change (e.g., Dice on annotated lesions pre-/post-registration), and no difference-map visualizations on Medical-Diff-VQA cases with evolving pathology are described. Without these, performance gains cannot be confidently attributed to the proposed saliency-conditioned framework.
Authors: We agree that explicit validation of the registration step's effect on change signals is necessary to support the central claim. The small regularizer is intended only to mitigate nuisance motion, but the current manuscript does not include the requested ablations or visualizations. In revision we will add an ablation removing the registration module, quantitative measures of preserved change (e.g., Dice on any available lesion annotations), and difference-map visualizations on Medical-Diff-VQA cases showing evolving pathology. These additions will allow clearer attribution of gains to the saliency-conditioned components. revision: yes
-
Referee: [Abstract] Abstract (evaluation claims): the statement that the model 'delivers strong BLEU, ROUGE-L, CIDEr, and METEOR scores' is presented without any baseline comparisons, statistical significance tests, ablation results, or error analysis. This absence prevents assessment of whether the reported numbers constitute an advance over prior longitudinal VQA methods and therefore undermines the claim that the results support the framework as principled.
Authors: The full manuscript reports comparisons against prior longitudinal VQA methods on Medical-Diff-VQA together with ablation studies. However, the abstract itself does not reference these results or provide context. We will revise the abstract to note the comparative performance, cite the relevant tables/sections for baselines and ablations, and include mention of statistical testing where performed. A concise error analysis summary will also be added to the abstract or a new subsection if space allows. revision: yes
Circularity Check
No significant circularity; derivation remains self-contained
full rationale
The abstract and described method introduce a registration module, DINO-based masks, and auxiliary losses explicitly drawn from external DINO-v3 work. No equations, fitted parameters renamed as predictions, or load-bearing self-citations appear in the provided text. The performance claims on Medical-Diff-VQA are presented as empirical outcomes rather than reductions to inputs by construction. The central framework is therefore not forced by definition or prior author work.
Assumptions & free parameters
assumptions (2)
- domain assumption A frozen DINO model produces useful saliency masks for chest X-ray change detection without task-specific fine-tuning.
- domain assumption The auxiliary losses (mask rebuilding, Gram-style consistency, KoLeo uniformity) improve representation geometry for the VQA task.
Cite this review
Pith. "Pith review of Attention Consistent Longitudinal Medical Visual Question Answering Guided by Vision Foundation Models." pith.science (2026). https://pith.science/paper/R3UXKPWM
@misc{pith2026260606534,
author = {Pith},
title = {Pith review of: Attention Consistent Longitudinal Medical Visual Question Answering Guided by Vision Foundation Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/R3UXKPWM}},
note = {Machine review of arXiv:2606.06534}
}
read the original abstract
Longitudinal medical visual question answering (VQA) requires reasoning about anatomical differences between an image of a current time point and an image of a referred time point. We propose an attention-guided encoder-decoder for this task with chest X-rays. Instead of conventional direct contrast, we propose to include a lightweight affine registration module to reduce nuisance motion by co-registering the current image to the reference image with a small registration regularizer. The registered image pair is fed into the image encoder, followed by a frozen DINO-based mask generator and a trainable adaptive mask generator to produce masks applied to the original image pairs. The masked image pairs are again fed into the image encoder and concatenated with text features as the input to a multimodal transformer-based decoder to generate final answers. To facilitate learning stabilization and clarify the change signal, inspired by DINO-v3, we include additional auxiliary objectives, including a mask rebuilding loss, a pairwise Gram-style consistency loss, and a KoLeo uniformity loss, which enhances the geometry of the representation. On the Medical-Diff-VQA benchmark, the model delivers strong BLEU, ROUGE-L, CIDEr, and METEOR scores while offering intrinsic interpretability through the shared saliency mask. These results support saliency-conditioned generation with mild pre-alignment as a principled framework for longitudinal reasoning in medical VQA. Our training strategy also illustrates the potential of a paradigm in utilizing image foundation models in biomedicine: optimizing both supervised and unsupervised learning objectives simultaneously.
Figures
Reference graph
Works this paper leans on
-
[1]
METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments. InProceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan, 2005. Association for Computational Linguistics. 6
2005
-
[2]
Saliency-driven ex- plainable deep learning in medical imaging: Bridging visual explainability and statistical quantitative analysis
Yusuf Brima and Marcellin Atemkeng. Saliency-driven ex- plainable deep learning in medical imaging: Bridging visual explainability and statistical quantitative analysis. 17(1):18. 1
-
[3]
Swin-unet: Unet-like pure transformer for medical image segmentation,
Hu Cao, Yueyue Wang, Joy Chen, Dongsheng Jiang, Xi- aopeng Zhang, Qi Tian, and Manning Wang. Swin-unet: Unet-like pure transformer for medical image segmentation,
-
[4]
Yuille, and Yuyin Zhou
Jieneng Chen, Yongyi Lu, Qihang Yu, Xiangde Luo, Ehsan Adeli, Yan Wang, Le Lu, Alan L. Yuille, and Yuyin Zhou. Transunet: Transformers make strong encoders for medical image segmentation, 2021. 2
2021
-
[5]
Pretraining vision-language model for difference visual question answering in longitudinal chest x-rays, 2024
Yeongjae Cho, Taehee Kim, Heejun Shin, Sungzoon Cho, and Dongmyung Shin. Pretraining vision-language model for difference visual question answering in longitudinal chest x-rays, 2024. 1, 2, 6
2024
-
[6]
Co-saliency detection with co-attention fully convo- lutional network, 2020
Guangshuai Gao, Wenting Zhao, Qingjie Liu, and Yunhong Wang. Co-saliency detection with co-attention fully convo- lutional network, 2020. 2
2020
-
[7]
Goldberger, Luis A
Ary L. Goldberger, Luis A. N. Amaral, Leon Glass, Jef- frey M. Hausdorff, Plamen Ch. Ivanov, Roger G. Mark, Joseph E. Mietus, George B. Moody, Chung-Kang Peng, and H. Eugene Stanley. Physiobank, physiotoolkit, and phys- ionet.Circulation, 101(23):e215–e220, 2000. 6
2000
-
[8]
Unetr: Transformers for 3d medical image segmentation, 2021
Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath, Dong Yang, Andriy Myronenko, Bennett Landman, Holger Roth, and Daguang Xu. Unetr: Transformers for 3d medical image segmentation, 2021. 2
2021
Show all 51 references
-
[9]
Medical-Diff-VQA: A Large- Scale Medical Dataset for Difference Visual Question An- swering on Chest X-Ray Images
Xinyue Hu, Lin Gu, Qiyuan An, Mengliang Zhang, liangchen liu, Kazuma Kobayashi, Tatsuya Harada, Ronald Summers, and Yingying Zhu. Medical-Diff-VQA: A Large- Scale Medical Dataset for Difference Visual Question An- swering on Chest X-Ray Images. 1, 2, 6
-
[10]
Summers, and Yingying Zhu
Xinyue Hu, Lin Gu, Qiyuan An, Mengliang Zhang, Liangchen Liu, Kazuma Kobayashi, Tatsuya Harada, Ronald M. Summers, and Yingying Zhu. Expert knowledge- aware image difference graph representation learning for difference-aware medical visual question answering. InPro- ceedings o...
2023
-
[11]
Lungren, and Serena Yeung
Shih-Cheng Huang, Liyue Shen, Matthew P. Lungren, and Serena Yeung. Gloria: A multimodal global-local represen- tation learning framework for label-efficient medical image recognition. In2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 3922–3931, 2021. 1
2021
-
[12]
Elsabawy, Alaa Ebraheem Elnakeeb, and Nora Elrashidy
Alaa Hussien, Abdelkareem Elkhateb, Mai Saeed, Nourhan M. Elsabawy, Alaa Ebraheem Elnakeeb, and Nora Elrashidy. Explainable self-supervised learning for medical image diagnosis based on DINO V2 model and semantic search. 15(1):32174. 3
-
[13]
Jaeger, Simon A
Fabian Isensee, Paul F. Jaeger, Simon A. A. Kohl, Jens Petersen, and Klaus H. Maier-Hein. nnU-Net: A self- configuring method for deep learning-based biomedical im- age segmentation. 18(2):203–211. 2
-
[14]
Spatial transformer networks, 2016
Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu. Spatial transformer networks, 2016. 2
2016
-
[15]
One map does not fit all: Evaluating saliency map explanation on multi-modal medical images, 2021
Weina Jin, Xiaoxiao Li, and Ghassan Hamarneh. One map does not fit all: Evaluating saliency map explanation on multi-modal medical images, 2021. 1
2021
-
[16]
Alistair E. W. Johnson, Tom J. Pollard, Seth J. Berkowitz, Nathaniel R. Greenbaum, Matthew P. Lungren, Chih-ying Deng, Roger G. Mark, and Steven Horng. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. 6(1):317. 6
-
[17]
Alistair E. W. Johnson, Tom J. Pollard, Nathaniel R. Green- baum, Matthew P. Lungren, Chih ying Deng, Yifan Peng, Zhiyong Lu, Roger G. Mark, Seth J. Berkowitz, and Steven Horng. Mimic-cxr-jpg, a large publicly available database of labeled chest radiographs, 2019. 6
2019
-
[18]
Schroeder, and Tolga Tasdizen
Ricardo Bigolin Lanfredi, Ambuj Arora, Trafton Drew, Joyce D. Schroeder, and Tolga Tasdizen. Comparing radiol- ogists’ gaze and saliency maps generated by interpretability methods for chest x-rays, 2023. 1
2023
-
[19]
Llava-med: Training a large language- and-vision assistant for biomedicine in one day, 2023
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. Llava-med: Training a large language- and-vision assistant for biomedicine in one day, 2023. 1
2023
-
[20]
Meddinov3: How to adapt vision foundation models for medical image segmentation?, 2025
Yuheng Li, Yizhou Wu, Yuxiang Lai, Mingzhe Hu, and Xi- aofeng Yang. Meddinov3: How to adapt vision foundation models for medical image segmentation?, 2025. 3
2025
-
[21]
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. ROUGE: A package for automatic evaluation of summaries. InText Summarization Branches Out, pages 74–81, Barcelona, Spain, 2004. Association for Computa- tional Linguistics. 6
2004
-
[22]
Medical visual question answering: A survey
Zhihong Lin, Donghao Zhang, Qingyi Tao, Danli Shi, Gho- lamreza Haffari, Qi Wu, Mingguang He, and Zongyuan Ge. Medical visual question answering: A survey. 143:102611. 1
-
[23]
Swin trans- former: Hierarchical vision transformer using shifted win- dows.CoRR, abs/2103.14030, 2021
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin trans- former: Hierarchical vision transformer using shifted win- dows.CoRR, abs/2103.14030, 2021. 4
2021 arXiv
-
[24]
Thakoor, Padraig Corcoran, Ying Chen, and Hantao Liu
Jianxun Lou, Huasheng Wang, Xinbo Wu, John Cho Hui Ng, Richard White, Kaveri A. Thakoor, Padraig Corcoran, Ying Chen, and Hantao Liu. Chest x-ray visual saliency modeling: Eye-tracking dataset and saliency prediction model.IEEE Transactions on Neural Networks and Learning Syst...
2025
-
[25]
Spot the Difference: Difference Visual Ques- tion Answering with Residual Alignment
Zilin Lu, Yutong Xie, Qingjie Zeng, Mengkang Lu, Qi Wu, and Yong Xia. Spot the Difference: Difference Visual Ques- tion Answering with Residual Alignment . Inproceedings of Medical Image Computing and Computer Assisted Interven- tion – MICCAI 2024. Springer Nature Switzerland,...
2024
-
[26]
Segment anything in medical images
Jun Ma, Yuting He, Feifei Li, Lin Han, Chenyu You, and Bo Wang. Segment anything in medical images. 15(1):654. 2
-
[27]
Unveiling differences: A vision encoder-decoder model for difference medical visual question answering
Luis-Jesus Marhuenda, Miquel Obrador-Reina, Mohamed Aas-Alas, Alberto Albiol, and Roberto Paredes. Unveiling differences: A vision encoder-decoder model for difference medical visual question answering. InMedical Imaging with Deep Learning, 2025. 2, 4, 6
2025
-
[28]
Lon- gitudinal change detection on chest x-rays using geometric correlation maps
Dong Yul Oh, Jihang Kim, and Kyong Joon Lee. Lon- gitudinal change detection on chest x-rays using geometric correlation maps. page 748–756, Berlin, Heidelberg, 2019. Springer-Verlag. 1
2019
-
[29]
Dinov2: Learning robust visual features with- out supervision, 2024
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mah- moud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michae...
2024
-
[30]
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311– 318, Philadelphia, Pennsylvania, USA, 2002. Associ...
2002
-
[31]
Castro, Anton Schwaighofer, Matthew P
Fernando P ´erez-Garc´ıa, Harshita Sharma, Sam Bond- Taylor, Kenza Bouzid, Valentina Salvatelli, Maximilian Ilse, Shruthi Bannur, Daniel C. Castro, Anton Schwaighofer, Matthew P. Lungren, Maria Teodora Wetscherek, Noel Codella, Stephanie L. Hyland, Javier Alvarez-Valle, and Oz...
2025
-
[32]
Language models are unsuper- vised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsuper- vised multitask learners. 2019. 4
2019
-
[33]
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junt- ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao- Yuan Wu, Ross Girshick, Piotr Doll´ar, and Christoph Feic...
2024 arXiv
-
[34]
Imagenet-21k pretraining for the masses
Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zel- nik. Imagenet-21k pretraining for the masses. InProceed- ings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2021. 4
2021
-
[35]
Adriel Saporta, Xiaotong Gui, Ashwin Agrawal, Anuj Pa- reek, Steven Q. H. Truong, Chanh D. T. Nguyen, Van-Doan Ngo, Jayne Seekins, Francis G. Blankenberg, Andrew Y . Ng, Matthew P. Lungren, and Pranav Rajpurkar. Benchmarking saliency methods for chest X-ray interpretation. 4(10):867–
-
[36]
Peeken, Daniel Rueckert, and Benedikt Wiestler
Daniel Scholz, Ayhan Can Erdur, Viktoria Ehm, Anke Meyer-Baese, Jan C. Peeken, Daniel Rueckert, and Benedikt Wiestler. MM-DINOv2: Adapting Foundation Models for Multi-Modal Medical Image Analysis . Inproceedings of Medical Image Computing and Computer Assisted Interven- tion –...
2025
-
[37]
Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra. Grad-cam: Visual explanations from deep networks via gradient-based localization.International Journal of Com- puter Vision, 128(2):336–359, 2019. 2
2019
-
[38]
Oriane Sim ´eoni, Huy V . V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Micha ¨el Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timoth´ee Darcet, Th´eo Moutakanni, Leonel Sen...
2025
-
[39]
General pur- pose image encoder dinov2 for medical image registration,
Xinrui Song, Xuanang Xu, and Pingkun Yan. General pur- pose image encoder dinov2 for medical image registration,
-
[40]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2023. 4
2023
-
[41]
Lawrence Zitnick, and Devi Parikh
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evalua- tion, 2015. 6
2015
-
[42]
Co-attention aligned mutual cross-attention for cloth- changing person re-identification
Qizao Wang, Xuelin Qian, Yanwei Fu, and Xiangyang Xue. Co-attention aligned mutual cross-attention for cloth- changing person re-identification. page 351–368, Berlin, Heidelberg, 2022. Springer-Verlag. 2
2022
-
[43]
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chau- mond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Ma...
2020
-
[44]
Sabel, and Tobias Lasser
Alessandro Wollek, Robert Graf, Sa ˇsa ˇCeˇcatka, Nicola Fink, Theresa Willem, Bastian O. Sabel, and Tobias Lasser. Attention-based saliency maps improve interpretability of pneumothorax classification.Radiology: Artificial Intelli- gence, 5(2):e220187, 2023. PMID: 37035429. 2
2023
-
[45]
Segdino: An efficient design for medical and natural image segmentation with dino-v3, 2025
Sicheng Yang, Hongqiu Wang, Zhaohu Xing, Sixiang Chen, and Lei Zhu. Segdino: An efficient design for medical and natural image segmentation with dino-v3, 2025. 3
2025
-
[46]
Exploring visual relationship for image captioning, 2018
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. Exploring visual relationship for image captioning, 2018. 2, 6
2018
-
[47]
Describing and localizing multiple changes with transformers, 2021
Kodai Nakashima Ryota Suzuki Kenji Iwata Hi- rokatsu Kataoka Yue Qiu, Shintaro Yamamoto and Yutaka Satoh. Describing and localizing multiple changes with transformers, 2021. 2, 6
2021
-
[48]
Mazomenos
Ka-Wai Yung, Jayaram Sivaraj, Danail Stoyanov, Stavros Loukogeorgakis, and Evangelos B. Mazomenos. Region- Specific Retrieval Augmentation for Longitudinal Visual Question Answering: A Mix-and-Match Paradigm . In proceedings of Medical Image Computing and Computer Assisted Int...
2024
-
[49]
Lungren, Akshay Chaudhari, Ser- ena Yeung-Levy, Curtis P
Juan Manuel Zambrano Chaves, Shih-Cheng Huang, Yanbo Xu, Hanwen Xu, Naoto Usuyama, Sheng Zhang, Fei Wang, Yujia Xie, Mahmoud Khademi, Ziyi Yang, Hany Awadalla, Julia Gong, Houdong Hu, Jianwei Yang, Chunyuan Li, Jian- feng Gao, Yu Gu, Cliff Wong, Mu Wei, Tristan Naumann, Muhao ...
-
[50]
PMC-VQA: Vi- sual instruction tuning for medical visual question answer- ing, 2024
Xiaoman Zhang, Chaoyi Wu, Ziheng Zhao, Weixiong Lin, Ya Zhang, Yanfeng Wang, and Weidi Xie. PMC-VQA: Vi- sual instruction tuning for medical visual question answer- ing, 2024. 1
2024
-
[51]
Medical sam 2: Segment medical images as video via segment anything model 2, 2024
Jiayuan Zhu, Abdullah Hamdi, Yunli Qi, Yueming Jin, and Junde Wu. Medical sam 2: Segment medical images as video via segment anything model 2, 2024. 2
2024
Reviewed June 28, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.