REVIEW 4 major objections 7 minor 56 references
MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion
T0 review · 4 major / 7 minor · reviewed 2026-07-31 · grok-4.5
Pith's one-line read Medical image fusion guided by diagnostic intent texts, not uniform rules, yields clearer composites and better brain-tumor segmentation.
desk verdict Solid systems paper on intent-conditioned DiT fusion with real multi-dataset evidence; the clinical-intent story is only partly isolated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Multi-scale Latent Adapter (MLA) plus the timestep-truncated medical semantic consistency loss: MLA extracts multi-scale 2D features from source latents before flattening and injects them into matching transformer depths; the semantic loss applies BioMedCLIP image–text alignment only after an early-noise cutoff so physical manifold reconstruction stays stable while late steps lock to the intent text.
What would settle it
On a held-out clinical cohort with independent expert labels, check whether MIND fused images still beat strong non-text baselines on radiologist diagnostic accuracy or lesion segmentation Dice when the guiding texts are wrong, generic, or replaced by human-written intents; a collapse of the claimed gains would falsify the intent-proxy claim.
Extended reading notes
Core claim
Guiding a diffusion transformer with intent-driven fusion texts, multi-scale latent spatial injection, and a timestep-truncated medical semantic consistency loss produces fused medical images that retain more source information, stay aligned with stated diagnostic goals, and improve brain-tumor segmentation relative to uniform-rule and prior text-driven fusion methods.
Load-bearing premise
That language-model fusion texts and CLIP-style image–text similarity are faithful stand-ins for real clinical diagnostic intent and medical image quality.
Editorial extensions
If this is right
- Fused CT/PET/SPECT–MRI and FLAIR–T1CE images retain higher entropy, mutual information, and contrast than eight published fusion methods on the reported benchmarks.
- nnU-Net tumor segmentation on BraTS improves, with the best mean rank across edema, non-enhancing, and enhancing subregions.
- Changing the fusion text retargets what structures and metabolic cues appear in the output, enabling interactive control.
- The same pipeline generalizes to non-radiology GFP–phase-contrast cell images when prompts are adapted.
- Intent-conditioned fusion becomes a building block for text-steerable clinical decision-support imaging.
Reading between the lines
- If intent texts are the control knob, hospital systems could store per-specialty prompt templates (e.g., bone vs soft tissue vs perfusion) instead of training separate fusion networks per modality pair.
- The early-noise truncation idea may transfer to other medical generative tasks where semantic losses currently fight pixel fidelity, such as MRI reconstruction or lesion inpainting.
- Failure modes will likely cluster where BioMedGPT misreads rare pathology or BioMedCLIP rewards superficial color/texture match; auditing those pairs is the next empirical stress test.
- Latency still sits in multi-second diffusion territory, so clinical bedside use would need distillation or fewer ODE steps before interactive reading-room deployment.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MIND proposes a Diffusion Transformer framework for medical image fusion conditioned on intent-driven fusion texts generated by BioMedGPT. The architecture freezes an SDXL VAE and Phi-3 DiT backbone, injects multi-scale 2D anatomical priors via a Multi-scale Latent Adapter (MLA) before sequence modeling, and trains with continuous flow matching plus a timestep-truncated medical semantic consistency loss (BioMedCLIP cosine, Eqs. 11–13) and anatomical reconstruction terms (Eq. 14). Experiments on Harvard (CT/PET/SPECT–MRI), BraTS FLAIR–T1CE, and GFP–PC report strong fusion metrics (Table 1, Table 5), ablations of MLA/L_sem/truncation (Table 4, Fig. 9), hyperparameter and allocation studies, ODE-stability checks, text-robustness (Table 6), and nnU-Net tumor segmentation on max-tumor 2D BraTS slices (Table 3, Fig. 8). The paper claims superior fusion quality, significantly improved downstream segmentation, and flexible interactive fusion for clinical decision support.
Significance. If the results hold under stronger clinical isolation, this is a solid systems contribution: it adapts DiT/flow-matching fusion to medical settings with an explicit spatial adapter and a carefully staged semantic loss, and it provides unusually thorough empirical support (eight fusion metrics across three Harvard tasks, external GFP, efficiency table, component ablations, allocation variants, wrong-text robustness, and a downstream segmentation endpoint). The intent-driven vs process-driven framing and the MLA residual injection are concrete engineering ideas others can reuse. The main significance risk is that the clinical-utility and “intent-driven” claims rest heavily on VLM proxies and a partially confounded segmentation protocol rather than on isolated causal evidence or human expert evaluation.
major comments (4)
- [§4.3.3, Table 3] Table 3 / §4.3.3: The claim that MIND “significantly improves downstream brain tumor segmentation accuracy” is only partly supported. Mean Rank 2.000 is driven by ED (0.786) and ET (0.616), while NET (0.706) is below TextFusion (0.731) and MR-T1CE-only (0.728). More importantly, there is no ablation that freezes the DiT+MLA stack and swaps intent-driven texts for process-driven/null texts, then re-trains nnU-Net. Without that control, gains cannot be attributed to intent guidance versus MLA, flow matching, or reconstruction losses. Please add this isolation experiment or soften the causal language.
- [§3.4, Eqs. (11)–(13)] §3.4, Eqs. (11)–(13) and Fig. 3: The medical semantic consistency loss treats BioMedCLIP cosine similarity (with threshold θ=0.85) as a surrogate for diagnostic correctness, and BioMedGPT texts as faithful encodings of clinical intent. Table 4 and Table 6 show CLIP and metric movement under these proxies, but no radiologist preference study, lesion-localization task, or pathology-verified labels validate that higher CLIP/L_sem corresponds to better clinical content rather than text–image surface match. This is load-bearing for the “intent-driven intelligent clinical decision support” claim. At minimum, report expert ratings on a subset or a task-based interactive protocol; otherwise narrow the claim to metric/CLIP-controllable fusion.
- [§4.1, Table 3] §4.1 Data Pre-processing: BraTS volumes are reduced to a single 2D slice per case via arg max of tumor mask area. Downstream Dice therefore measures 2D max-tumor-slice segmentation, not standard 3D BraTS evaluation. This choice is understandable for a 2D fusion backbone but should be stated explicitly in the abstract/claims, and preferably supplemented with multi-slice or 3D aggregation so that “brain tumor segmentation accuracy” is not over-read as full volumetric clinical performance.
- [Appendix A, Eq. (14)] Appendix A / training setup: Under data scarcity the paper uses complementary synthetic degradations of clean MRI as self-supervised GT pairs, plus VLM-generated texts, with no absolute multimodal fusion GT. That is a reasonable practical choice, but it interacts with L_rec (Eq. 14), which anchors reconstruction toward the anatomical source I_A. Please clarify how much reported Harvard/BraTS superiority depends on this self-supervised regime versus true multimodal supervision, and whether functional-modality fidelity is systematically under-penalized relative to anatomical structure.
minor comments (7)
- [Abstract] Abstract and §1: “significantly improves” should be qualified (which sub-regions, vs which baselines) once Table 3 is clarified.
- [Table 1] Table 1 SPECT-MRI: MIND AG (6.974) is not best; several baselines exceed it. The narrative of comprehensive superiority should acknowledge metric-level trade-offs more evenly.
- [§3.3–§3.4, Appendix D] Eq. (5) and surrounding text: “machanism” → “mechanism”; also check “wights” in §3.4 and “Rubustness” in Appendix D.
- [Fig. 9] Fig. 9 caption uses β in panel labels while the text discusses Φ, L_sem, and γ; align notation with Table 4.
- [Table 2] Table 2: report number of ODE/function evaluations and hardware parity conditions so inference-time comparisons to DDFM/Text-DiFuse are interpretable.
- [§4.2] §4.2: PyTorch “2.12.0” looks implausible at time of writing; verify version string.
- [§2.2] Related work could more clearly separate medical-specific text-fusion baselines from general IR/VIS methods when claiming novelty of intent-driven (vs process-driven) prompts.
Circularity Check
Empirical fusion system; only mild circularity is BioMedCLIP used both as L_sem training signal and as reported CLIP metric—main EN/MI/Dice claims remain external.
-
fitted input called prediction
[§3.4 Eqs. 11–13; Table 4 CLIP column; §4.5.1 ablation]
"We first define the cosine similarity between the generated image and the textual instruction as: S_cos(Î_1,T)=E_CLIP−I(Î_1)·E_CLIP−T(T)/… where E_CLIP−I and E_CLIP−T denote the frozen image and text encoders of BioMedCLIP [48]. … L_sem = E_t∼U(0,1)[I(t>τ)·t·ℓ_sem(Î_1,T)]. … Specifically, the CLIP score improves by 15.49%, 8.86%, and 23.88% across the three subsets."
L_sem directly maximizes BioMedCLIP image–text cosine similarity (with threshold θ). Reporting the same BioMedCLIP cosine as ‘CLIP’ / MedCLIP-S and attributing large CLIP gains to L_sem is statistically forced for that metric: the evaluation coordinate is the training objective. This does not make EN/MI/Dice circular, but CLIP-score ‘predictions’ of semantic locking are not independent evidence.
full rationale
MIND is a systems/ML paper, not a closed-form derivation. The generative objective (continuous flow matching, Eq. 9), physical reconstruction toward anatomical sources (L_rec, Eq. 14), Multi-scale Latent Adapter injection, and BioMedGPT intent texts are design choices evaluated against eight external SOTA methods on held-out Harvard/BraTS/GFP splits and an independent nnU-Net segmentation endpoint (Tables 1, 3, 5). Those primary metrics (EN, MI, SD, AG, SF, Dice) are not algebraic restatements of the training losses. The only mild circularity is that BioMedCLIP cosine similarity is both the training penalty (Eqs. 11–13) and a reported alignment score (Table 4 CLIP column, ablation CLIP gains); CLIP improvements are therefore partly by construction and should not be read as independent evidence of clinical intent. That does not force the fusion-quality or segmentation results. No self-definitional identity, uniqueness theorem from overlapping authors, or ansatz-smuggled-via-citation chain carries the central claim. Score 2 reflects dual-use of the VLM proxy only.
Assumptions & free parameters
free parameters (6)
- adapter scale α =
1.0
- semantic timestep truncation τ =
0.5
- λ_L1, λ_SSIM, λ_sem =
0.15, 0.05, 0.1
- semantic cosine threshold θ =
0.85
- LoRA rank and learning rate schedule =
rank 64; lr 1e-7; 20 epochs
- MLA scale count S and linear allocation s(l) =
S=3; layers 0–10/11–21/22–31
assumptions (7)
- domain assumption Continuous flow matching with OT path x_t = t x_1 + (1-t) x_0 is an adequate generative objective for pixel-faithful medical fusion.
- domain assumption Frozen SDXL VAE latents preserve clinically relevant anatomy and functional signal at 8× downsampling.
- domain assumption Source image pairs are rigidly co-registered and 2D slices (max-tumor for BraTS) represent the clinical fusion task.
- ad hoc to paper BioMedCLIP image–text cosine similarity is a valid surrogate for medical semantic correctness of fused outputs.
- ad hoc to paper Intent descriptions from BioMedGPT under the paper’s prompts correctly encode diagnostic goals without modality name leakage.
- ad hoc to paper Complementary synthetic degradations of clean MRI yield a valid self-supervised GT for learning fusion under data scarcity.
- standard math Standard analysis tools for ODEs / Lipschitz continuity and transformer depth–frequency progression justify truncation smoothness and linear scale allocation.
invented entities (3)
-
Multi-scale Latent Adapter (MLA)
-
Timestep-truncated multimodal medical semantic consistency loss
-
Intent-driven fusion text paradigm (vs process-driven prompts)
Cite this review
Pith. "Pith review of MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion." pith.science (2026). https://pith.science/paper/VLXYSZVV
@misc{pith2026260728565,
author = {Pith},
title = {Pith review of: MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion},
year = {2026},
howpublished = {\url{https://pith.science/paper/VLXYSZVV}},
note = {Machine review of arXiv:2607.28565}
}
read the original abstract
Medical image fusion aims to integrate complementary information from diverse imaging modalities to support clinical diagnosis. Existing methods typically apply uniform fusion rules globally, lacking a deep understanding of diagnostic intents and pathological structures. To address these limitations, we propose MIND, a Multimodal Intent-Driven Network via Diffusion Transformers (DiTs) for medical image fusion. Specifically, we utilize BioMedGPT to generate intent-driven fusion texts from source images, guiding the fusion process with pathology-aware diagnostic intents. To combat the loss of 2D spatial continuity caused by 1D sequence flattening in DiTs, we design a Multi-scale Latent Adapter. This module explicitly extracts source image features before serialization, injecting them into the network via strict dimensional alignment to effectively supplement image features. To resolve the semantic shift caused by decoupling image outputs from diagnostic intents, we design a medical semantic consistency loss. This loss ensures deep semantic locking between fused images and fusion texts while maintaining the stability of the underlying physical manifold reconstruction. Comprehensive experiments on the Harvard, BraTS, and GFP datasets reveal that MIND delivers superior fusion quality, significantly improves downstream brain tumor segmentation accuracy, and enables flexible interactive fusion, holding significant promise for intent-driven intelligent clinical decision support systems.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, and et al. 2024. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone. arXiv:2404.14219 [cs.CL]
arXiv 2024
-
[2]
Berthold Bein. 2006. Entropy. Best Practice & Research Clinical Anaesthesiology 20, 1 (2006), 101–109
2006
-
[3]
Zihan Cao, Yu Zhong, Ziqi Wang, and Liang-Jian Deng. 2025. MMAIF: Multi-task and Multi-degradation All-in-One for Image Fusion with Language Guidance. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 11744–11754
2025
-
[4]
CC Chaithra, NL Taranath, LM Darshan, and CK Subbaraya. 2018. A survey on image fusion techniques and performance metrics. In 2018 Second International Conference on Electronics, Communication and Aerospace Technology (ICECA) . IEEE, 995–999
2018
-
[5]
Chunyang Cheng, Tianyang Xu, Xiao-Jun Wu, Hui Li, Xi Li, Zhangyong Tang, and Josef Kittler. 2025. TextFusion: Unveiling the power of textual semantics for controllable image fusion. Information Fusion 117 (2025), 102790
2025
-
[6]
Allen A Goldstein. 1977. Optimization of Lipschitz continuous functions. Mathe- matical Programming 13, 1 (1977), 14–22
1977
-
[7]
Yu Han, Yunze Cai, Yin Cao, and Xiaoming Xu. 2013. A new image fusion performance metric based on visual information fidelity. Information Fusion 14 (2013), 127–135
2013
-
[8]
Tiankai Hang, Shuyang Gu, Chen Li, Jianmin Bao, Dong Chen, Han Hu, Xin Geng, and Baining Guo. 2023. Efficient Diffusion Training via Min-SNR Weighting Strategy. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 7441–7451
2023
Show all 56 references
-
[9]
Dan He, Weisheng Li, Guofen Wang, Yuping Huang, and Shiqiang Liu. 2025. DM- FNet: Unified Multimodal Medical Image Fusion via Diffusion Process-Trained Encoder-Decoder. IEEE Transactions on Multimedia 27 (2025), 9415–9428
2025
-
[10]
Haithem Hermessi, Olfa Mourali, and Ezzeddine Zagrouba. 2021. Multimodal medical image fusion review: Theoretical background and recent advances.Signal Processing 183 (2021), 108036
2021
-
[11]
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi
-
[12]
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations (ICLR)
2022
-
[13]
Junjia Huang, Pengxiang Yan, Jiyang Liu, Jie Wu, Zhao Wang, Yitong Wang, Liang Lin, and Guanbin Li. 2025. DreamFuse: Adaptive Image Fusion with Diffusion Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 17292–17301
2025
-
[14]
Fabian Isensee, Paul F Jaeger, Simon AA Kohl, Jens Petersen, and Klaus H Maier- Hein. 2021. nnU-Net: a self-configuring method for deep learning-based biomed- ical image segmentation. Nature Methods 18, 2 (2021), 203–211
2021
-
[15]
Dasarathy
Alex Pappachen James and Belur V. Dasarathy. 2014. Medical image fusion: A survey of the state of the art. Information Fusion 19 (2014), 4–19. Special Issue on Information Fusion in Medical Image Computing and Systems
2014
-
[16]
Olga A Koroleva, Matthew L Tomlinson, David Leader, Peter Shaw, and John H Doonan. 2005. High-throughput protein localization in Arabidopsis using Agrobacterium-mediated transient expression of GFP-ORF fusions. The Plant Journal 41, 1 (2005), 162–174
2005
-
[17]
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger. 2004. Estimating mutual information. Phys. Rev. E 69, 6 (2004), 066138
2004
-
[18]
Huafeng Li, Dayong Su, Qing Cai, and Yafei Zhang. 2025. BSAFusion: A Bidirec- tional Stepwise Feature Alignment Network for Unaligned Medical Image Fusion. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , Vol. 39. 4725–4733. Preprint, August, 2025 Yunz...
2025
-
[19]
Jiayang Li, Chengjie Jiang, Junjun Jiang, Pengwei Liang, Jiayi Ma, and Liqiang Nie. 2025. Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach. IEEE Transactions on Pattern Analysis and Machine Intelligence (2025), 1–18
2025
-
[20]
Weisheng Li, Pengtao Jia, Dan He, Shiqiang Liu, Guofen Wang, and Yuping Huang. 2026. SAFusion: Scenario-Adaptive Network for Multimodal Medical Image Fusion. IEEE Journal of Biomedical and Health Informatics (2026), 1–14
2026
-
[21]
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2023. Flow Matching for Generative Modeling. arXiv:2210.02747 [cs.LG]
2023 arXiv
-
[22]
Jinyuan Liu, Xingyuan Li, Zirui Wang, Zhiying Jiang, Wei Zhong, Wei Fan, and Bin Xu. 2025. PromptFusion: Harmonized Semantic Prompt Learning for Infrared and Visible Image Fusion. IEEE/CAA Journal of Automatica Sinica 12, 3 (2025), 502–515
2025
-
[23]
Yu Liu, Xun Chen, Juan Cheng, and Hu Peng. 2017. A medical image fusion method based on convolutional neural networks. In 2017 20th International Con- ference on Information Fusion (FUSION) . 1–7
2017
-
[24]
Jiayi Ma, Yong Ma, and Chang Li. 2019. Infrared and visible image fusion methods and applications: A survey. Information Fusion 45 (2019), 153–178
2019
-
[25]
Jiayi Ma, Wei Yu, Pengwei Liang, Chang Li, and Junjun Jiang. 2019. FusionGAN: A generative adversarial network for infrared and visible image fusion. Information Fusion 48 (2019), 11–26
2019
-
[26]
Menze, Andras Jakab, Stefan Bauer, and et al
Bjoern H. Menze, Andras Jakab, Stefan Bauer, and et al. 2015. The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS). IEEE Transactions on Medical Imaging 34, 10 (2015), 1993–2024
2015
-
[27]
Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. 2024. T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion Models. InProceedings of the AAAI Conference on Artificial Intelligence (AAAI), Vol. 38...
2024
-
[28]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. arXiv:2307.01952 [cs.CV]
2023 arXiv
-
[29]
G Poornima and L Anand. 2025. Medical image fusion model using CT and MRI images based on dual scale weighted fusion based residual attention network with encoder-decoder architecture. Biomedical Signal Processing and Control 108 (2025), 107932
2025
-
[30]
Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy. 2021. Do vision transformers see like convolutional neural networks? Advances in Neural Information Processing Systems 34 (2021), 12116– 12128
2021
-
[31]
Mojtaba Safari, Ali Fatemi, and Louis Archambault. 2023. MedFusionGAN: mul- timodal medical image fusion using an unsupervised deep generative adversarial network. BMC Medical Imaging 23, 1 (2023), 203
2023
-
[32]
Robert Shapley, Peter Lennie, et al. 1985. Spatial frequency analysis in the visual system. Annual Review of Neuroscience 8, 1 (1985), 547–581
1985
-
[33]
Stefan Siegmund, Christine Nowak, and Josef Diblík. 2016. A generalized Picard- Lindelöf theorem. Electronic Journal of Qualitative Theory of Differential Equations 2016, 28 (2016), 1–8
2016
-
[34]
Yifei Sun, Yuzhi He, Junhao Jia, Jinhong Wang, Ruiquan Ge, Changmiao Wang, and Hongxia Xu. 2026. WDT-MD: Wavelet Diffusion Transformers for Microa- neurysm Detection in Fundus Images. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 9242–9250
2026
-
[35]
Wei Tan, Prayag Tiwari, Hari Mohan Pandey, Catarina Moreira, and Amit Kumar Jaiswal. 2025. Multimodal medical image fusion algorithm in the era of big data. Neural Computing and Applications 37, 28 (2025), 22995–23015
2025
-
[36]
Wang, E.P
Z. Wang, E.P. Simoncelli, and A.C. Bovik. 2003. Multiscale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, Vol. 2. 1398–1402 Vol.2
2003
-
[37]
Zeyu Wang, Libo Zhao, Jizheng Zhang, Rui Song, Haiyu Song, Jiana Meng, and Shidong Wang. 2025. Multi-text guidance is important: Multi-modality image fusion via large generative vision-language model. International Journal of Computer Vision 133, 7 (2025), 4646–4668
2025
-
[38]
Caifeng Xia, Hongwei Gao, Wei Yang, and Jiahui Yu. 2025. MSDT: Multiscale Diffusion Transformer for Multimodality Image Fusion. IEEE Transactions on Emerging Topics in Computational Intelligence 9, 3 (2025), 2269–2283
2025
-
[39]
Haozhe Xiang, Han Zhang, Yu Cheng, Xiongwen Quan, and Wanwan Huang
-
[40]
Shitao Xiao, Yueze Wang, Junjie Zhou, Huaying Yuan, Xingrun Xing, Ruiran Yan, Chaofan Li, Shuting Wang, Tiejun Huang, and Zheng Liu. 2025. Omnigen: Unified image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 13294–13304
2025
-
[41]
Xinyu Xie, Xiaozhi Zhang, Xinglong Tang, Jiaxi Zhao, Dongping Xiong, Lijun Ouyang, Bin Yang, Hong Zhou, Bingo Wing-Kuen Ling, and Kok Lay Teo. 2025. MACTFusion: Lightweight Cross Transformer for Adaptive Multimodal Medical Image Fusion. IEEE Journal of Biomedical and Health In...
2025
-
[42]
Xydeas and V
C.S. Xydeas and V. Petrović. 2000. Objective image fusion performance measure. Electronics Letters 36, 4 (2000), 308–309
2000
-
[43]
Wu, and Mengye Lyu
Huaishui Yang, Shaojun Liu, Yilong Liu, Lingyan Zhang, Shoujin Huang, Jiayu Zheng, Jingzhe Liu, Hua Guo, Ed X. Wu, and Mengye Lyu. 2025. An Unsupervised Learning Approach for Reconstructing 3T-Like Images From 0.3T MRI Without Paired Training Data. IEEE Transactions on Medical...
2025
-
[44]
Xunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang, and Jiayi Ma. 2024. Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 27026–27035
2024
-
[45]
Jun Yue, Leyuan Fang, Shaobo Xia, Yue Deng, and Jiayi Ma. 2023. Dif-Fusion: Toward High Color Fidelity in Infrared and Visible Image Fusion With Diffusion Models. IEEE Transactions on Image Processing 32 (2023), 5705–5720
2023
-
[46]
Hao Zhang, Lei Cao, and Jaiyi Ma. 2024. Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model. InAdvances in Neural Information Processing Systems , Vol. 37. 39552–39572
2024
-
[47]
Davison, Hui Ren, Jing Huang, Chen Chen, Yuyin Zhou, Sunyang Fu, Wei Liu, Tianming Liu, Xiang Li, Yong Chen, Lifang He, James Zou, Quanzheng Li, Hongfang Liu, and Lichao Sun
Kai Zhang, Rong Zhou, Eashan Adhikarla, Zhiling Yan, Yixin Liu, Jun Yu, Zhengliang Liu, Xun Chen, Brian D. Davison, Hui Ren, Jing Huang, Chen Chen, Yuyin Zhou, Sunyang Fu, Wei Liu, Tianming Liu, Xiang Li, Yong Chen, Lifang He, James Zou, Quanzheng Li, Hongfang Liu, and Lichao ...
2024 doi
-
[48]
Lungren, Tristan Naumann, Sheng Wang, and Hoifung Poon
Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, Cliff Wong, Andrea Tupini, Yu Wang, Matt Mazzola, Swadheen Shukla, Lars Liden, Jianfeng Gao, Angela Crabtree, Brian Piening, Carlo Bifulco, Matthew P....
2025
-
[49]
Yu Zhang, Yu Liu, Peng Sun, Han Yan, Xiaolin Zhao, and Li Zhang. 2020. IFCNN: A general image fusion framework based on convolutional neural network. In- formation Fusion 54 (2020), 99–118
2020
-
[50]
Hang Zhao, Orazio Gallo, Iuri Frosio, and Jan Kautz. 2016. Loss functions for image restoration with neural networks. IEEE Transactions on Computational Imaging 3, 1 (2016), 47–57
2016
-
[51]
Zixiang Zhao, Haowen Bai, Jiangshe Zhang, Yulun Zhang, Shuang Xu, Zudi Lin, Radu Timofte, and Luc Van Gool. 2023. CDDFuse: Correlation-Driven Dual- Branch Feature Decomposition for Multi-Modality Image Fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pa...
2023
-
[52]
Zixiang Zhao, Haowen Bai, Yuanzhi Zhu, Jiangshe Zhang, Shuang Xu, Yulun Zhang, Kai Zhang, Deyu Meng, Radu Timofte, and Luc Van Gool. 2023. DDFM: Denoising Diffusion Model for Multi-Modality Image Fusion. In Proceedings of the IEEE/CVF International Conference on Computer Visio...
2023
-
[53]
Zixiang Zhao, Lilun Deng, Haowen Bai, Yukun Cui, Zhipeng Zhang, Yulun Zhang, Haotong Qin, Dongdong Chen, Jiangshe Zhang, Peng Wang, and Luc Van Gool. 2024. Image fusion via vision-language model. In Proceedings of the 41st International Conference on Machine Learning (ICML) . JMLR.org
2024
-
[54]
Tao Zhou, Qi Li, Huiling Lu, Qianru Cheng, and Xiangxiang Zhang. 2023. GAN review: Models and medical image fusion applications. Information Fusion 91 (2023), 134–148
2023
-
[2021]
In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP)
CLIPScore: A Reference-free Evaluation Metric for Image Captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP). 7514–7528
2021
-
[2025]
IEEE Journal of Biomedical and Health Informatics (2025), 1–14
SMFusion: Semantic-Preserving Fusion of Multimodal Medical Images for Enhanced Clinical Diagnosis. IEEE Journal of Biomedical and Health Informatics (2025), 1–14
2025
Reviewed July 31, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.