Pith. sign in

REVIEW 3 major objections 3 minor 1 cited by

MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Fused infrared-visible images made readable for visible-trained models

desk verdict The submission body is completely unrelated to the abstract: no MambaTrans method, experiments, or equations exist in the paper—only a plausible abstract and the full text of a different paper on video restoration. read the letter →

arxiv 2508.07803 v1 pith:22DFRMUR submitted 2025-08-11 cs.CV

classification cs.CV
keywords multimodalimagefusioninfrared-visibleimagestranslationstatespacemodelobjectdetectionsemanticsegmentationlargelanguagepriorsdownstreamvisualtasks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the distribution gap between fused infrared-visible images and the visible images used to train off-the-shelf detectors and segmenters can be closed by a dedicated translator, MambaTrans. The translator is conditioned on text descriptions from a multimodal large language model and on masks from a semantic segmentation model, and it is trained using the downstream detection loss as an explicit training signal. The claim is that, after translation, frozen pre-trained detection and segmentation models perform better on multimodal fused images than on the raw fused images, sometimes even at the level of visible-only inputs. If true, this would let practitioners apply existing visible-trained models to multimodal sensor data without retraining any downstream model.

What carries the argument

The Multi-Model State Space Block (MSSB) is the central component. It combines a mask-image-text cross-attention mechanism with a 3D-Selective Scan Module, which scans feature maps along spatial and channel dimensions to model long-range dependencies among the image, the semantic mask, and the text description, thereby aligning the translated fused image with what a visible-trained detector or segmenter expects.

What would settle it

A decisive test: train MambaTrans using a single detector's loss, then evaluate on a held-out detection dataset with a different detector architecture. If the translated images do not improve that unseen detector, or if performance improves only on the exact detector used in training, the central claim of general compatibility with pre-trained models collapses.

Watch

Extended reading notes

Core claim

MambaTrans proposes a trainable image translator that converts fused infrared-visible images into a form compatible with pre-trained models originally trained only on visible images. The translator takes the fused image, a text description generated by a multimodal LLM, and semantic segmentation masks as inputs, and its core component, the Multi-Model State Space Block, uses mask-image-text cross-attention and a 3D-Selective Scan Module to capture long-range dependencies across the three modalities. During training, the translator minimizes the detection loss of the target pre-trained detector, so the translation is explicitly optimized to make the fused image more amenable to that detector.

Load-bearing premise

The load-bearing assumption is that optimizing the translator with the target detectors' loss functions will improve performance on those same detectors across datasets, rather than merely overfitting to the specific detectors or the exact loss landscape used during training.

Editorial extensions

If this is right

  • If MambaTrans works as claimed, visible-trained detectors and segmenters can be applied to fused infrared-visible imagery without parameter changes, extending the usability of existing models.
  • Because the translator is trained once with frozen downstream models, the same translated images could be fed to multiple downstream models, avoiding per-task retraining.
  • The use of LLM descriptions and segmentation masks as conditioning inputs suggests that semantic priors can steer translation toward features that downstream models actually rely on.
  • The 3D selective scan is claimed to capture long-range dependencies jointly across image, mask, and text, offering a potential explanation for why translation may generalize beyond simple pixel-level adjustment.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper does not explicitly report: train MambaTrans with one detector and evaluate with a different detector architecture; the degree of transfer would separate genuine distribution alignment from overfitting to the training detector's loss landscape.
  • The dependence on a multimodal LLM raises an ablation question: how much of the gain comes from the text prior versus the mask prior? Removing each input and measuring the drop would isolate their contributions.
  • If the method generalizes across different fusion algorithms, it would establish a general 'modality adapter' design: a translator that maps any sensor modality into the input distribution of a frozen backbone.
  • The paper leaves open whether the translated images also benefit human perception, since evaluation is through downstream models rather than perceptual quality.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The manuscript is titled and abstracted as "MambaTrans," a multimodal fusion image translation method that uses LLM descriptions and semantic segmentation masks to adapt infrared-visible fused images to downstream detection and segmentation models trained on visible images. The abstract claims that MambaTrans's Multi-Model State Space Block, 3D-Selective Scan Module, and mask-image-text cross-attention, trained with detection loss, improve downstream task performance on public datasets. However, the submitted full text is not about MambaTrans at all. It is the complete text of a separate paper, "DiTVR: Zero-Shot Diffusion Transformer for Video Restoration" by Sicheng Gao, Nancy Mehta, Zongwei Wu, and Radu Timofte, including its own abstract, introduction, methodology, experiments, tables, figures, and references. No equations, architecture descriptions, training details, datasets, baselines, or experimental results for MambaTrans appear anywhere in the manuscript. The central claim is therefore entirely unsupported by the submitted document.

Significance. If the MambaTrans idea were actually developed and validated, it could be of interest to the multimodal fusion and downstream-task adaptation community: adapting fused infrared-visible images to detectors and segmenters trained on visible images is a legitimate and practically relevant problem. However, the manuscript provides no method, no experiments, and no analysis for MambaTrans. The only concrete technical content is an unrelated video-restoration paper, which has no bearing on the abstract's claims. No code, proofs, or reproducible artifacts are included. As submitted, the significance of the claimed contribution cannot be assessed because the contribution itself is absent.

major comments (3)
  1. [Full Text (entire manuscript body)] The body of the manuscript is the complete text of a different paper, "DiTVR: Zero-Shot Diffusion Transformer for Video Restoration," including its own abstract, introduction, related work, equations (1)–(12), methodology sections, experimental tables (Tables 1–4), and reference list. None of these describe MambaTrans, the Multi-Model State Space Block, the 3D-Selective Scan Module, or mask-image-text cross-attention. This is not a presentation issue; it is the absence of the proposed method and its evaluation. The central claim of the abstract is therefore unsupported by any evidence in the manuscript.
  2. [Abstract (last two sentences)] The abstract states that "MambaTrans minimizes detection loss during training" and then claims "favorable results in pre-trained models without adjusting their parameters." If the same detection models and losses used to train the translator are also used as evaluators in downstream experiments, the claimed improvement could be a form of fitting the translator to the evaluator. This is a correctness risk that would need to be addressed with held-out detectors or datasets. However, because the manuscript contains no training or evaluation protocol for MambaTrans, the concern cannot even be checked; the absence of any experiments is the primary problem.
  3. [Abstract vs. Full Text (no experimental support)] The abstract claims "Experiments on public datasets show that MambaTrans effectively improves multimodal image performance in downstream tasks." No such experiments appear in the body. The tables and figures in the body measure PSNR, SSIM, LPIPS, warping error, and frame similarity for video super-resolution and denoising on Vid4, SPMC, and DAVIS. These datasets, metrics, and tasks are unrelated to infrared-visible image fusion or to object detection and semantic segmentation downstream evaluation. Thus the claimed empirical validation is entirely missing.
minor comments (3)
  1. [Header/front matter] The manuscript header shows arXiv:2508.07811v1 while the submitted identifier is arXiv:2508.07803; this mismatch is consistent with the body being a different paper's text.
  2. [Methodology, DiTVR text] Within the unrelated DiTVR text, the Methodology section contains placeholders such as "(Sec. to Sec. )" and "(Sec. )" where section references should appear, indicating that even the included paper is not in polished form.
  3. [References] The reference list is entirely for the DiTVR paper and contains no citations related to multimodal fusion, infrared-visible image translation, object detection priors, or semantic segmentation masks, further confirming that the submitted body does not support the abstract.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity can be established: the manuscript body is an unrelated paper (DiTVR) and contains no MambaTrans method, equations, or experiments to exhibit a reduction.

full rationale

The submitted full text is entirely the text of 'DiTVR: Zero-Shot Diffusion Transformer for Video Restoration' by Sicheng Gao et al., with its own abstract, methodology, equations, experiments, and references. The MambaTrans abstract claims a multimodal fusion translator using LLM descriptions, segmentation masks, and detection-loss training, but none of these components, no derivation, and no experimental protocol appear in the body. Under the hard rule that circularity may only be claimed when the paper's own equations or self-citations exhibit the reduction, no such reduction can be quoted. The reader-identified concern—that minimizing detection loss during training while evaluating on detectors could be a fitted-input-called-prediction issue—cannot be substantiated because the manuscript provides neither the training-loss specification nor the evaluation setup. This is a structural absence of the claimed argument rather than circularity. Therefore the score is 0, with no circular steps identified; the appropriate concern is manuscript integrity or missing content, not circular reasoning.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

Only the abstract is usable for the MambaTrans claim because the full text is an unrelated paper. No free parameters or invented entities (new particles, forces, etc.) can be identified from the abstract. The two domain assumptions above are the main unstated premises of the claimed approach.

assumptions (2)
  • domain assumption The performance gap of downstream models on fused images is caused by pixel distribution differences, and translating fused images to be more visible-like will improve downstream task performance.
    The abstract motivates MambaTrans by 'significant pixel distribution differences between visible and multimodal fusion images' degrading downstream performance. This causal assumption is not demonstrated in the abstract.
  • domain assumption Text descriptions from a multimodal LLM and masks from a segmentation model provide sufficient complementary information to guide the translation so that downstream models improve.
    The abstract states these are used as input, but no evidence is given that they carry useful, non-redundant signal for translation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks." pith.science (2026). https://pith.science/paper/22DFRMUR

@misc{pith2026250807803,
  author       = {Pith},
  title        = {Pith review of: MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/22DFRMUR}},
  note         = {Machine review of arXiv:2508.07803}
}
read the original abstract

The goal of multimodal image fusion is to integrate complementary information from infrared and visible images, generating multimodal fused images for downstream tasks. Existing downstream pre-training models are typically trained on visible images. However, the significant pixel distribution differences between visible and multimodal fusion images can degrade downstream task performance, sometimes even below that of using only visible images. This paper explores adapting multimodal fused images with significant modality differences to object detection and semantic segmentation models trained on visible images. To address this, we propose MambaTrans, a novel multimodal fusion image modality translator. MambaTrans uses descriptions from a multimodal large language model and masks from semantic segmentation models as input. Its core component, the Multi-Model State Space Block, combines mask-image-text cross-attention and a 3D-Selective Scan Module, enhancing pure visual capabilities. By leveraging object detection prior knowledge, MambaTrans minimizes detection loss during training and captures long-term dependencies among text, masks, and images. This enables favorable results in pre-trained models without adjusting their parameters. Experiments on public datasets show that MambaTrans effectively improves multimodal image performance in downstream tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. InfraNet: Quality-Aware RGB Guidance for Efficient Infrared Object Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    QualGate-regulated RGB guidance during training produces efficient IR-only and dual-modal detectors that match or beat equal-fusion baselines under low light and adverse weather.

Reference graph

Works this paper leans on

47 extracted references · 40 canonical work pages · cited by 1 Pith paper

  1. [1]

    A deep learning framework for infrared and visible image fusion without strict registration

    Huafeng Li, Junyu Liu, Yafei Zhang, and Yu Liu. A deep learning framework for infrared and visible image fusion without strict registration. International Journal of Computer Vision, 132 0 (5): 0 1625--1644, 2024 a

  2. [2]

    All-weather multi-modality image fusion: Unified framework and 100k benchmark

    Xilai Li, Wuyang Liu, Xiaosong Li, Fuqiang Zhou, Huafeng Li, and Feiping Nie. All-weather multi-modality image fusion: Unified framework and 100k benchmark. arXiv preprint arXiv:2402.02090, 2024 b

  3. [3]

    Infrared and visible image fusion based on domain transform filtering and sparse representation

    Xilai Li, Haishu Tan, Fuqiang Zhou, Gao Wang, and Xiaosong Li. Infrared and visible image fusion based on domain transform filtering and sparse representation. Infrared Physics & Technology, 131: 0 104701, 2023 a

  4. [4]

    Infrared and visible image fusion: From data compatibility to task adaption

    Jinyuan Liu, Guanyao Wu, Zhu Liu, Di Wang, Zhiying Jiang, Long Ma, Wei Zhong, and Xin Fan. Infrared and visible image fusion: From data compatibility to task adaption. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024 a

  5. [5]

    Simultaneous tri-modal medical image fusion and super-resolution using conditional diffusion model

    Yushen Xu, Xiaosong Li, Yuchan Jie, and Haishu Tan. Simultaneous tri-modal medical image fusion and super-resolution using conditional diffusion model. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 635--645. Springer, 2024

  6. [7]

    Fast r-cnn

    Ross Girshick. Fast r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 1440--1448, 2015

  7. [8]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \'a r, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer vision--ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13, pages 740--755. Springer, 2014

  8. [9]

    Visualizing data using t-sne

    Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9 0 (11), 2008

Show all 47 references
  1. [10]

    Tsjnet: A multi-modality target and semantic awareness joint-driven image fusion network

    Yuchan Jie, Yushen Xu, Xiaosong Li, and Haishu Tan. Tsjnet: A multi-modality target and semantic awareness joint-driven image fusion network. arXiv preprint arXiv:2402.01212, 2024

  2. [11]

    Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection

    Jinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu, Risheng Liu, Wei Zhong, and Zhongxuan Luo. Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection. In Proceedings of the IEEE/CVF conference on compu...

  3. [12]

    Fs-diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

    Yuchan Jie, Yushen Xu, Xiaosong Li, Fuqiang Zhou, Jianming Lv, and Huafeng Li. Fs-diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution. Information Fusion, 121: 0 103146, 2025

  4. [13]

    A task-guided, implicitly-searched and metainitialized deep model for image fusion

    Risheng Liu, Zhu Liu, Jinyuan Liu, Xin Fan, and Zhongxuan Luo. A task-guided, implicitly-searched and metainitialized deep model for image fusion. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024 b

  5. [14]

    Image fusion in the loop of high-level vision tasks: A semantic-aware real-time infrared and visible image fusion network

    Linfeng Tang, Jiteng Yuan, and Jiayi Ma. Image fusion in the loop of high-level vision tasks: A semantic-aware real-time infrared and visible image fusion network. Information Fusion, 82: 0 28--42, 2022 a

  6. [15]

    Mrfs: Mutually reinforcing image fusion and segmentation

    Hao Zhang, Xuhui Zuo, Jie Jiang, Chunchao Guo, and Jiayi Ma. Mrfs: Mutually reinforcing image fusion and segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26974--26983, 2024

  7. [16]

    Image-to-image translation: Methods and applications

    Yingxue Pang, Jianxin Lin, Tao Qin, and Zhibo Chen. Image-to-image translation: Methods and applications. IEEE Transactions on Multimedia, 24: 0 3859--3881, 2021

  8. [17]

    Source-free open compound domain adaptation in semantic segmentation

    Yuyang Zhao, Zhun Zhong, Zhiming Luo, Gim Hee Lee, and Nicu Sebe. Source-free open compound domain adaptation in semantic segmentation. IEEE Transactions on Circuits and Systems for Video Technology, 32 0 (10): 0 7019--7032, 2022

  9. [18]

    Infragan: A gan architecture to transfer visible images to infrared domain

    Mehmet Akif \"O zkano g lu and Sedat Ozer. Infragan: A gan architecture to transfer visible images to infrared domain. Pattern Recognition Letters, 155: 0 69--76, 2022

  10. [19]

    Cnn-based thermal infrared person detection by domain adaptation

    Christian Herrmann, Miriam Ruf, and J \"u rgen Beyerer. Cnn-based thermal infrared person detection by domain adaptation. In Autonomous Systems: Sensors, Vehicles, Security, and the Internet of Everything, volume 10643, pages 38--43. SPIE, 2018

  11. [20]

    Hallucidet: hallucinating rgb modality for person detection through privileged information

    Heitor Rapela Medeiros, Fidel A Guerrero Pena, Masih Aminbeidokhti, Thomas Dubail, Eric Granger, and Marco Pedersoli. Hallucidet: hallucinating rgb modality for person detection through privileged information. In Proceedings of the IEEE/CVF Winter Conference on Applications of...

  12. [21]

    Modality translation for object detection adaptation without forgetting prior knowledge

    Heitor Rapela Medeiros, Masih Aminbeidokhti, Fidel Alejandro Guerrero Pe \ n a, David Latortue, Eric Granger, and Marco Pedersoli. Modality translation for object detection adaptation without forgetting prior knowledge. In European Conference on Computer Vision, pages 51--68. ...

  13. [22]

    Mamba: Linear-time sequence modeling with selective state spaces

    Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023

  14. [23]

    Vmamba: Visual state space model

    Yue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang, Qixiang Ye, Jianbin Jiao, and Yunfan Liu. Vmamba: Visual state space model. Advances in neural information processing systems, 37: 0 103031--103063, 2024 c

  15. [24]

    Mambair: A simple baseline for image restoration with state-space model

    Hang Guo, Jinmin Li, Tao Dai, Zhihao Ouyang, Xudong Ren, and Shu-Tao Xia. Mambair: A simple baseline for image restoration with state-space model. In European conference on computer vision, pages 222--241. Springer, 2024

  16. [25]

    Vision mamba: efficient visual representation learning with bidirectional state space model

    Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. Vision mamba: efficient visual representation learning with bidirectional state space model. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org, 2024

  17. [26]

    Glu variants improve transformer

    Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020

  18. [27]

    Faster r-cnn: Towards real-time object detection with region proposal networks

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28, 2015

  19. [28]

    Piafusion: A progressive infrared and visible image fusion network based on illumination aware

    Linfeng Tang, Jiteng Yuan, Hao Zhang, Xingyu Jiang, and Jiayi Ma. Piafusion: A progressive infrared and visible image fusion network based on illumination aware. Information Fusion, 83: 0 79--92, 2022 b

  20. [29]

    Every sam drop counts: Embracing semantic priors for multi-modality image fusion and beyond

    Guanyao Wu, Haoyu Liu, Hongming Fu, Yichuan Peng, Jinyuan Liu, Xin Fan, and Risheng Liu. Every sam drop counts: Embracing semantic priors for multi-modality image fusion and beyond. arXiv preprint arXiv:2503.01210, 2025

  21. [30]

    Dcevo: Discriminative cross-dimensional evolutionary learning for infrared and visible image fusion

    Jinyuan Liu, Bowei Zhang, Qingyun Mei, Xingyuan Li, Yang Zou, Zhiying Jiang, Long Ma, Risheng Liu, and Xin Fan. Dcevo: Discriminative cross-dimensional evolutionary learning for infrared and visible image fusion. arXiv preprint arXiv:2503.17673, 2025

  22. [31]

    Where elegance meets precision: towards a compact, automatic, and flexible framework for multi-modality image fusion and applications

    Jinyuan Liu, Guanyao Wu, Zhu Liu, Long Ma, Risheng Liu, and Xin Fan. Where elegance meets precision: towards a compact, automatic, and flexible framework for multi-modality image fusion and applications. In Proceedings of the Thirty-Third International Joint Conference on Arti...

  23. [32]

    Equivariant multi-modality image fusion

    Zixiang Zhao, Haowen Bai, Jiangshe Zhang, Yulun Zhang, Kai Zhang, Shuang Xu, Dongdong Chen, Radu Timofte, and Luc Van Gool. Equivariant multi-modality image fusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 25912--25921, 2024

  24. [33]

    Cddfuse: Correlation-driven dual-branch feature decomposition for multi-modality image fusion

    Zixiang Zhao, Haowen Bai, Jiangshe Zhang, Yulun Zhang, Shuang Xu, Zudi Lin, Radu Timofte, and Luc Van Gool. Cddfuse: Correlation-driven dual-branch feature decomposition for multi-modality image fusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern r...

  25. [34]

    Probing synergistic high-order interaction in infrared and visible image fusion

    Naishan Zheng, Man Zhou, Jie Huang, Junming Hou, Haoying Li, Yuan Xu, and Feng Zhao. Probing synergistic high-order interaction in infrared and visible image fusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 26384--26395, 2024

  26. [35]

    Coconet: Coupled contrastive learning network with multi-level feature ensemble for multi-modality image fusion

    Jinyuan Liu, Runjia Lin, Guanyao Wu, Risheng Liu, Zhongxuan Luo, and Xin Fan. Coconet: Coupled contrastive learning network with multi-level feature ensemble for multi-modality image fusion. International Journal of Computer Vision, 132 0 (5): 0 1748--1775, 2024 e

  27. [36]

    Contrastive learning for unpaired image-to-image translation

    Taesung Park, Alexei A Efros, Richard Zhang, and Jun-Yan Zhu. Contrastive learning for unpaired image-to-image translation. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part IX 16, pages 319--345. Springer, 2020 a

  28. [37]

    The pascal visual object classes challenge: A retrospective

    Mark Everingham, SM Ali Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes challenge: A retrospective. International journal of computer vision, 111: 0 98--136, 2015

  29. [38]

    Assessment of image fusion procedures using entropy, image quality, and multispectral classification

    J Wesley Roberts, Jan A Van Aardt, and Fethi Babikker Ahmed. Assessment of image fusion procedures using entropy, image quality, and multispectral classification. Journal of Applied Remote Sensing, 2 0 (1): 0 023522, 2008

  30. [39]

    Detail preserved fusion of visible and infrared images using regional saliency extraction and multi-scale image decomposition

    Guangmang Cui, Huajun Feng, Zhihai Xu, Qi Li, and Yueting Chen. Detail preserved fusion of visible and infrared images using regional saliency extraction and multi-scale image decomposition. Optics Communications, 341: 0 199--209, 2015

  31. [40]

    Fsim: A feature similarity index for image quality assessment

    Lin Zhang, Lei Zhang, Xuanqin Mou, and David Zhang. Fsim: A feature similarity index for image quality assessment. IEEE transactions on Image Processing, 20 0 (8): 0 2378--2386, 2011

  32. [41]

    Image fusion metric based on mutual information and tsallis entropy

    N Cvejic, CN Canagarajah, and DR Bull. Image fusion metric based on mutual information and tsallis entropy. Electronics letters, 42 0 (11): 0 626--627, 2006

  33. [42]

    Image quality measures and their performance

    Ahmet M Eskicioglu and Paul S Fisher. Image quality measures and their performance. IEEE Transactions on communications, 43 0 (12): 0 2959--2965, 2002

  34. [43]

    Contrastive learning for unpaired image-to-image translation

    Taesung Park, Alexei A Efros, Richard Zhang, and Jun-Yan Zhu. Contrastive learning for unpaired image-to-image translation. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part IX 16, pages 319--345. Springer, 2020 b

  35. [44]

    Doubao vision pro 32k

    Doubao. Doubao vision pro 32k. https://console.volcengine.com/ark/region:ark+cn-beijing/model/detail?Id=doubao-vision-pro-32k

  36. [45]

    Fastinst: A simple query-based model for real-time instance segmentation

    Junjie He, Pengyu Li, Yifeng Geng, and Xuansong Xie. Fastinst: A simple query-based model for real-time instance segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 23663--23672, 2023

  37. [46]

    Cascade r-cnn: Delving into high quality object detection

    Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6154--6162, 2018

  38. [47]

    Mask dino: Towards a unified transformer-based framework for object detection and segmentation

    Feng Li, Hao Zhang, Huaizhe Xu, Shilong Liu, Lei Zhang, Lionel M Ni, and Heung-Yeung Shum. Mask dino: Towards a unified transformer-based framework for object detection and segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, page...

  39. [48]

    Yolov11: An overview of the key architectural enhancements

    Rahima Khanam and Muhammad Hussain. Yolov11: An overview of the key architectural enhancements. arXiv preprint arXiv:2410.17725, 2024 b

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.