REVIEW 4 major objections 6 minor 79 references
M2I2HA models RGB-thermal detection with hypergraph attention, reporting top accuracy on DroneVehicle and FLIR while keeping inference real-time.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 09:05 UTC pith:EZBMFK2T
load-bearing objection Solid engineering package, but the SOTA claim collapses on the paper's own tables (LLVIP, VEDAI), leaving sub-point margins over COMO on DroneVehicle/FLIR without uncertainty. the 4 major comments →
M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper proposes M2I2HA, a YOLO-based multi-modal detector that represents image features as vertices of a hypergraph, where a single edge can link many nodes. An Intra-Hypergraph Enhancement module constructs high-order hyperedges inside each modality to capture global many-to-many dependencies; an Inter-Hypergraph Fusion module builds cross-modal hyperedges that connect RGB and thermal nodes, aligning the two feature spaces and fusing them through gated attention. A third component, M3FDFP, adaptively merges features from multiple levels to counter the information bottleneck. On four public benchmarks (DroneVehicle, FLIR-aligned, LLVIP, VEDAI), M2I2HA reports the best overall mAP on Dron
What carries the argument
A hypergraph H=(V,E) in which hyperedges connect more than two vertices, allowing direct modeling of high-order many-to-many associations. The paper builds on this by adding low-rank decomposition and Top-K sparsification to the intra-modal hypergraph attention (LR-C3AH), and a cross-hypergraph update scheme (CHEG, CHNN, GateFusion) for inter-modal alignment. The M3FDFP block then channels enhanced features into the neck with learnable per-modality weights.
Load-bearing premise
The from-scratch training protocol assumes that initializing all compared models without pre-trained weights penalizes them equally, so that the reported margins over baselines reflect the fusion architecture rather than the comparison setup.
What would settle it
Re-run the four benchmarks with each baseline using its officially pre-trained backbone and standard fine-tuning; if COMO or CFT then matches or exceeds M2I2HA's reported mAP on DroneVehicle or FLIR, the claimed advantage would be an artifact of the from-scratch protocol rather than the hypergraph design.
If this is right
- On DroneVehicle, the YOLOv8s-based M2I2HA reports mAP 63.4% and mAP@.75 75.7%, edging the Mamba-based COMO (63.3%, 75.1%) and improving small-object and localization metrics.
- The YOLOv8n variant reaches 237.4 Hz on the DroneVehicle test set, demonstrating that high-order hypergraph attention can be deployed in real-time pipelines.
- Ablations show that the Inter-Hypergraph Fusion module contributes more to mAP than the Intra-Hypergraph Enhancement module, and that both together outperform switching on either alone.
- On the small VEDAI dataset, M2I2HA comes close to SuperYOLO's 75.4% mAP@.5 without using a super-resolution auxiliary task, indicating the fusion design can handle tiny targets.
Where Pith is reading between the lines
- The hypergraph construction is not tied to a specific sensor pair, so the same intra/inter modules should transfer to depth, point clouds, or optical flow by swapping the second feature encoder; the authors list this as future work.
- The paper's ablations toggle entire modules but do not compare against a pairwise-attention variant of the same block; replacing hyperedges with standard attention edges would isolate the value of high-order connectivity.
- Converged values of the learnable scalars α, β, γ in M3FDFP would reveal whether the network leans on inter-modal features, intra-modal enhancement, or raw features in different scenarios, a diagnostic the paper does not provide.
- If the from-scratch protocol understates transformer baselines, the real-world gap between methods may be smaller or larger; an independent benchmark with official pre-trained weights would settle whether hypergraph fusion is the deciding factor.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes M2I2HA, a YOLO-based RGB-thermal fusion detector. Two parallel CNN branches extract features; an Intra-Hypergraph Enhancement module (FuseSEBlock, LR-C3AH low-rank sparse hypergraph attention, DSC3k) models high-order intra-modal dependencies; an Inter-Hypergraph Fusion module aligns and fuses cross-modal high-level features; and an M3FDFP block adaptively redistributes multi-level features. The method is evaluated on DroneVehicle, FLIR-aligned, LLVIP, and VEDAI against SuperYOLO, ICAFusion, CFT, GHOST, GM-DETR, and COMO, with ablations on DroneVehicle. The authors claim state-of-the-art performance on multiple public datasets.
Significance. The high-order hypergraph mechanism is a plausible and timely alternative to pairwise attention and sequential SSM scanning, and the architecture is efficient (about 225 Hz for the YOLOv8s variant). The paper is transparent about the from-scratch protocol, reports computational overhead, and includes ablations. However, the SOTA claim is not established: the paper's own Tables 4–5 show other methods leading on two of the four benchmarks, and the margins over the closest competitor on the other two are small (0.1–0.8 mAP) with no multi-seed statistics. The core architectural direction is worthy of further work, but the current claims outrun the evidence.
major comments (4)
- [Abstract and Tables 4–5] The unqualified SOTA claim is contradicted by the paper's own data. On VEDAI (Table 5), SuperYOLO is clearly ahead (75.4/47.5 vs 73.5/43.6) and COMO has a higher overall mAP (44.2 vs 43.6). On LLVIP (Table 4), CFT has higher mAP (60.0 vs 59.7/58.8); the text itself says Ours-n is 'second-best'. The only clear wins are DroneVehicle and FLIR, with margins over COMO of 0.1–0.8 mAP and no variance reported. The abstract and introduction should either claim SOTA only on the datasets/metrics where it actually holds or the method must be improved.
- [Sec. 4.1, second paragraph] All baselines are retrained from scratch without pretrained weights, and the paper concedes this can make some comparative methods 'slightly lower than their officially reported figures.' For CFT, GM-DETR, and COMO, whose published pipelines rely on pretrained backbones and carefully tuned schedules, this protocol can plausibly handicap them more than M2I2HA, which shares a YOLOv8s backbone with COMO. Since the reported wins over COMO are only 0.1–0.8 mAP, this is load-bearing. Please add comparisons under official/pretrained conditions, or clearly argue and demonstrate that the from-scratch protocol preserves relative rankings, and report multiple seeds/confidence intervals.
- [Sec. 3.2 / 3.4 / 3.5 (C3,C4,C5)] The Inter-Hypergraph Fusion module is described in Eq. (9) as taking only high-level features H^rgb_5 and H^ir_5 and outputting Y, yet Sec. 3.2 says it yields {C3,C4,C5} and Sec. 3.5 uses C_i for i=3,4,5. The mechanism producing lower-level cross-modal outputs is not specified. This leaves the fusion path underspecified and is a reproducibility gap.
- [Introduction vs Table 2] The introduction states that M2I2HA improves accuracy 'without increasing model parameters or sacrificing computational efficiency and inference speed.' Table 2 shows the opposite relative to COMO: 37.6M vs 24.3M parameters, 68.7 vs 47.0 GFLOPs, and 225.0 vs 226.5 Hz for the YOLOv8s variants. This claim should be removed or qualified to the specific baseline against which it is true.
minor comments (6)
- [Sec. 3.4 title and Sec. 4.7] 'Intre-Hypergraph' is a typo for 'Inter-Hypergraph' in the section title and in Sec. 4.7.
- [Tables 3 and 4] The arrows for mAP@.75 and mAP in Tables 3 and 4 are shown as downward arrows (↓), but higher values are better; they should be upward arrows (↑).
- [Sec. 4.5 vs Table 5] The text says 'second-best mAP@0.5 with 74.5%' but Table 5 lists 73.5% for Ours (YOLOv8s).
- [Abstract vs Sec. 3.5] The abstract and highlights call the module 'M2-FullPAD' while Sec. 3.5 uses 'M3FDFP'; unify the naming.
- [Eq. (4)] In the FuseSEBlock equations, F is used in the AvgPool2d line after F' is defined; clarify whether F' should appear there.
- [Sec. 3.3 (DSC3k)] The DSC3k block is said to be 'retained from original implementation' but no citation or definition is provided; please specify its source.
Circularity Check
No significant circularity: performance is measured on external public benchmarks, modules are defined by explicit trainable operations, and no prediction reduces to a fitted value or a self-citation chain.
full rationale
M2I2HA is an empirically evaluated architecture paper rather than a derivation from first principles. The claimed results are benchmark measurements on external public datasets (DroneVehicle, FLIR, LLVIP, VEDAI) against published baselines, so there is no fitted parameter being renamed as a prediction. The hypergraph modules are specified by explicit equations (Eqs. 4–13) with ordinary trainable weights (U, V_dyn, W, alpha, beta, gamma); the outputs are architectural feature transformations, not quantities defined by the target mAP numbers. The DPI/mutual-information paragraph in Section 3.5 is motivational and does not derive the M3FDFP block. The paper builds on external prior work (YOLOv13's HyperACE, COMO, CFT, etc.) and does not invoke a load-bearing self-citation; the single overlapping-author citation [18] is only a general remark about CNN limitations. Concerns about the from-scratch training protocol and the fact that some baselines are not state-of-the-art on every dataset are comparison-fairness and evidence-quality issues, not circularity of the paper's derivation. Therefore no specific circular step can be exhibited.
Axiom & Free-Parameter Ledger
free parameters (4)
- Fusion weights α, β, γ (M3FDFP) =
Not reported (learned during training)
- Sparsity ratio γ (LR-C3AH) =
Not reported
- Low-rank dimension r (hyperedge prototypes) =
Not reported
- Top-K hyperedge count =
Not reported
axioms (5)
- domain assumption Data Processing Inequality implies deeper network layers contain strictly less input information, and dense re-introduction of earlier features mitigates this bottleneck
- domain assumption High-level features (P5) are sufficient for cross-modal fusion because lower levels are more affected by spatial misalignment
- domain assumption Training all methods from scratch without pretrained weights is a fair and informative comparison protocol
- domain assumption Softmax attention over learned hyperedge prototype vectors (Eqs. 6-8, 11-13) constitutes genuine high-order hypergraph modeling rather than an instance of pairwise attention
- standard math Hypergraph incidence matrix and node/hyperedge degree definitions (Eqs. 1-3) correctly describe the data structures used
read the original abstract
Recent advances in multi-modal detection have significantly improved detection accuracy in challenging environments (e.g., low light, overexposure). By integrating RGB with modalities such as thermal and depth, multi-modal fusion increases data redundancy and system robustness. However, significant challenges remain in effectively extracting task-relevant information both within and across modalities, as well as in achieving precise cross-modal alignment. While CNNs excel at feature extraction, they are limited by constrained receptive fields, strong inductive biases, and difficulty in capturing long-range dependencies. Transformer-based models offer global context but suffer from quadratic computational complexity and are confined to pairwise correlation modeling. Mamba and other State Space Models (SSMs), on the other hand, are hindered by their sequential scanning mechanism, which flattens 2D spatial structures into 1D sequences, disrupting topological relationships and limiting the modeling of complex higher-order dependencies. To address these issues, we propose a multi-modal perception network based on hypergraph theory called M2I2HA. Our architecture includes an Intra-Hypergraph Enhancement module to capture global many-to-many high-order relationships within each modality, and an Inter-Hypergraph Fusion module to align, enhance, and fuse cross-modal features by bridging configuration and spatial gaps between data sources. We further introduce a M2-FullPAD module to enable adaptive multi-level fusion of multi-modal enhanced features within the network, meanwhile enhancing data distribution and flow across the architecture. Extensive object detection experiments on multiple public datasets against baselines demonstrate that M2I2HA achieves state-of-the-art performance in multi-modal object detection tasks.
Figures
Reference graph
Works this paper leans on
-
[1]
J. Redmon, A. Farhadi, Yolov3: An incremental improvement, arXiv preprint arXiv:1804.02767 (2018)
Pith/arXiv arXiv 2018
-
[2]
Varghese, M
R. Varghese, M. Sambath, Yolov8: A novel object detection algorithm with enhanced performance and robustness, in: 2024 International con- ference on advances in data engineering and intelligent computing sys- tems (ADICS), IEEE, 2024, pp. 1–6
2024
-
[3]
Y. Tian, Q. Ye, D. Doermann, Yolov12: Attention-centric real-time object detectors, arXiv preprint arXiv:2502.12524 (2025)
Pith/arXiv arXiv 2025
-
[4]
M. Lei, S. Li, Y. Wu, H. Hu, Y. Zhou, X. Zheng, G. Ding, S. Du, Z. Wu, Y. Gao, Yolov13: Real-time object detection with hypergraph-enhanced adaptive visual perception, arXiv preprint arXiv:2506.17733 (2025)
Pith/arXiv arXiv 2025
-
[5]
R. Sapkota, R. H. Cheppally, A. Sharda, M. Karkee, Yolo26: Key ar- chitectural enhancements and performance benchmarking for real-time object detection, arXiv preprint arXiv:2509.25164 (2025)
arXiv 2025
-
[6]
K. He, G. Gkioxari, P. Dollár, R. Girshick, Mask r-cnn, in: Proceedings oftheIEEEinternationalconferenceoncomputervision, 2017, pp.2961– 2969
2017
-
[7]
S. Ren, K. He, R. Girshick, J. Sun, Faster r-cnn: Towards real-time object detection with region proposal networks, Advances in neural in- formation processing systems 28 (2015)
2015
-
[8]
C. Li, D. Song, R. Tong, M. Tang, Illumination-aware faster r-cnn for ro- bust multispectral pedestrian detection, Pattern Recognition 85 (2019) 161–171. 35
2019
-
[9]
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, J. Dai, Deformable detr: De- formable transformers for end-to-end object detection, arXiv preprint arXiv:2010.04159 (2020)
Pith/arXiv arXiv 2010
-
[10]
X. Dai, Y. Chen, J. Yang, P. Zhang, L. Yuan, L. Zhang, Dynamic detr: End-to-end object detection with dynamic attention, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 2988–2997
2021
-
[11]
Y. Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, Y. Liu, J. Chen, Detrs beat yolos on real-time object detection, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 16965–16974
2024
-
[12]
W. Lv, Y. Zhao, Q. Chang, K. Huang, G. Wang, Y. Liu, Rt-detrv2: Im- proved baseline with bag-of-freebies for real-time detection transformer, arXiv preprint arXiv:2407.17140 (2024)
Pith/arXiv arXiv 2024
-
[13]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transfer- able visual models from natural language supervision, in: International conference on machine learning, PmLR, 2021, pp. 8748–8763
2021
-
[14]
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggar- wal, O. K. Mohammed, S. Singhal, S. Som, et al., Image as a foreign language: Beit pretraining for vision and vision-language tasks, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 19175–19186
2023
-
[15]
Y. Chen, H. Xie, H. Shin, Multi-layer fusion techniques using a cnn for multispectral pedestrian detection, IET Computer Vision 12 (8) (2018) 1179–1187
2018
-
[16]
Zhang, J
J. Zhang, J. Lei, W. Xie, Z. Fang, Y. Li, Q. Du, Superyolo: Super reso- lution assisted object detection in multimodal remote sensing imagery, IEEE Transactions on Geoscience and Remote Sensing 61 (2023) 1–15
2023
-
[17]
Kalamkar, G
S. Kalamkar, G. M. Amalanathan, Mda-vit: Multimodal image fusion using dual attention vision transformer, Multimedia Tools and Applica- tions 84 (21) (2025) 23701–23723. 36
2025
-
[18]
W. Pan, J. Shen, B. Wang, S. Wang, Z. Sun, Open-set recognition based on the combination of deep learning and hypothesis testing for detecting unknown nuclear faults, Nuclear Engineering and Design 429 (2024) 113654
2024
-
[19]
A. Dosovitskiy, An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020)
Pith/arXiv arXiv 2010
-
[20]
Touvron, M
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, H. Jégou, Training data-efficient image transformers & distillation through atten- tion, in: International conference on machine learning, PMLR, 2021, pp. 10347–10357
2021
-
[21]
Chandrasiri, P
M. Chandrasiri, P. D. Talagala, Cross-vit: Cross-attention vision trans- former for image duplicate detection, in: 2023 8th International Con- ference on Information Technology Research (ICITR), IEEE, 2023, pp. 1–6
2023
-
[22]
C. Liu, X. Ma, X. Yang, Y. Zhang, Y. Dong, Como: Cross-mamba interaction and offset-guided fusion for multimodal object detection, In- formation Fusion (2025) 103414
2025
-
[23]
A. Gu, T. Dao, Mamba: Linear-time sequence modeling with selective state spaces, in: First conference on language modeling, 2024
2024
-
[24]
F.Gao, X.Jin, X.Zhou, J.Dong, Q.Du, Msfmamba: Multi-scalefeature fusion state space model for multi-source remote sensing image classifi- cation, IEEE Transactions on Geoscience and Remote Sensing (2025)
2025
-
[25]
Z. Wan, P. Zhang, Y. Wang, S. Yong, S. Stepputtis, K. Sycara, Y. Xie, Sigma: Siamese mamba network for multi-modal semantic segmenta- tion, in: 2025 IEEE/CVF Winter Conference on Applications of Com- puter Vision (WACV), IEEE, 2025, pp. 1734–1744
2025
-
[26]
L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, X. Wang, Vision mamba: Efficient visual representation learning with bidirectional state space model, arXiv preprint arXiv:2401.09417 (2024)
Pith/arXiv arXiv 2024
-
[27]
J. Qu, J. Liu, C. Yu, Adaptive multi-scale hypernet with bi-direction residual attention module for scene text detection, Journal of Informa- tion Hiding and Privacy Protection 3 (2) (2021) 83. 37
2021
-
[28]
Ferens, Y
R. Ferens, Y. Keller, Hyperpose: Hypernetwork-infused camera pose localization and an extended cambridge landmarks dataset, in: Pro- ceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 11547–11557
2025
-
[29]
X. Li, Y. Li, C. Shen, A. Dick, A. Van Den Hengel, Contextual hyper- graph modeling for salient object detection, in: Proceedings of the IEEE international conference on computer vision, 2013, pp. 3328–3335
2013
-
[30]
Y. Feng, J. Huang, S. Du, S. Ying, J.-H. Yong, Y. Li, G. Ding, R. Ji, Y. Gao, Hyper-yolo: When visual object detection meets hypergraph computation, IEEE Transactions on Pattern Analysis and Machine In- telligence (2024)
2024
-
[31]
A. A. Aguileta, R. F. Brena, O. Mayora, E. Molino-Minero-Re, L. A. Trejo, Multi-sensor fusion for activity recognition—a survey, Sensors 19 (17) (2019) 3808
2019
-
[32]
A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al., Deepseek-v3 technical report, arXiv preprint arXiv:2412.19437 (2024)
Pith/arXiv arXiv 2024
-
[33]
X. Wang, J. Wu, P. Zhang, Z. Yu, Language-driven cross-attention for visible–infrared image fusion using clip, Sensors 25 (16) (2025) 5083
2025
-
[34]
Gohil, S
P. Gohil, S. Thoduka, P. G. Plöger, Sensor fusion and multimodal learn- ing for robotic grasp verification using neural networks, in: 2022 26th International Conference on Pattern Recognition (ICPR), IEEE, 2022, pp. 5111–5117
2022
-
[35]
Ramachandram, G
D. Ramachandram, G. W. Taylor, Deep multimodal learning: A survey on recent advances and trends, IEEE signal processing magazine 34 (6) (2017) 96–108
2017
-
[36]
C. He, Q. Liu, H. Li, H. Wang, Multimodal medical image fusion based on ihs and pca, Procedia Engineering 7 (2010) 280–285
2010
-
[37]
L. A. Maglanoc, T. Kaufmann, R. Jonassen, E. Hilland, D. Beck, N. I. Landrø, L. T. Westlye, Multimodal fusion of structural and functional brain imaging in depression using linked independent component anal- ysis, Human brain mapping 41 (1) (2020) 241–255. 38
2020
-
[38]
B. Lei, S. Chen, D. Ni, T. Wang, Discriminative learning for alzheimer’s disease diagnosis via canonical correlation analysis and multimodal fu- sion, Frontiers in aging neuroscience 8 (2016) 77
2016
-
[39]
F. Ma, X. Xu, S.-L. Huang, L. Zhang, Maximum likelihood estima- tion for multimodal learning with missing modality, arXiv preprint arXiv:2108.10513 (2021)
Pith/arXiv arXiv 2021
-
[40]
Singhal, C
A. Singhal, C. R. Brown, Dynamic bayes net approach to multimodal sensor fusion, in: Sensor Fusion and Decentralized Control in Au- tonomous Robotic Systems, Vol. 3209, SPIE, 1997, pp. 2–10
1997
-
[41]
L. D. Phi, B. P. N. Thanh, Q. T. Van, A classification method based on concatenation features for diagnosing skin diseases, Journal of Informa- tion Hiding and Multimedia Signal Processing 16 (1) (2025) 401–413
2025
-
[42]
T. Jain, D. Gopalani, Y. Kumar Meena, Informative task classification with concatenated embeddings using deep learning on crisismmd, Inter- national Journal of Computers and Applications 47 (2) (2025) 123–140
2025
- [43]
-
[44]
F. Yuan, Y. Chen, W. He, J. Zeng, Feature fusion-guided network with sparse prior constraints for unsupervised hyperspectral image quality improvement, IEEE Transactions on Geoscience and Remote Sensing (2025)
2025
-
[45]
H. Zhao, W. Li, D. Huang, J. Huang, L. Zhang, M-gan: multiattribute learning and multimodal feature fusion-based generative adversarial net- work for text-to-image synthesis, The Visual Computer 41 (5) (2025) 3017–3035
2025
-
[46]
J. Peng, K. Lv, G. Wang, W. Xiao, T. Ran, L. Yuan, Mlsa-yolo: A multi-level feature fusion and scale-adaptive framework for small object detection, The Journal of Supercomputing 81 (4) (2025) 528
2025
-
[47]
Ozdemir, I
B. Ozdemir, I. Pacal, An innovative deep learning framework for skin cancer detection employing convnextv2 and focal self-attention mecha- nisms, Results in Engineering 25 (2025) 103692. 39
2025
-
[48]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, Advances in neural information processing systems 30 (2017)
2017
-
[49]
Yi, K.-C
M.-H. Yi, K.-C. Kwak, J.-H. Shin, Hyfuser: hybrid multimodal trans- formerforemotionrecognitionusingdualcrossmodalattention, Applied Sciences 15 (3) (2025) 1053
2025
-
[50]
Y. Guo, F. Wang, H. Chu, S. Wen, Cross-modal attention and geo- metric contextual aggregation network for 6dof object pose estimation, Neurocomputing 617 (2025) 128891
2025
-
[51]
X. Song, X. Zhang, J. Ji, Y. Liu, P. Wei, Cross-modal contrastive atten- tion model for medical report generation, in: Proceedings of the 29th international conference on computational linguistics, 2022, pp. 2388– 2397
2022
-
[52]
R. G. Praveen, J. Alam, Recursive joint cross-modal attention for multi- modal fusion in dimensional emotion recognition, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 4803–4813
2024
-
[53]
W. Zhou, S. Dong, M. Fang, L. Yu, Cacfnet: Cross-modal attention cas- caded fusion network for rgb-t urban scene parsing, IEEE Transactions on Intelligent Vehicles 9 (1) (2023) 1919–1929
2023
-
[54]
X. Xu, J. Chen, D. Thakur, D. Hong, Multi-modal disease segmentation with continual learning and adaptive decision fusion, Information Fusion 118 (2025) 102962
2025
-
[55]
Zhang, J
B. Zhang, J. Ma, X. Fu, G. Dai, Logic augmented multi-decision fu- sion framework for stance detection on social media, Information Fusion (2025) 103214
2025
-
[56]
C. Fu, F. Qian, K. Su, Y. Su, Z. Wang, J. Shi, Z. Liu, C. Liu, C. T. Ishi, Himul-lgg: A hierarchical decision fusion-based local–global graph neu- ral network for multimodal emotion recognition in conversation, Neural Networks 181 (2025) 106764
2025
-
[57]
M. Han, B. Shen, J. Ruan, Multi-modal data fusion for 3d object detec- tion using dual-attention mechanism, Sensors 25 (20) (2025) 6360. 40
2025
-
[58]
S. Xu, D. Zhou, J. Fang, J. Yin, Z. Bin, L. Zhang, Fusionpainting: Multimodal fusion with adaptive attention for 3d object detection, in: 2021 IEEE International Intelligent Transportation Systems Conference (ITSC), IEEE, 2021, pp. 3047–3054
2021
-
[59]
X. Chen, H. Ma, J. Wan, B. Li, T. Xia, Multi-view 3d object detection network for autonomous driving, in: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2017, pp. 1907–1915
2017
-
[60]
Y. Li, A. W. Yu, T. Meng, B. Caine, J. Ngiam, D. Peng, J. Shen, Y. Lu, D. Zhou, Q. V. Le, et al., Deepfusion: Lidar-camera deep fusion for multi-modal 3d object detection, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 17182– 17191
2022
-
[61]
Y. Wang, W. Huang, F. Sun, T. Xu, Y. Rong, J. Huang, Deep mul- timodal fusion by channel exchanging, Advances in neural information processing systems 33 (2020) 4835–4845
2020
-
[62]
F. Qingyun, H. Dapeng, W. Zhaokui, Cross-modality fusion transformer for multispectral object detection, arXiv preprint arXiv:2111.00273 (2021)
Pith/arXiv arXiv 2021
-
[63]
Ikram, I
S. Ikram, I. Sarwar, A. Ikram, M. Abdullah-AI-Wahud, et al., A transformer-basedmultimodalobjectdetectionsystemforreal-worldap- plications, IEEE Access (2025)
2025
-
[64]
C. Meng, H. Motevalli, Link prediction in social networks using hyper- motif representation on hypergraph, Multimedia Systems 30 (3) (2024) 123
2024
-
[65]
W. Ye, Q. Qian, Forecasting the molecular interactions: A hypergraph- based neural network for molecular relational learning, Knowledge- Based Systems 300 (2024) 112177
2024
-
[66]
La Gatta, V
V. La Gatta, V. Moscato, M. Pennone, M. Postiglione, G. Sperlí, Mu- sic recommendation via hypergraph embedding, IEEE transactions on neural networks and learning systems 34 (10) (2022) 7887–7899. 41
2022
-
[67]
D. Sakong, V. H. Vu, T. T. Huynh, P. L. Nguyen, H. Yin, Q. V. H. Nguyen, T. T. Nguyen, Heterogeneous hypergraph embedding for rec- ommendation systems, arXiv preprint arXiv:2407.03665 (2024)
Pith/arXiv arXiv 2024
-
[68]
R. Zhang, Y. Zou, J. Ma, Hyper-sagnn: a self-attention based graph neuralnetworkforhypergraphs, arXivpreprintarXiv:1911.02613(2019)
Pith/arXiv arXiv 1911
-
[69]
S. Bai, F. Zhang, P. H. Torr, Hypergraph convolution and hypergraph attention, Pattern Recognition 110 (2021) 107637
2021
-
[70]
S. Sinha, H. Bharadhwaj, A. Srinivas, A. Garg, D2rl: Deep dense ar- chitectures in reinforcement learning, arXiv preprint arXiv:2010.09163 (2020)
Pith/arXiv arXiv 2010
-
[71]
Huang, Z
G. Huang, Z. Liu, L. Van Der Maaten, K. Q. Weinberger, Densely con- nected convolutional networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 4700–4708
2017
-
[72]
Y. Sun, B. Cao, P. Zhu, Q. Hu, Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning, IEEE Transactions on Circuits and Systems for Video Technology 32 (10) (2022) 6700–6713
2022
-
[73]
Zhang, E
H. Zhang, E. Fromont, S. Lefevre, B. Avignon, Multispectral fusion for object detection with cyclic fuse-and-refine blocks, in: 2020 IEEE International conference on image processing (ICIP), IEEE, 2020, pp. 276–280
2020
-
[74]
X. Jia, C. Zhu, M. Li, W. Tang, W. Zhou, Llvip: A visible-infrared paired dataset for low-light vision, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 3496–3504
2021
-
[75]
Razakarivony, F
S. Razakarivony, F. Jurie, Vehicle detection in aerial imagery: A small target detection benchmark, Journal of Visual Communication and Im- age Representation 34 (2016) 187–203
2016
-
[76]
J. Shen, Y. Chen, Y. Liu, X. Zuo, H. Fan, W. Yang, Icafusion: Iterative cross-attention guided feature fusion for multispectral object detection, Pattern Recognition 145 (2024) 109913
2024
-
[77]
Zhang, J
J. Zhang, J. Lei, W. Xie, Y. Li, G. Yang, X. Jia, Guided hybrid quan- tization for object detection in remote sensing imagery via one-to-one 42 self-teaching, IEEE transactions on geoscience and remote sensing 61 (2023) 1–15
2023
-
[78]
Y. Xiao, F. Meng, Q. Wu, L. Xu, M. He, H. Li, Gm-detr: General- izedmuiltispectraldetectiontransformerwithefficientfusionencoderfor visible-infrared detection, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5541–5549
2024
-
[79]
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, C. L. Zitnick, Microsoft coco: Common objects in context, in: European conference on computer vision, Springer, 2014, pp. 740–755. 43
2014
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.