REVIEW 11 cited by
Cross-Modality Fusion Transformer for Multispectral Object Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Cross-Modality Fusion Transformer for Multispectral Object Detection
read the original abstract
Multispectral image pairs can provide the combined information, making object detection applications more reliable and robust in the open world. To fully exploit the different modalities, we present a simple yet effective cross-modality feature fusion approach, named Cross-Modality Fusion Transformer (CFT) in this paper. Unlike prior CNNs-based works, guided by the transformer scheme, our network learns long-range dependencies and integrates global contextual information in the feature extraction stage. More importantly, by leveraging the self attention of the transformer, the network can naturally carry out simultaneous intra-modality and inter-modality fusion, and robustly capture the latent interactions between RGB and Thermal domains, thereby significantly improving the performance of multispectral object detection. Extensive experiments and ablation studies on multiple datasets demonstrate that our approach is effective and achieves state-of-the-art detection performance. Our code and models are available at https://github.com/DocF/multispectral-object-detection.
Forward citations
Cited by 11 Pith papers
-
WD-FQDet: Multispectral Detection Transformer via Wavelet Decomposition and Frequency-aware Query Learning
WD-FQDet decouples modality-shared and modality-specific features in infrared-visible images via wavelet-based frequency decomposition and frequency-aware query selection to achieve state-of-the-art detection performance.
-
SMART-Ship: A Comprehensive Synchronized Multi-modal Aligned Remote Sensing Targets Dataset and Benchmark for Berthed Ships Analysis
SMART-Ship introduces a new synchronized multi-modal remote sensing dataset with fine-grained annotations for berthed ships and benchmarks for five interpretation tasks.
-
Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark
DHNet with patch alignment and dual hypergraph fusion reaches SOTA RGBT video object detection on VT-VOD50 and the new large-scale DVT-VOD1000 benchmark.
-
InfraNet: Quality-Aware RGB Guidance for Efficient Infrared Object Detection
QualGate-regulated RGB guidance during training produces efficient IR-only and dual-modal detectors that match or beat equal-fusion baselines under low light and adverse weather.
-
FreqKD: Frequency-Decoupled Cross-Modal Knowledge Distillation for Infrared Object Detection
FreqKD uses strict MSE on low-frequency features and relaxed log-MSE (weight 0.1) on high-frequency features for RGB-to-IR distillation, reporting 2.4 mAP50 gain on KAIST pedestrian detection with transfer to other da...
-
LER-YOLO: Reliability-Aware Expert Routing for Misaligned RGB-Infrared UAV Detection
LER-YOLO reports 89.7% AP50 on the MBU benchmark for misaligned RGB-IR UAV detection by routing among RGB-dominant, IR-dominant, and fusion experts using a spatial reliability map.
-
Dual Sparse Aggregation Transformer for Multispectral Object Detection
DSAFormer applies spatial and channel sparse multi-head cross-attention plus a learnable fusion block to reduce redundant token interactions and improve multispectral object detection on four public datasets.
-
Efficient RGB-T Object Detection via Sparse Cross-Modality Fusion
A two-stage RGB-T detector performs lightweight modality-specific proposal generation followed by sparse fusion-based refinement to match accuracy of heavier models at lower parameter and compute cost.
-
Bridging the RGB-IR Gap: Consensus and Discrepancy Modeling for Text-Guided Multispectral Detection
A text-guided fusion method for RGB-IR object detection aligns modalities via semantic bridging and incorporates both consensus and discrepancy cues through dynamic recalibration.
-
COXNet: Cross-Layer Fusion with Adaptive Alignment and Scale Integration for RGBT Tiny Object Detection
COXNet proposes cross-layer fusion, dynamic alignment, and GeoShape-based label assignment for RGBT tiny object detection, reporting 3.32% mAP50 gain on the RGBTDronePerson dataset.
-
M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention
M2I2HA adds intra-modal and cross-modal hypergraph attention modules to a YOLO-style detector and reports the best average precision on DroneVehicle and FLIR, while on LLVIP and VEDAI prior methods score higher on the...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.