REVIEW 7 cited by
LR-FPN: Enhancing Remote Sensing Object Detection with Location Refined Feature Pyramid Network
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Remote sensing target detection aims to identify and locate critical targets within remote sensing images, finding extensive applications in agriculture and urban planning. Feature pyramid networks (FPNs) are commonly used to extract multi-scale features. However, existing FPNs often overlook extracting low-level positional information and fine-grained context interaction. To address this, we propose a novel location refined feature pyramid network (LR-FPN) to enhance the extraction of shallow positional information and facilitate fine-grained context interaction. The LR-FPN consists of two primary modules: the shallow position information extraction module (SPIEM) and the contextual interaction module (CIM). Specifically, SPIEM first maximizes the retention of solid location information of the target by simultaneously extracting positional and saliency information from the low-level feature map. Subsequently, CIM injects this robust location information into different layers of the original FPN through spatial and channel interaction, explicitly enhancing the object area. Moreover, in spatial interaction, we introduce a simple local and non-local interaction strategy to learn and retain the saliency information of the object. Lastly, the LR-FPN can be readily integrated into common object detection frameworks to improve performance significantly. Extensive experiments on two large-scale remote sensing datasets (i.e., DOTAV1.0 and HRSC2016) demonstrate that the proposed LR-FPN is superior to state-of-the-art object detection approaches. Our code and models will be publicly available.
Forward citations
Cited by 7 Pith papers
-
EPANet: Efficient Path Aggregation Network for Underwater Fish Detection
EPANet combines cross-scale feature pyramid connections and a diversified bottleneck block to improve underwater fish detection accuracy and speed.
-
Fab-ME: A Vision State-Space and Attention-Enhanced Framework for Fabric Defect Detection
Fab-ME modifies YOLOv8s with a VMamba-based state-space module in the neck and an enhanced channel attention module, reporting 59.4 percent mAP@0.5 on the Tianchi fabric defect dataset versus a 57.4 percent baseline.
-
BAFPN: Bi directional alignment of features to improve localization accuracy
A bidirectional feature alignment network improves oriented object detection accuracy on the DOTAv1.5 aerial benchmark.
-
CCi-YOLOv8n: Enhanced Fire Detection with CARAFE and Context-Guided Modules
A YOLOv8n variant combining CARAFE, Context-Guided Downsampling, and iRMB reports small accuracy gains over YOLOv8n on two fire-detection datasets.
-
YOLO-FDA: Integrating Hierarchical Attention and Detail Enhancement for Surface Defect Detection
A YOLOv5 variant with BiFPN, directional detail enhancement, and two attention fusion modules reports state-of-the-art mAP on GC10-DET and DAGM2007.
-
Cross-modal Context Fusion and Adaptive Graph Convolutional Network for Multimodal Conversational Emotion Recognition
A model combining co-attention transformers, a BiGRU, and graph convolution for speaker relationships reports higher accuracies on IEMOCAP and MELD, though the evaluation has significant gaps.
-
Dynamic Attention and Bi-directional Fusion for Safety Helmet Wearing Detection
A YOLOv8-based helmet detector combining an attention head, weighted bidirectional feature fusion, and Wise-IoU loss reports 1.7% mAP gain over YOLOv8 on the SHWD dataset.
Discussion (0). Continue with ORCID to comment.