REVIEW 3 cited by
PillarMamba: Learning Local-Global Context for Roadside Point Cloud via Hybrid State Space Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Serving the Intelligent Transport System (ITS) and Vehicle-to-Everything (V2X) tasks, roadside perception has received increasing attention in recent years, as it can extend the perception range of connected vehicles and improve traffic safety. However, roadside point cloud oriented 3D object detection has not been effectively explored. To some extent, the key to the performance of a point cloud detector lies in the receptive field of the network and the ability to effectively utilize the scene context. The recent emergence of Mamba, based on State Space Model (SSM), has shaken up the traditional convolution and transformers that have long been the foundational building blocks, due to its efficient global receptive field. In this work, we introduce Mamba to pillar-based roadside point cloud perception and propose a framework based on Cross-stage State-space Group (CSG), called PillarMamba. It enhances the expressiveness of the network and achieves efficient computation through cross-stage feature fusion. However, due to the limitations of scan directions, state space model faces local connection disrupted and historical relationship forgotten. To address this, we propose the Hybrid State-space Block (HSB) to obtain the local-global context of roadside point cloud. Specifically, it enhances neighborhood connections through local convolution and preserves historical memory through residual attention. The proposed method outperforms the state-of-the-art methods on the popular large scale roadside benchmark: DAIR-V2X-I. The code will be released soon.
Forward citations
Cited by 3 Pith papers
-
Defer to Plan: Adaptive Multi-Agent Fusion for End-to-End V2X Driving
Moving multi-agent fusion from perception to planning, via an autoregressive decoder with MoE tokenization, yields 79.72 driving score on V2Xverse vs CoDriving's 77.15.
-
RoadMamba: A Dual Branch Visual State Space Model for Road Surface Classification
A dual-branch state space model combining whole-image and windowed local scanning with attention fusion achieves 92.81% top-1 accuracy on the 27-class RSCD road surface dataset, ahead of the compared Mamba, Transforme...
-
RoadFormer : Local-Global Feature Fusion for Road Surface Classification in Autonomous Driving
RoadFormer, a hybrid convolutional-transformer network with a foreground-background training module, reports top-1 accuracies of 92.52% and 96.50% on the RSCD pavement datasets.
Discussion (0). Continue with ORCID to comment.