Pith. sign in

REVIEW 3 cited by

PillarMamba: Learning Local-Global Context for Roadside Point Cloud via Hybrid State Space Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.05397 v1 pith:LZBXKMYC submitted 2025-05-08 cs.CV

classification cs.CV
keywords roadsidecloudpointcontextmodelperceptionspacestate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Serving the Intelligent Transport System (ITS) and Vehicle-to-Everything (V2X) tasks, roadside perception has received increasing attention in recent years, as it can extend the perception range of connected vehicles and improve traffic safety. However, roadside point cloud oriented 3D object detection has not been effectively explored. To some extent, the key to the performance of a point cloud detector lies in the receptive field of the network and the ability to effectively utilize the scene context. The recent emergence of Mamba, based on State Space Model (SSM), has shaken up the traditional convolution and transformers that have long been the foundational building blocks, due to its efficient global receptive field. In this work, we introduce Mamba to pillar-based roadside point cloud perception and propose a framework based on Cross-stage State-space Group (CSG), called PillarMamba. It enhances the expressiveness of the network and achieves efficient computation through cross-stage feature fusion. However, due to the limitations of scan directions, state space model faces local connection disrupted and historical relationship forgotten. To address this, we propose the Hybrid State-space Block (HSB) to obtain the local-global context of roadside point cloud. Specifically, it enhances neighborhood connections through local convolution and preserves historical memory through residual attention. The proposed method outperforms the state-of-the-art methods on the popular large scale roadside benchmark: DAIR-V2X-I. The code will be released soon.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Defer to Plan: Adaptive Multi-Agent Fusion for End-to-End V2X Driving

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Moving multi-agent fusion from perception to planning, via an autoregressive decoder with MoE tokenization, yields 79.72 driving score on V2Xverse vs CoDriving's 77.15.

  2. RoadMamba: A Dual Branch Visual State Space Model for Road Surface Classification

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A dual-branch state space model combining whole-image and windowed local scanning with attention fusion achieves 92.81% top-1 accuracy on the 27-class RSCD road surface dataset, ahead of the compared Mamba, Transforme...

  3. RoadFormer : Local-Global Feature Fusion for Road Surface Classification in Autonomous Driving

    cs.CV 2025-06 conditional novelty 4.0 of 10

    RoadFormer, a hybrid convolutional-transformer network with a foreground-background training module, reports top-1 accuracies of 92.52% and 96.50% on the RSCD pavement datasets.

Pith tools