Pith. sign in

REVIEW 21 cited by

Sparse4D: Multi-view 3D Object Detection with Sparse Spatial-Temporal Fusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.10581 v2 pith:7OKT3673 submitted 2022-11-19 cs.CV

Sparse4D: Multi-view 3D Object Detection with Sparse Spatial-Temporal Fusion

classification cs.CV
keywords detectionmethodssparsefeaturesdifferentmulti-viewsparse4danchor
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Bird-eye-view (BEV) based methods have made great progress recently in multi-view 3D detection task. Comparing with BEV based methods, sparse based methods lag behind in performance, but still have lots of non-negligible merits. To push sparse 3D detection further, in this work, we introduce a novel method, named Sparse4D, which does the iterative refinement of anchor boxes via sparsely sampling and fusing spatial-temporal features. (1) Sparse 4D Sampling: for each 3D anchor, we assign multiple 4D keypoints, which are then projected to multi-view/scale/timestamp image features to sample corresponding features; (2) Hierarchy Feature Fusion: we hierarchically fuse sampled features of different view/scale, different timestamp and different keypoints to generate high-quality instance feature. In this way, Sparse4D can efficiently and effectively achieve 3D detection without relying on dense view transformation nor global attention, and is more friendly to edge devices deployment. Furthermore, we introduce an instance-level depth reweight module to alleviate the ill-posed issue in 3D-to-2D projection. In experiment, our method outperforms all sparse based methods and most BEV based methods on detection task in the nuScenes dataset.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Streaming Gaussian Encoding for 4D Panoptic Occupancy Tracking

    cs.CV 2026-06 unverdicted novelty 7.0

    Introduces a streaming Gaussian encoder maintaining persistent volumetric representations via ego-motion compensation and confidence-guided updates for improved 4D panoptic occupancy tracking from cameras.

  2. PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations

    cs.CV 2026-05 unverdicted novelty 7.0

    PointForward uses sparse world-space 3D queries and scene graphs to deliver consistent single-pass reconstruction of dynamic driving scenes via point-aligned representations.

  3. Efficient Multi-View 3D Object Detection by Dynamic Token Selection and Fine-Tuning

    cs.CV 2026-04 unverdicted novelty 7.0

    Dynamic token selection and training only 1.6 million parameters instead of over 300 million reduces computation by 48-55% and improves accuracy over prior state-of-the-art on the NuScenes dataset.

  4. UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving

    cs.CV 2026-06 unverdicted novelty 6.0

    UniTeD unifies perception and planning in autonomous driving via shared temporal diffusion with TTM and ARS modules, reporting SOTA results on benchmarks.

  5. VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving

    cs.CV 2026-06 unverdicted novelty 6.0

    VLGA introduces geometry as a fourth modality in VLA models via pointmap regression loss, reporting SOTA open-loop and closed-loop driving metrics on nuScenes and Bench2Drive.

  6. SimPB++: Simultaneously Detecting 2D and 3D Objects from Multiple Cameras

    cs.CV 2026-05 unverdicted novelty 6.0

    SimPB++ unifies multi-view 2D perspective and 3D BEV object detection in one model via an interactive hybrid decoder, reporting state-of-the-art results on nuScenes and long-range detection up to 150 m on Argoverse2.

  7. CAM3DNet: Comprehensively mining the multi-scale features for 3D Object Detection with Multi-View Cameras

    cs.CV 2026-04 unverdicted novelty 6.0

    CAM3DNet outperforms prior camera-based 3D detectors on nuScenes, Waymo and Argoverse by using three new modules to better mine multi-scale spatiotemporal features from 2D queries and pyramid maps.

  8. Kerr-Schild Double Copy of the Randall-Sundrum Black String

    hep-th 2026-04 unverdicted novelty 6.0

    Kerr-Schild double copy of the RS II black string produces a sourceless Maxwell single copy and a warp-induced massive scalar zeroth copy, with an alternative splitting giving inequivalent gauge and scalar fields.

  9. DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale

    cs.CV 2026-04 unverdicted novelty 6.0

    DVGT-2 is a streaming vision-geometry-action model that jointly reconstructs dense 3D geometry and plans trajectories online, achieving better reconstruction than prior batch methods while transferring directly to pla...

  10. AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

    cs.RO 2026-01 conditional novelty 6.0

    Conditioning speed planning on the predicted path and relabeling synthetic cut-ins yields SOTA Bench2Drive scores (DS 89.07, SR 73.18%).

  11. AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

    cs.RO 2026-01 unverdicted novelty 6.0

    A cascaded end-to-end driving model conditions longitudinal planning on the lateral path via anchor-based regression and path-conditioned 1D displacement prediction, achieving SOTA driving score of 89.07 and 73.18% su...

  12. SparseCoop: Cooperative Perception with Kinematic-Grounded Queries

    cs.CV 2025-12 conditional novelty 6.0

    SparseCoop delivers state-of-the-art 3D detection and tracking performance on V2X-Seq and Griffin datasets using only sparse kinematic queries instead of dense BEV features, with lower transmission cost and latency ro...

  13. BePo: Dual Representation for 3D Occupancy Prediction

    cs.CV 2025-06 unverdicted novelty 6.0

    BePo proposes a dual BEV and sparse-points representation with cross-attention fusion for more accurate and efficient 3D occupancy prediction on autonomous driving benchmarks.

  14. Can BEV Perception Gracefully Degrade under Sensor Failures?

    cs.CV 2026-05 unverdicted novelty 5.0

    Grace-BEV enables graceful degradation in BEV perception under sensor failures by using a TrustGate Router for modality trustworthiness and FailSafe Fusion Block for dynamic integration, with modality dropout training...

  15. Grounded 3D-Aware Spatial Vision-Language Modeling

    cs.CV 2026-05 unverdicted novelty 5.0

    GR3D is a VLM that combines explicit 2D, implicit 2D, and monocular 3D grounding mechanisms to improve performance on spatial understanding benchmarks.

  16. VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction

    cs.CV 2026-05 unverdicted novelty 5.0

    VGGT-Occ embeds geometric tokens via PA-DA and uses sequential coarse-to-fine gated fusion to reach 33.00% IoU and 21.08% mIoU on SurroundOcc-nuScenes while using only ~41M parameters in the occupancy head.

  17. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

    cs.CV 2026-04 unverdicted novelty 5.0

    SemLT3D introduces semantic-guided expert distillation with a language MoE module and CLIP projection to enrich features for long-tailed classes in camera-only 3D detection.

  18. Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors

    cs.CV 2026-04 conditional novelty 5.0

    SEPatch3D accelerates ViT-based 3D object detectors up to 57% faster than StreamPETR via dynamic patch sizing and cross-granularity enhancement while keeping comparable accuracy on nuScenes and Argoverse 2.

  19. Not All Agents Matter: From Global Attention Dilution to Risk-Prioritized Game Planning

    cs.CV 2026-04 unverdicted novelty 5.0

    GameAD models autonomous driving as a risk-prioritized game among agents via Risk-Aware Topology Anchoring, Minimax Risk-Aware Sparse Attention and related components, yielding safer trajectories than prior end-to-end...

  20. Kerr-Schild Double Copy of the Randall-Sundrum Black String

    hep-th 2026-04 unverdicted novelty 5.0

    Kerr-Schild double copy of the RSII black string gives a holographic-coordinate-independent sourceless single-copy gauge field and a zeroth copy with warp-induced mass m²=12/l², while an alternative split is inequivalent.

  21. SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model

    cs.CV 2025-11 unverdicted novelty 5.0

    A sparse transformer predicts multi-frame 3D occupancy from images without BEV or VAE tokenization and reports SOTA results on nuScenes for 1-3s forecasting under arbitrary trajectories.