REVIEW 21 cited by
Sparse4D: Multi-view 3D Object Detection with Sparse Spatial-Temporal Fusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Sparse4D: Multi-view 3D Object Detection with Sparse Spatial-Temporal Fusion
read the original abstract
Bird-eye-view (BEV) based methods have made great progress recently in multi-view 3D detection task. Comparing with BEV based methods, sparse based methods lag behind in performance, but still have lots of non-negligible merits. To push sparse 3D detection further, in this work, we introduce a novel method, named Sparse4D, which does the iterative refinement of anchor boxes via sparsely sampling and fusing spatial-temporal features. (1) Sparse 4D Sampling: for each 3D anchor, we assign multiple 4D keypoints, which are then projected to multi-view/scale/timestamp image features to sample corresponding features; (2) Hierarchy Feature Fusion: we hierarchically fuse sampled features of different view/scale, different timestamp and different keypoints to generate high-quality instance feature. In this way, Sparse4D can efficiently and effectively achieve 3D detection without relying on dense view transformation nor global attention, and is more friendly to edge devices deployment. Furthermore, we introduce an instance-level depth reweight module to alleviate the ill-posed issue in 3D-to-2D projection. In experiment, our method outperforms all sparse based methods and most BEV based methods on detection task in the nuScenes dataset.
Forward citations
Cited by 21 Pith papers
-
Streaming Gaussian Encoding for 4D Panoptic Occupancy Tracking
Introduces a streaming Gaussian encoder maintaining persistent volumetric representations via ego-motion compensation and confidence-guided updates for improved 4D panoptic occupancy tracking from cameras.
-
PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations
PointForward uses sparse world-space 3D queries and scene graphs to deliver consistent single-pass reconstruction of dynamic driving scenes via point-aligned representations.
-
Efficient Multi-View 3D Object Detection by Dynamic Token Selection and Fine-Tuning
Dynamic token selection and training only 1.6 million parameters instead of over 300 million reduces computation by 48-55% and improves accuracy over prior state-of-the-art on the NuScenes dataset.
-
UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving
UniTeD unifies perception and planning in autonomous driving via shared temporal diffusion with TTM and ARS modules, reporting SOTA results on benchmarks.
-
VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving
VLGA introduces geometry as a fourth modality in VLA models via pointmap regression loss, reporting SOTA open-loop and closed-loop driving metrics on nuScenes and Bench2Drive.
-
SimPB++: Simultaneously Detecting 2D and 3D Objects from Multiple Cameras
SimPB++ unifies multi-view 2D perspective and 3D BEV object detection in one model via an interactive hybrid decoder, reporting state-of-the-art results on nuScenes and long-range detection up to 150 m on Argoverse2.
-
CAM3DNet: Comprehensively mining the multi-scale features for 3D Object Detection with Multi-View Cameras
CAM3DNet outperforms prior camera-based 3D detectors on nuScenes, Waymo and Argoverse by using three new modules to better mine multi-scale spatiotemporal features from 2D queries and pyramid maps.
-
Kerr-Schild Double Copy of the Randall-Sundrum Black String
Kerr-Schild double copy of the RS II black string produces a sourceless Maxwell single copy and a warp-induced massive scalar zeroth copy, with an alternative splitting giving inequivalent gauge and scalar fields.
-
DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale
DVGT-2 is a streaming vision-geometry-action model that jointly reconstructs dense 3D geometry and plans trajectories online, achieving better reconstruction than prior batch methods while transferring directly to pla...
-
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
Conditioning speed planning on the predicted path and relabeling synthetic cut-ins yields SOTA Bench2Drive scores (DS 89.07, SR 73.18%).
-
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
A cascaded end-to-end driving model conditions longitudinal planning on the lateral path via anchor-based regression and path-conditioned 1D displacement prediction, achieving SOTA driving score of 89.07 and 73.18% su...
-
SparseCoop: Cooperative Perception with Kinematic-Grounded Queries
SparseCoop delivers state-of-the-art 3D detection and tracking performance on V2X-Seq and Griffin datasets using only sparse kinematic queries instead of dense BEV features, with lower transmission cost and latency ro...
-
BePo: Dual Representation for 3D Occupancy Prediction
BePo proposes a dual BEV and sparse-points representation with cross-attention fusion for more accurate and efficient 3D occupancy prediction on autonomous driving benchmarks.
-
Can BEV Perception Gracefully Degrade under Sensor Failures?
Grace-BEV enables graceful degradation in BEV perception under sensor failures by using a TrustGate Router for modality trustworthiness and FailSafe Fusion Block for dynamic integration, with modality dropout training...
-
Grounded 3D-Aware Spatial Vision-Language Modeling
GR3D is a VLM that combines explicit 2D, implicit 2D, and monocular 3D grounding mechanisms to improve performance on spatial understanding benchmarks.
-
VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction
VGGT-Occ embeds geometric tokens via PA-DA and uses sequential coarse-to-fine gated fusion to reach 33.00% IoU and 21.08% mIoU on SurroundOcc-nuScenes while using only ~41M parameters in the occupancy head.
-
SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
SemLT3D introduces semantic-guided expert distillation with a language MoE module and CLIP projection to enrich features for long-tailed classes in camera-only 3D detection.
-
Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors
SEPatch3D accelerates ViT-based 3D object detectors up to 57% faster than StreamPETR via dynamic patch sizing and cross-granularity enhancement while keeping comparable accuracy on nuScenes and Argoverse 2.
-
Not All Agents Matter: From Global Attention Dilution to Risk-Prioritized Game Planning
GameAD models autonomous driving as a risk-prioritized game among agents via Risk-Aware Topology Anchoring, Minimax Risk-Aware Sparse Attention and related components, yielding safer trajectories than prior end-to-end...
-
Kerr-Schild Double Copy of the Randall-Sundrum Black String
Kerr-Schild double copy of the RSII black string gives a holographic-coordinate-independent sourceless single-copy gauge field and a zeroth copy with warp-induced mass m²=12/l², while an alternative split is inequivalent.
-
SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
A sparse transformer predicts multi-frame 3D occupancy from images without BEV or VAE tokenization and reports SOTA results on nuScenes for 1-3s forecasting under arbitrary trajectories.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.