Pith. sign in

REVIEW 7 cited by

Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.11926 v2 pith:QVXUTYRI submitted 2023-03-21 cs.CV

classification cs.CV
keywords multi-viewobjectstreampetrachievesdetectionframemethodmodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we propose a long-sequence modeling framework, named StreamPETR, for multi-view 3D object detection. Built upon the sparse query design in the PETR series, we systematically develop an object-centric temporal mechanism. The model is performed in an online manner and the long-term historical information is propagated through object queries frame by frame. Besides, we introduce a motion-aware layer normalization to model the movement of the objects. StreamPETR achieves significant performance improvements only with negligible computation cost, compared to the single-frame baseline. On the standard nuScenes benchmark, it is the first online multi-view method that achieves comparable performance (67.6% NDS & 65.3% AMOTA) with lidar-based methods. The lightweight version realizes 45.0% mAP and 31.7 FPS, outperforming the state-of-the-art method (SOLOFusion) by 2.3% mAP and 1.8x faster FPS. Code has been available at https://github.com/exiawsh/StreamPETR.git.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A contrastive learning framework with instance-level and perspective-level losses consistently improves multiple BEV detection models on nuScenes by up to 2.4 mAP.

  2. NoiseController: Towards Consistent Multi-view Video Generation via Noise Decomposition and Collaboration

    cs.CV 2025-04 conditional novelty 6.0 of 10

    NoiseController decomposes initial diffusion noise into scene-level foreground/background and shared/residual components, then collaborates them across views and frames, improving multi-view video consistency on nuScenes.

  3. Distant Object Localisation from Noisy Image Segmentation Sequences

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A bootstrap particle filter localizes distant static objects in 3D from noisy image segmentation sequences and GNSS camera poses, shown in simulation and one real drone sequence.

  4. Achieving Precise and Reliable Locomotion with Differentiable Simulation-Based System Identification

    cs.RO 2025-08 unverdicted novelty 5.0 of 10

    Estimating robot dynamics parameters from trajectory data alone inside a differentiable simulator, inside the reinforcement learning loop, is claimed to improve trajectory following in bipedal locomotion.

  5. Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A humanoid-specific multimodal occupancy perception system with a new dataset, sensor layout, and a fusion network that claims state-of-the-art results on its own benchmark.

  6. Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model

    cs.RO 2024-12 conditional novelty 5.0 of 10

    A reactive closed-loop driving simulator that uses a diffusion renderer with retrieval from real recordings, plus a nuPlan behavioral controller, to generate sensor images in response to an end-to-end driving model's actions.

  7. CrossTracker: Robust Multi-modal 3D Multi-Object Tracking via Cross Correction

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A two-stage camera-LiDAR 3D tracker that generates coarse dual-stream trajectories and then refines them via cross-modal correction, achieving state-of-the-art KITTI results.

Pith tools