Pith. sign in

REVIEW 3 major objections 6 minor 57 references

Pretraining a camera–radar driving backbone by forecasting future LiDAR makes the low-cost sensor stack learn reusable scene dynamics.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 17:40 UTC pith:VSNX5FTY

load-bearing objection Useful CR forecasting pretraining with real multi-task gains, but the bolded “CRISP” rows vs BEVFormer/UniAD-CRISP transfer rows are not cleanly defined and that muddies the transfer claim. the 3 major comments →

arxiv 2607.04541 v1 pith:VSNX5FTY submitted 2026-07-05 cs.CV cs.AIcs.LGcs.RO

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining

classification cs.CV cs.AIcs.LGcs.RO
keywords camera-radar fusionbird's-eye-viewforecasting pretrainingworld modelpoint cloud predictionautonomous drivingmodality gatingtemporal attention
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Most camera–radar models for autonomous driving are trained only for one labeled job at a time, so they do not learn a general scene representation. CRISP instead pretrains a shared bird’s-eye-view backbone by asking it to forecast future LiDAR point clouds from past multi-view images and radar only. LiDAR is used solely as privileged training supervision; the deployed system stays camera–radar. Three design pieces make the pretraining work for this sensor mix: an ego-aware radar encoder, radar-primed temporal attention so Doppler and range help memory propagate over time, and gated multimodal rendering that lets each BEV token admit camera or radar evidence selectively. On nuScenes the pretrained backbone improves long-horizon geometry forecasting and transfers to detection, tracking, mapping, motion forecasting, occupancy prediction, and planning, arguing that predictive pretraining is a practical way to scale driving representations without LiDAR at inference.

Core claim

The paper claims that a camera–radar spatiotemporal BEV backbone, optimized by forecasting future LiDAR occupancy from historical CR inputs, learns transferable multimodal geometry and motion features that beat strong camera-only and camera–radar baselines on long-horizon point-cloud forecasting and on a broad suite of perception, prediction, and planning tasks while remaining LiDAR-free at deployment.

What carries the argument

CRISP’s gated camera–radar BEV backbone: radar-enhanced temporal self-attention primes queries with range/Doppler before temporal aggregation, and multimodal feature rendering with modality innovation gating selectively admits residual camera and radar updates into each BEV token, supervised only by a future LiDAR occupancy decoder during pretraining.

Load-bearing premise

The load-bearing premise is that forecasting future LiDAR geometry from past camera and radar is a rich enough training signal to teach general driving dynamics, even when real future behavior depends on traffic lights, rules, or intent that geometry and Doppler alone do not show.

What would settle it

If, on held-out nuScenes-style sequences, a CRISP-pretrained CR backbone fails to improve long-horizon Chamfer distance or downstream NDS/AMOTA/planning collision rate over the same architecture trained from scratch or with camera-only forecasting pretraining, the claim that predictive CR pretraining yields reusable dynamics-aware features would be refuted.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. CRISP proposes a camera–radar spatiotemporal BEV backbone pretrained by forecasting future LiDAR point clouds from historical multi-view images and radar, with LiDAR used only as privileged pretraining supervision. The architecture extends BEVFormer/ViDAR-style encoding with an enhanced radar encoder (ego-motion bias + residual spatial calibration), radar-enhanced temporal self-attention (radar-primed queries and temporal gating), and multimodal feature rendering via independent Modality Innovation Gates over residual camera/radar innovations. On nuScenes, the paper reports improved long-horizon point-cloud forecasting versus CO and some LC baselines, and claims broad transfer to 3D detection, tracking, online mapping, motion forecasting, future occupancy, and planning under CR-only inference, with component ablations and a label-efficiency study.

Significance. If the multi-task transfer results hold under a clearly defined protocol, this is a solid systems contribution: it connects forecasting-based world-model pretraining with practical CR sensing, a configuration that is under-explored relative to CO/LC pretraining. Strengths include a coherent privileged-supervision design, explicit radar injection into temporal BEV updates rather than late fusion only, cumulative and design ablations (Tables VIII–XI), label-efficiency evidence (Fig. 5), and honest discussion of intent/traffic-signal failure modes (Fig. 8, §VI). The work is incremental relative to ViDAR and recent CR multi-task systems, but the combination is timely and potentially useful for scalable CR representation learning.

major comments (3)
  1. §V-B.2 and Tables II–VII: the multi-task transfer claim is load-bearing, but the tables are not self-consistent with the stated protocol. The text defines BEVFormer-CRISP / UniAD-CRISP as backbone replacement with original heads retained, yet each table also reports a separate bolded “CRISP” row that often substantially outperforms those transfer rows (e.g., Table II: CRISP 53.2 mAP / 61.3 NDS vs BEVFormer-CRISP 49.8 / 59.1; Table IV: 48.6 vs 44.6 AMOTA; Table VII: 0.80 m / 0.15% vs UniAD-CRISP 0.85 / 0.20). The manuscript never defines the architecture, heads, training schedule, or whether this standalone CRISP row is the forecasting model, a different multi-task system, or an uncontrolled stronger recipe. Without that definition, the bold numbers cannot be read as evidence for the claimed controlled transfer protocol.
  2. §V-B.2 prose vs Table II: there is a direct text/table identity error. The detection paragraph states “BEVFormer-CRISP achieves 53.2 mAP and 61.3 NDS,” but Table II assigns those numbers to the bold “CRISP” row and lists BEVFormer-CRISP as 49.8 / 59.1. This is not a cosmetic typo: it changes which system is credited with the headline detection result and further undermines confidence in the transfer tables.
  3. Tables II–VII baselines: SpaRC-AD is the main published CR multi-task comparator, but it is not clear that it is matched on backbone capacity, image resolution, radar preprocessing, training schedule, or whether it uses the same UniAD-style heads/evaluation. Given that the paper’s central claim is broad CR transfer rather than forecasting alone, a capacity- and protocol-matched CR baseline (or an ablated CRISP trained from scratch with the same heads) is needed to separate pretraining benefit from architecture/capacity effects already partially isolated only in the reduced ablation setting.
minor comments (6)
  1. §IV-C / Fig. 2 caption: the figure caption refers to a “future occupancy decoder,” while the body text mostly says “future prediction decoder” / point-cloud forecasting; align terminology.
  2. Eqs. (5)–(14): several Linear/ϕ operators are layer-specific but notationally dense; a short pseudocode block for one encoder layer would improve reproducibility.
  3. Table I: HERMES is marked with † as language-augmented; still, reporting only selected horizons makes the comparison hard to interpret—state missing entries explicitly as unavailable.
  4. §V-A.2: pretraining uses λ_dense but the numerical value is not stated in the provided text; please specify all loss weights and w_τ.
  5. Fig. 5 and ablation setting: reduced BEV grid / 1/8 pretrain subset is reasonable for cost, but state clearly that absolute ablation numbers are not comparable to full-model Tables I–VII.
  6. Minor writing issues: “A Vs” spacing, occasional “mA VE” line-break artifacts, and repeated “JOURNAL OF LATEX CLASS FILES” headers should be cleaned for camera-ready.

Circularity Check

0 steps flagged

No derivation circularity: CRISP is empirical CR pretraining with independent LiDAR-supervised forecasting and external downstream labels.

full rationale

CRISP does not claim a first-principles derivation that reduces to its inputs. The pretraining objective (Eq. 18) is standard supervised occupancy forecasting against privileged LiDAR geometry on historical CR inputs; evaluation of forecasting (Table I, Chamfer Distance) and of transfer (Tables II–VII) uses held-out nuScenes validation labels that are not fitted parameters of the pretext. Architectural inheritance from BEVFormer, ViDAR, and UniAD is methodological reuse of external recipes, not a self-citation uniqueness chain that forces the CR transfer claim. Self-citation of CRKD (Song/Skinner co-authors) appears only as related CR fusion work and is not load-bearing for the forecasting or transfer results. Ambiguity about what the standalone bolded “CRISP” table rows implement versus BEVFormer-CRISP/UniAD-CRISP is a reporting/protocol clarity issue, not circular reduction of a prediction to a fitted input. No self-definitional equations, fitted-as-prediction steps, or uniqueness-by-self-citation were found.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 3 invented entities

As an empirical AV pretraining paper, the claim rests less on free physical constants than on inherited BEV/world-model machinery, hand-chosen architecture hyperparameters, and the domain premise that LiDAR future geometry is a good privileged teacher for CR-only deployment. The invented entities are architectural modules whose value is evidenced only by the paper’s own ablations and transfers.

free parameters (6)
  • pretraining learning rate
    Set to 2e-4 for 24 epochs (§V-A.2); standard optimizer hyperparameter that affects final backbone quality.
  • BEV grid / range / channels
    200×200 over ±51.2 m with C=256 and 6 encoder layers chosen by practice from BEVFormer/UniAD/ViDAR; defines representation capacity.
  • radar voxel size and sweep count
    [0.512,0.512,8] voxels and 6 radar sweeps (§V-A.2) control radar sparsity and temporal aggregation into the BEV feature.
  • MIG channel groups
    8 groups matched to deformable-attention heads; controls granularity of gated camera/radar residual admission.
  • forecasting loss weights w_τ and λ_dense
    Current+3 future frames supervised with dense-loss weight λ_dense (Eq. 18); balances sparse surface classification vs dense voxel reconstruction without a unique theoretical value.
  • ablation data fractions / reduced encoder
    Ablations use 1/8 pretrain subset, 1/4 detection subset, 160×160 BEV, 3 layers (§V-D.1); free experimental choices that mediate component-effect estimates.
axioms (5)
  • domain assumption BEVFormer-style deformable temporal self-attention and camera spatial cross-attention are a valid base for spatiotemporal BEV encoding.
    CRISP inherits and modifies BEVFormer TSA/SCA (§III-A, §IV-C) rather than re-deriving view transformation.
  • domain assumption ViDAR-style future LiDAR occupancy/point-cloud forecasting is a useful pretraining signal for driving representations.
    Objective and decoder are adopted unchanged from ViDAR (§III-B, §IV-D); transfer claims depend on this proxy being task-relevant.
  • domain assumption Radar range and Doppler provide complementary motion/geometry cues that should condition temporal BEV propagation and residual fusion.
    Core design premise in Introduction and §IV-B/C; motivates radar-primed queries and MIG.
  • domain assumption nuScenes synchronized CR+LiDAR sequences are representative enough to support claims about practical CR pretraining transfer.
    All quantitative claims are on nuScenes (§V); Discussion §VI notes limited scale/diversity.
  • ad hoc to paper Independent sigmoid gates over residual innovations are preferable to competitive softmax fusion for complementary sensors.
    MIG design choice (§IV-C.2, Eq. 13–14); supported by ablation Table XI but not derived from first principles.
invented entities (3)
  • Modality Innovation Gate (MIG) no independent evidence
    purpose: Selectively admit camera and radar residual updates into each BEV token without overwriting the residual path.
    New fusion operator relative to ViDAR latent rendering; evidence is internal ablations (Table XI), not external independent measurement.
  • Radar-enhanced temporal self-attention (radar-primed query + temporal gate) no independent evidence
    purpose: Inject radar range/Doppler into temporal offset/attention prediction and replace uniform history/current averaging.
    Architectural invention of the paper (§IV-C.1, Eqs. 5–10); validated only within CRISP experiments.
  • Enhanced radar encoder (ego-motion feature bias + residual spatial calibration) no independent evidence
    purpose: Make Doppler features ego-state-aware and improve radar–camera BEV feature compatibility before fusion.
    Paper-specific radar front-end (§IV-B, Eqs. 3–4); ablation Table IX is the only support.

pith-pipeline@v1.1.0-grok45 · 29510 in / 3887 out tokens · 48317 ms · 2026-07-11T17:40:20.428612+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining." pith.science (2026). https://pith.science/paper/VSNX5FTY

@misc{pith2026260704541,
  author       = {Pith},
  title        = {Pith review of: CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VSNX5FTY}},
  note         = {Machine review of arXiv:2607.04541}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Camera-radar (CR) fusion is a practical sensing configuration for autonomous driving, but existing models are typically trained with task-specific supervision, limiting reusable representation learning. We present CRISP, a spatiotemporal CR backbone pretrained through forecasting-based representation learning. Given historical multi-view images and radar sweeps, CRISP learns a unified bird's-eye-view (BEV) representation by predicting future LiDAR point clouds. LiDAR is used only as privileged supervision during pretraining; the deployed model requires only camera and radar. To make forecasting-based pretraining effective for CR fusion, CRISP introduces an enhanced radar encoder, radar-enhanced temporal self-attention, and multimodal feature rendering with modality innovation gating. These components inject radar range and Doppler cues into BEV temporal propagation and allow BEV tokens to selectively incorporate camera and radar evidence. Experiments on nuScenes show that CRISP improves long-horizon point cloud forecasting and transfers effectively to downstream tasks, including 3D detection, tracking, online mapping, motion forecasting, future occupancy prediction, and planning, suggesting that predictive CR pretraining is a promising path toward scalable driving representations under practical sensor configurations. The project website is https://umfieldrobotics.github.io/CRISP.

Figures

Figures reproduced from arXiv: 2607.04541 by Jingyu Song, Katherine A. Skinner, Yi Liu.

Figure 1
Figure 1. Figure 1: Overview of the motivation of CRISP. We use low-cost camera-radar [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 1
Figure 1. Figure 1: CRISP is designed to capture multimodal geometry, [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of CRISP. Historical multi-view images and aggregated radar observations are encoded into a shared CR BEV backbone. Radar first [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Radar-enhanced temporal self-attention. The calibrated radar BEV () [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Multimodal feature rendering. Camera and radar cross-attention [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Label-efficiency of CRISP pretraining on nuScenes 3D detection. [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Qualitative comparison between CRISP and ViDAR for future world-geometry prediction in a low-light scene. We show the front camera views at the [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Qualitative comparison between CRISP and ViDAR for future world-geometry prediction. We show the front camera views at the current frame (0s) [PITH_FULL_IMAGE:figures/full_fig_p014_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Failure case analysis for future world-geometry prediction. We show the front camera views at the current frame (0s) and future timestamps for [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

57 extracted references · 5 linked inside Pith

  1. [1]

    nuscenes: A multi- modal dataset for autonomous driving,

    H. Caesar, V . Bankiti, A. H. Lang, S. V ora, V . E. Liong, Q. Xu, A. Krishnan, Y . Pan, G. Baldan, and O. Beijbom, “nuscenes: A multi- modal dataset for autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020

  2. [2]

    Scalability in perception for autonomous driving: Waymo open dataset,

    P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V . Patnaik, P. Tsui, J. Guo, Y . Zhou, Y . Chai, B. Caineet al., “Scalability in perception for autonomous driving: Waymo open dataset,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 2446–2454

  3. [3]

    Planning-oriented autonomous driving,

    Y . Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wanget al., “Planning-oriented autonomous driving,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023

  4. [4]

    Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,

    L. Wang, X. Zhang, Z. Song, J. Bi, G. Zhang, H. Wei, L. Tang, L. Yang, J. Li, C. Jia, and L. Zhao, “Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,”IEEE Transactions on Intelligent Vehicles, vol. 8, no. 7, pp. 3781–3798, 2023

  5. [5]

    Deep learning-based perception systems for autonomous driving: A comprehensive survey,

    L.-H. Wen and K.-H. Jo, “Deep learning-based perception systems for autonomous driving: A comprehensive survey,”Neurocomputing, vol. 489, pp. 255–270, 2022

  6. [6]

    Towards deep radar perception for autonomous driving: Datasets, methods, and challenges,

    Y . Zhou, L. Liu, H. Zhao, M. L ´opez-Ben´ıtez, L. Yu, and Y . Yue, “Towards deep radar perception for autonomous driving: Datasets, methods, and challenges,”Sensors, vol. 22, no. 11, p. 4208, 2022

  7. [7]

    Crkd: Enhanced camera-radar object detection with cross-modality knowledge distillation,

    L. Zhao, J. Song, and K. A. Skinner, “Crkd: Enhanced camera-radar object detection with cross-modality knowledge distillation,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 15 470–15 480

  8. [8]

    Multi-modal 3d object detection in autonomous driving: a survey,

    Y . Wang, Q. Mao, H. Zhu, J. Deng, Y . Zhang, J. Ji, H. Li, and Y . Zhang, “Multi-modal 3d object detection in autonomous driving: a survey,” International Journal of Computer Vision, pp. 1–31, 2023

  9. [9]

    Vision meets robotics: The kitti dataset,

    A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The kitti dataset,”The International Journal of Robotics Research, vol. 32, no. 11, pp. 1231–1237, 2013

  10. [10]

    Unifying voxel-based representation with transformer for 3d object detection,

    Y . Li, Y . Chen, X. Qi, Z. Li, J. Sun, and J. Jia, “Unifying voxel-based representation with transformer for 3d object detection,” inAdvances in Neural Information Processing Systems (NeurIPS), 2022

  11. [11]

    Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation,

    Z. Liu, H. Tang, A. Amini, X. Yang, H. Mao, D. L. Rus, and S. Han, “Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation,” in2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 2774–2781

  12. [12]

    Futr3d: A unified sensor fusion framework for 3d detection,

    X. Chen, T. Zhang, Y . Wang, Y . Wang, and H. Zhao, “Futr3d: A unified sensor fusion framework for 3d detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), June 2023

  13. [13]

    Radar and camera fusion for object detection and tracking: A comprehensive survey,

    K. Shi, S. He, Z. Shi, A. Chen, Z. Xiong, J. Chen, and J. Luo, “Radar and camera fusion for object detection and tracking: A comprehensive survey,”arXiv preprint arXiv:2410.19872, 2024

  14. [14]

    Centerfusion: Center-based radar and camera fusion for 3d object detection,

    R. Nabati and H. Qi, “Centerfusion: Center-based radar and camera fusion for 3d object detection,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021, pp. 1527–1536

  15. [15]

    BEV-Guided Multi-Modality Fusion for Driving Perception,

    Y . Man, L.-Y . Gui, and Y .-X. Wang, “BEV-Guided Multi-Modality Fusion for Driving Perception,” inProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), 2023

  16. [16]

    Crn: Camera radar net for accurate, robust, efficient 3d perception,

    Y . Kim, J. Shin, S. Kim, I.-J. Lee, J. W. Choi, and D. Kum, “Crn: Camera radar net for accurate, robust, efficient 3d perception,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023

  17. [17]

    Bevcar: Camera-radar fusion for bev map and object segmentation,

    J. Schramm, N. V ¨odisch, K. Petek, B. R. Kiran, S. Yogamani, W. Bur- gard, and A. Valada, “Bevcar: Camera-radar fusion for bev map and object segmentation,” in2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2024, pp. 1435–1442

  18. [18]

    Licrocc: Teach radar for accurate semantic occupancy prediction using lidar and camera,

    Y . Ma, J. Mei, X. Yang, L. Wen, W. Xu, J. Zhang, X. Zuo, B. Shi, and Y . Liu, “Licrocc: Teach radar for accurate semantic occupancy prediction using lidar and camera,”IEEE Robotics and Automation Letters, 2024

  19. [19]

    Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,

    J. Philion and S. Fidler, “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,” inProceedings of the European Conference on Computer Vision (ECCV), 2020. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 16

  20. [20]

    Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,

    Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y . Qiao, and J. Dai, “Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,” inProceedings of the European Conference on Computer Vision (ECCV). Springer, 2022, pp. 1–18

  21. [21]

    Gd-mae: generative decoder for mae pre-training on lidar point clouds,

    H. Yang, T. He, J. Liu, H. Chen, B. Wu, B. Lin, X. He, and W. Ouyang, “Gd-mae: generative decoder for mae pre-training on lidar point clouds,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 9403–9414

  22. [22]

    Unipad: A universal pre-training paradigm for autonomous driving,

    H. Yang, S. Zhang, D. Huang, X. Wu, H. Zhu, T. He, S. Tang, H. Zhao, Q. Qiu, B. Linet al., “Unipad: A universal pre-training paradigm for autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 15 238– 15 250

  23. [23]

    Visual point cloud forecasting enables scalable autonomous driving,

    Z. Yang, L. Chen, Y . Sun, and H. Li, “Visual point cloud forecasting enables scalable autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 14 673–14 684

  24. [24]

    Forging spatial intelligence: A roadmap of multi-modal data pre- training for autonomous systems,

    S. Wang, L. Kong, X. Liu, H. Shi, W. Li, J. Zhu, and S. C. H. Hoi, “Forging spatial intelligence: A roadmap of multi-modal data pre- training for autonomous systems,”arXiv preprint arXiv:2512.24385, 2025

  25. [25]

    Masked autoencoder for self-supervised pre-training on lidar point clouds,

    G. Hess, J. Jaxing, E. Svensson, D. Hagerman, C. Petersson, and L. Svensson, “Masked autoencoder for self-supervised pre-training on lidar point clouds,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023, pp. 350–359

  26. [26]

    Bev-mae: Bird’s eye view masked autoencoders for point cloud pre-training in autonomous driving scenarios,

    Z. Lin, Y . Wang, S. Qi, N. Dong, and M.-H. Yang, “Bev-mae: Bird’s eye view masked autoencoders for point cloud pre-training in autonomous driving scenarios,” inProceedings of the AAAI Conference on Artificial Intelligence (AAAI), vol. 38, no. 4, 2024, pp. 3531–3539

  27. [27]

    Is pseudo- lidar needed for monocular 3d object detection?

    D. Park, R. Ambrus, V . Guizilini, J. Li, and A. Gaidon, “Is pseudo- lidar needed for monocular 3d object detection?” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 3142–3152

  28. [28]

    Pimae: Point cloud and image interactive masked autoencoders for 3d object detection,

    A. Chen, K. Zhang, R. Zhang, Z. Wang, Y . Lu, Y . Guo, and S. Zhang, “Pimae: Point cloud and image interactive masked autoencoders for 3d object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 5291– 5301

  29. [29]

    Point cloud forecasting as a proxy for 4d occupancy forecasting,

    T. Khurana, P. Hu, D. Held, and D. Ramanan, “Point cloud forecasting as a proxy for 4d occupancy forecasting,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 1116–1124

  30. [30]

    Uniworld: Au- tonomous driving pre-training via world models,

    C. Min, D. Zhao, L. Xiao, Y . Nie, and B. Dai, “Uniworld: Au- tonomous driving pre-training via world models,”arXiv preprint arXiv:2308.07234, 2023

  31. [31]

    Driveworld: 4d pre- trained scene understanding via world models for autonomous driving,

    C. Min, D. Zhao, L. Xiao, J. Zhao, X. Xu, Z. Zhu, L. Jin, J. Li, Y . Guo, J. Xing, L. Jing, Y . Nie, and B. Dai, “Driveworld: 4d pre- trained scene understanding via world models for autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 15 522–15 533

  32. [32]

    Crt-fusion: Camera, radar, temporal fusion using motion information for 3d object detection,

    J. Kim, M. Seong, and J. W. Choi, “Crt-fusion: Camera, radar, temporal fusion using motion information for 3d object detection,” inAdvances in Neural Information Processing Systems (NeurIPS), vol. 37, 2024, pp. 108 625–108 648

  33. [33]

    Cross-modality knowledge distillation network for monocular 3d object detection,

    Y . Hong, H. Dai, and Y . Ding, “Cross-modality knowledge distillation network for monocular 3d object detection,” inProceedings of the European Conference on Computer Vision (ECCV). Springer, 2022, pp. 87–104

  34. [34]

    X-align: Cross-modal cross-view alignment for bird’s-eye-view segmentation,

    S. Borse, M. Klingner, V . R. Kumar, H. Cai, A. Almuzairee, S. Yoga- mani, and F. Porikli, “X-align: Cross-modal cross-view alignment for bird’s-eye-view segmentation,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023, pp. 3287–3297

  35. [35]

    Unifying voxel-based representation with transformer for 3d object detection,

    Y . Li, Y . Chen, X. Qi, Z. Li, J. Sun, and J. Jia, “Unifying voxel-based representation with transformer for 3d object detection,” inAdvances in Neural Information Processing Systems (NeurIPS), vol. 35, 2022, pp. 18 442–18 455

  36. [36]

    Masked au- toencoders are scalable vision learners,

    K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked au- toencoders are scalable vision learners,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 16 000–16 009

  37. [37]

    Point-bert: Pre-training 3d point cloud transformers with masked point modeling,

    X. Yu, L. Tang, Y . Rao, T. Huang, J. Zhou, and J. Lu, “Point-bert: Pre-training 3d point cloud transformers with masked point modeling,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 19 313–19 322

  38. [38]

    Masked autoencoders for point cloud self-supervised learning,

    Y . Pang, W. Wang, F. E. Tay, W. Liu, Y . Tian, and L. Yuan, “Masked autoencoders for point cloud self-supervised learning,” inProceedings of the European Conference on Computer Vision (ECCV). Springer, 2022, pp. 604–621

  39. [39]

    Exploring geometry-aware contrast and clustering harmonization for self-supervised 3d object detection,

    H. Liang, C. Jiang, D. Feng, X. Chen, H. Xu, X. Liang, W. Zhang, Z. Li, and L. Van Gool, “Exploring geometry-aware contrast and clustering harmonization for self-supervised 3d object detection,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 3293–3302

  40. [40]

    Visionpad: A vision-centric pre-training paradigm for autonomous driving,

    H. Zhang, W. Zhou, Y . Zhu, X. Yan, J. Gao, D. Bai, Y . Cai, B. Liu, S. Cui, and Z. Li, “Visionpad: A vision-centric pre-training paradigm for autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2025, pp. 17 165–17 175

  41. [41]

    Bootstrapping autonomous driving radars with self-supervised learn- ing,

    Y . Hao, S. Madani, J. Guan, M. Alloulah, S. Gupta, and H. Hassanieh, “Bootstrapping autonomous driving radars with self-supervised learn- ing,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 15 012–15 023

  42. [42]

    Self-supervised sparse sensor fusion for long range percep- tion,

    E. Palladin, S. Brucker, F. Ghilotti, P. Narayanan, M. Bijelic, and F. Heide, “Self-supervised sparse sensor fusion for long range percep- tion,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025

  43. [43]

    Multi-sensor fusion in automated driving: A survey,

    Z. Wang, Y . Wu, and Q. Niu, “Multi-sensor fusion in automated driving: A survey,”IEEE Access, vol. 8, pp. 2847–2868, 2019

  44. [44]

    Radar-camera fusion for object detection and semantic segmentation in autonomous driving: A comprehensive review,

    S. Yao, R. Guan, X. Huang, Z. Li, X. Sha, Y . Yue, E. G. Lim, H. Seo, K. L. Man, X. Zhu, and Y . Yue, “Radar-camera fusion for object detection and semantic segmentation in autonomous driving: A comprehensive review,”IEEE Transactions on Intelligent Vehicles, vol. 9, no. 1, pp. 2094–2128, 2024

  45. [45]

    Cramnet: Camera-radar fusion with ray-constrained cross-attention for robust 3d object detection,

    J.-J. Hwang, H. Kretzschmar, J. Manela, S. Rafferty, N. Armstrong- Crews, T. Chen, and D. Anguelov, “Cramnet: Camera-radar fusion with ray-constrained cross-attention for robust 3d object detection,” in Proceedings of the European Conference on Computer Vision (ECCV). Springer, 2022

  46. [46]

    Mvfusion: Multi-view 3d object detection with semantic-aligned radar and camera fusion,

    Z. Wu, G. Chen, Y . Gan, L. Wang, and J. Pu, “Mvfusion: Multi-view 3d object detection with semantic-aligned radar and camera fusion,” in2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 2766–2773

  47. [47]

    Ea-lss: Edge-aware lift-splat-shot framework for 3d bev object detection,

    H. Hu, F. Wang, J. Su, Y . Wang, L. Hu, W. Fang, J. Xu, and Z. Zhang, “Ea-lss: Edge-aware lift-splat-shot framework for 3d bev object detection,”arXiv preprint arXiv:2303.17895, 2023

  48. [48]

    Rcbevdet: Radar-camera fusion in bird’s eye view for 3d object detection,

    Z. Lin, Z. Liu, Z. Xia, X. Wang, Y . Wang, S. Qi, Y . Dong, N. Dong, L. Zhang, and C. Zhu, “Rcbevdet: Radar-camera fusion in bird’s eye view for 3d object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 14 928–14 937

  49. [49]

    Sparc-ad: A baseline for radar-camera fusion in end-to-end autonomous driving,

    P. Wolters, J. Gilg, T. Teepe, and G. Rigoll, “Sparc-ad: A baseline for radar-camera fusion in end-to-end autonomous driving,” inProceedings of the IEEE/CVF International Conference on Computer Vision Work- shops (ICCVW), October 2025, pp. 1831–1841

  50. [50]

    Hermes: A unified self-driving world model for simultaneous 3d scene understanding and generation,

    X. Zhou, D. Liang, S. Tu, X. Chen, Y . Ding, D. Zhang, F. Tan, H. Zhao, and X. Bai, “Hermes: A unified self-driving world model for simultaneous 3d scene understanding and generation,”arXiv preprint arXiv:2501.14729, 2025

  51. [51]

    Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free,

    Z. Qiu, Z. Wang, B. Zheng, Z. Huang, K. Wen, S. Yang, R. Men, L. Yu, F. Huang, S. Huanget al., “Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free,” inAdvances in Neural Information Processing Systems (NeurIPS)

  52. [52]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778

  53. [53]

    Fcos3d: Fully convolutional one- stage monocular 3d object detection,

    T. Wang, X. Zhu, J. Pang, and D. Lin, “Fcos3d: Fully convolutional one- stage monocular 3d object detection,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 913– 922

  54. [54]

    Feature pyramid networks for object detection,

    T.-Y . Lin, P. Doll´ar, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2117–2125

  55. [55]

    Pointpillars: Fast encoders for object detection from point clouds,

    A. H. Lang, S. V ora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom, “Pointpillars: Fast encoders for object detection from point clouds,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  56. [56]

    Bevformer: learning bird’s-eye-view representation from lidar-camera JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 17 via spatiotemporal transformers,

    Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Q. Yu, and J. Dai, “Bevformer: learning bird’s-eye-view representation from lidar-camera JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 17 via spatiotemporal transformers,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 3, pp. 2020–2036, 2024

  57. [57]

    nuplan: A closed-loop ml- based planning benchmark for autonomous vehicles,

    H. Caesar, J. Kabzan, K. S. Tan, W. K. Fong, E. Wolff, A. Lang, L. Fletcher, O. Beijbom, and S. Omari, “nuplan: A closed-loop ml- based planning benchmark for autonomous vehicles,”arXiv preprint arXiv:2106.11810, 2021