REVIEW 3 major objections 6 minor 57 references
Pretraining a camera–radar driving backbone by forecasting future LiDAR makes the low-cost sensor stack learn reusable scene dynamics.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 17:40 UTC pith:VSNX5FTY
load-bearing objection Useful CR forecasting pretraining with real multi-task gains, but the bolded “CRISP” rows vs BEVFormer/UniAD-CRISP transfer rows are not cleanly defined and that muddies the transfer claim. the 3 major comments →
CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that a camera–radar spatiotemporal BEV backbone, optimized by forecasting future LiDAR occupancy from historical CR inputs, learns transferable multimodal geometry and motion features that beat strong camera-only and camera–radar baselines on long-horizon point-cloud forecasting and on a broad suite of perception, prediction, and planning tasks while remaining LiDAR-free at deployment.
What carries the argument
CRISP’s gated camera–radar BEV backbone: radar-enhanced temporal self-attention primes queries with range/Doppler before temporal aggregation, and multimodal feature rendering with modality innovation gating selectively admits residual camera and radar updates into each BEV token, supervised only by a future LiDAR occupancy decoder during pretraining.
Load-bearing premise
The load-bearing premise is that forecasting future LiDAR geometry from past camera and radar is a rich enough training signal to teach general driving dynamics, even when real future behavior depends on traffic lights, rules, or intent that geometry and Doppler alone do not show.
What would settle it
If, on held-out nuScenes-style sequences, a CRISP-pretrained CR backbone fails to improve long-horizon Chamfer distance or downstream NDS/AMOTA/planning collision rate over the same architecture trained from scratch or with camera-only forecasting pretraining, the claim that predictive CR pretraining yields reusable dynamics-aware features would be refuted.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. CRISP proposes a camera–radar spatiotemporal BEV backbone pretrained by forecasting future LiDAR point clouds from historical multi-view images and radar, with LiDAR used only as privileged pretraining supervision. The architecture extends BEVFormer/ViDAR-style encoding with an enhanced radar encoder (ego-motion bias + residual spatial calibration), radar-enhanced temporal self-attention (radar-primed queries and temporal gating), and multimodal feature rendering via independent Modality Innovation Gates over residual camera/radar innovations. On nuScenes, the paper reports improved long-horizon point-cloud forecasting versus CO and some LC baselines, and claims broad transfer to 3D detection, tracking, online mapping, motion forecasting, future occupancy, and planning under CR-only inference, with component ablations and a label-efficiency study.
Significance. If the multi-task transfer results hold under a clearly defined protocol, this is a solid systems contribution: it connects forecasting-based world-model pretraining with practical CR sensing, a configuration that is under-explored relative to CO/LC pretraining. Strengths include a coherent privileged-supervision design, explicit radar injection into temporal BEV updates rather than late fusion only, cumulative and design ablations (Tables VIII–XI), label-efficiency evidence (Fig. 5), and honest discussion of intent/traffic-signal failure modes (Fig. 8, §VI). The work is incremental relative to ViDAR and recent CR multi-task systems, but the combination is timely and potentially useful for scalable CR representation learning.
major comments (3)
- §V-B.2 and Tables II–VII: the multi-task transfer claim is load-bearing, but the tables are not self-consistent with the stated protocol. The text defines BEVFormer-CRISP / UniAD-CRISP as backbone replacement with original heads retained, yet each table also reports a separate bolded “CRISP” row that often substantially outperforms those transfer rows (e.g., Table II: CRISP 53.2 mAP / 61.3 NDS vs BEVFormer-CRISP 49.8 / 59.1; Table IV: 48.6 vs 44.6 AMOTA; Table VII: 0.80 m / 0.15% vs UniAD-CRISP 0.85 / 0.20). The manuscript never defines the architecture, heads, training schedule, or whether this standalone CRISP row is the forecasting model, a different multi-task system, or an uncontrolled stronger recipe. Without that definition, the bold numbers cannot be read as evidence for the claimed controlled transfer protocol.
- §V-B.2 prose vs Table II: there is a direct text/table identity error. The detection paragraph states “BEVFormer-CRISP achieves 53.2 mAP and 61.3 NDS,” but Table II assigns those numbers to the bold “CRISP” row and lists BEVFormer-CRISP as 49.8 / 59.1. This is not a cosmetic typo: it changes which system is credited with the headline detection result and further undermines confidence in the transfer tables.
- Tables II–VII baselines: SpaRC-AD is the main published CR multi-task comparator, but it is not clear that it is matched on backbone capacity, image resolution, radar preprocessing, training schedule, or whether it uses the same UniAD-style heads/evaluation. Given that the paper’s central claim is broad CR transfer rather than forecasting alone, a capacity- and protocol-matched CR baseline (or an ablated CRISP trained from scratch with the same heads) is needed to separate pretraining benefit from architecture/capacity effects already partially isolated only in the reduced ablation setting.
minor comments (6)
- §IV-C / Fig. 2 caption: the figure caption refers to a “future occupancy decoder,” while the body text mostly says “future prediction decoder” / point-cloud forecasting; align terminology.
- Eqs. (5)–(14): several Linear/ϕ operators are layer-specific but notationally dense; a short pseudocode block for one encoder layer would improve reproducibility.
- Table I: HERMES is marked with † as language-augmented; still, reporting only selected horizons makes the comparison hard to interpret—state missing entries explicitly as unavailable.
- §V-A.2: pretraining uses λ_dense but the numerical value is not stated in the provided text; please specify all loss weights and w_τ.
- Fig. 5 and ablation setting: reduced BEV grid / 1/8 pretrain subset is reasonable for cost, but state clearly that absolute ablation numbers are not comparable to full-model Tables I–VII.
- Minor writing issues: “A Vs” spacing, occasional “mA VE” line-break artifacts, and repeated “JOURNAL OF LATEX CLASS FILES” headers should be cleaned for camera-ready.
Circularity Check
No derivation circularity: CRISP is empirical CR pretraining with independent LiDAR-supervised forecasting and external downstream labels.
full rationale
CRISP does not claim a first-principles derivation that reduces to its inputs. The pretraining objective (Eq. 18) is standard supervised occupancy forecasting against privileged LiDAR geometry on historical CR inputs; evaluation of forecasting (Table I, Chamfer Distance) and of transfer (Tables II–VII) uses held-out nuScenes validation labels that are not fitted parameters of the pretext. Architectural inheritance from BEVFormer, ViDAR, and UniAD is methodological reuse of external recipes, not a self-citation uniqueness chain that forces the CR transfer claim. Self-citation of CRKD (Song/Skinner co-authors) appears only as related CR fusion work and is not load-bearing for the forecasting or transfer results. Ambiguity about what the standalone bolded “CRISP” table rows implement versus BEVFormer-CRISP/UniAD-CRISP is a reporting/protocol clarity issue, not circular reduction of a prediction to a fitted input. No self-definitional equations, fitted-as-prediction steps, or uniqueness-by-self-citation were found.
Axiom & Free-Parameter Ledger
free parameters (6)
- pretraining learning rate
- BEV grid / range / channels
- radar voxel size and sweep count
- MIG channel groups
- forecasting loss weights w_τ and λ_dense
- ablation data fractions / reduced encoder
axioms (5)
- domain assumption BEVFormer-style deformable temporal self-attention and camera spatial cross-attention are a valid base for spatiotemporal BEV encoding.
- domain assumption ViDAR-style future LiDAR occupancy/point-cloud forecasting is a useful pretraining signal for driving representations.
- domain assumption Radar range and Doppler provide complementary motion/geometry cues that should condition temporal BEV propagation and residual fusion.
- domain assumption nuScenes synchronized CR+LiDAR sequences are representative enough to support claims about practical CR pretraining transfer.
- ad hoc to paper Independent sigmoid gates over residual innovations are preferable to competitive softmax fusion for complementary sensors.
invented entities (3)
-
Modality Innovation Gate (MIG)
no independent evidence
-
Radar-enhanced temporal self-attention (radar-primed query + temporal gate)
no independent evidence
-
Enhanced radar encoder (ego-motion feature bias + residual spatial calibration)
no independent evidence
Cite this review
Pith. "Pith review of CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining." pith.science (2026). https://pith.science/paper/VSNX5FTY
@misc{pith2026260704541,
author = {Pith},
title = {Pith review of: CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining},
year = {2026},
howpublished = {\url{https://pith.science/paper/VSNX5FTY}},
note = {Machine review of arXiv:2607.04541}
}
read the original abstract
Camera-radar (CR) fusion is a practical sensing configuration for autonomous driving, but existing models are typically trained with task-specific supervision, limiting reusable representation learning. We present CRISP, a spatiotemporal CR backbone pretrained through forecasting-based representation learning. Given historical multi-view images and radar sweeps, CRISP learns a unified bird's-eye-view (BEV) representation by predicting future LiDAR point clouds. LiDAR is used only as privileged supervision during pretraining; the deployed model requires only camera and radar. To make forecasting-based pretraining effective for CR fusion, CRISP introduces an enhanced radar encoder, radar-enhanced temporal self-attention, and multimodal feature rendering with modality innovation gating. These components inject radar range and Doppler cues into BEV temporal propagation and allow BEV tokens to selectively incorporate camera and radar evidence. Experiments on nuScenes show that CRISP improves long-horizon point cloud forecasting and transfers effectively to downstream tasks, including 3D detection, tracking, online mapping, motion forecasting, future occupancy prediction, and planning, suggesting that predictive CR pretraining is a promising path toward scalable driving representations under practical sensor configurations. The project website is https://umfieldrobotics.github.io/CRISP.
Figures
Reference graph
Works this paper leans on
-
[1]
nuscenes: A multi- modal dataset for autonomous driving,
H. Caesar, V . Bankiti, A. H. Lang, S. V ora, V . E. Liong, Q. Xu, A. Krishnan, Y . Pan, G. Baldan, and O. Beijbom, “nuscenes: A multi- modal dataset for autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
2020
-
[2]
Scalability in perception for autonomous driving: Waymo open dataset,
P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V . Patnaik, P. Tsui, J. Guo, Y . Zhou, Y . Chai, B. Caineet al., “Scalability in perception for autonomous driving: Waymo open dataset,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 2446–2454
2020
-
[3]
Planning-oriented autonomous driving,
Y . Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wanget al., “Planning-oriented autonomous driving,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023
2023
-
[4]
Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,
L. Wang, X. Zhang, Z. Song, J. Bi, G. Zhang, H. Wei, L. Tang, L. Yang, J. Li, C. Jia, and L. Zhao, “Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,”IEEE Transactions on Intelligent Vehicles, vol. 8, no. 7, pp. 3781–3798, 2023
2023
-
[5]
Deep learning-based perception systems for autonomous driving: A comprehensive survey,
L.-H. Wen and K.-H. Jo, “Deep learning-based perception systems for autonomous driving: A comprehensive survey,”Neurocomputing, vol. 489, pp. 255–270, 2022
2022
-
[6]
Towards deep radar perception for autonomous driving: Datasets, methods, and challenges,
Y . Zhou, L. Liu, H. Zhao, M. L ´opez-Ben´ıtez, L. Yu, and Y . Yue, “Towards deep radar perception for autonomous driving: Datasets, methods, and challenges,”Sensors, vol. 22, no. 11, p. 4208, 2022
2022
-
[7]
Crkd: Enhanced camera-radar object detection with cross-modality knowledge distillation,
L. Zhao, J. Song, and K. A. Skinner, “Crkd: Enhanced camera-radar object detection with cross-modality knowledge distillation,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 15 470–15 480
2024
-
[8]
Multi-modal 3d object detection in autonomous driving: a survey,
Y . Wang, Q. Mao, H. Zhu, J. Deng, Y . Zhang, J. Ji, H. Li, and Y . Zhang, “Multi-modal 3d object detection in autonomous driving: a survey,” International Journal of Computer Vision, pp. 1–31, 2023
2023
-
[9]
Vision meets robotics: The kitti dataset,
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The kitti dataset,”The International Journal of Robotics Research, vol. 32, no. 11, pp. 1231–1237, 2013
2013
-
[10]
Unifying voxel-based representation with transformer for 3d object detection,
Y . Li, Y . Chen, X. Qi, Z. Li, J. Sun, and J. Jia, “Unifying voxel-based representation with transformer for 3d object detection,” inAdvances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[11]
Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation,
Z. Liu, H. Tang, A. Amini, X. Yang, H. Mao, D. L. Rus, and S. Han, “Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation,” in2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 2774–2781
2023
-
[12]
Futr3d: A unified sensor fusion framework for 3d detection,
X. Chen, T. Zhang, Y . Wang, Y . Wang, and H. Zhao, “Futr3d: A unified sensor fusion framework for 3d detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), June 2023
2023
-
[13]
Radar and camera fusion for object detection and tracking: A comprehensive survey,
K. Shi, S. He, Z. Shi, A. Chen, Z. Xiong, J. Chen, and J. Luo, “Radar and camera fusion for object detection and tracking: A comprehensive survey,”arXiv preprint arXiv:2410.19872, 2024
Pith/arXiv arXiv 2024
-
[14]
Centerfusion: Center-based radar and camera fusion for 3d object detection,
R. Nabati and H. Qi, “Centerfusion: Center-based radar and camera fusion for 3d object detection,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021, pp. 1527–1536
2021
-
[15]
BEV-Guided Multi-Modality Fusion for Driving Perception,
Y . Man, L.-Y . Gui, and Y .-X. Wang, “BEV-Guided Multi-Modality Fusion for Driving Perception,” inProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), 2023
2023
-
[16]
Crn: Camera radar net for accurate, robust, efficient 3d perception,
Y . Kim, J. Shin, S. Kim, I.-J. Lee, J. W. Choi, and D. Kum, “Crn: Camera radar net for accurate, robust, efficient 3d perception,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023
2023
-
[17]
Bevcar: Camera-radar fusion for bev map and object segmentation,
J. Schramm, N. V ¨odisch, K. Petek, B. R. Kiran, S. Yogamani, W. Bur- gard, and A. Valada, “Bevcar: Camera-radar fusion for bev map and object segmentation,” in2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2024, pp. 1435–1442
2024
-
[18]
Licrocc: Teach radar for accurate semantic occupancy prediction using lidar and camera,
Y . Ma, J. Mei, X. Yang, L. Wen, W. Xu, J. Zhang, X. Zuo, B. Shi, and Y . Liu, “Licrocc: Teach radar for accurate semantic occupancy prediction using lidar and camera,”IEEE Robotics and Automation Letters, 2024
2024
-
[19]
Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,
J. Philion and S. Fidler, “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,” inProceedings of the European Conference on Computer Vision (ECCV), 2020. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 16
2020
-
[20]
Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,
Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y . Qiao, and J. Dai, “Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,” inProceedings of the European Conference on Computer Vision (ECCV). Springer, 2022, pp. 1–18
2022
-
[21]
Gd-mae: generative decoder for mae pre-training on lidar point clouds,
H. Yang, T. He, J. Liu, H. Chen, B. Wu, B. Lin, X. He, and W. Ouyang, “Gd-mae: generative decoder for mae pre-training on lidar point clouds,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 9403–9414
2023
-
[22]
Unipad: A universal pre-training paradigm for autonomous driving,
H. Yang, S. Zhang, D. Huang, X. Wu, H. Zhu, T. He, S. Tang, H. Zhao, Q. Qiu, B. Linet al., “Unipad: A universal pre-training paradigm for autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 15 238– 15 250
2024
-
[23]
Visual point cloud forecasting enables scalable autonomous driving,
Z. Yang, L. Chen, Y . Sun, and H. Li, “Visual point cloud forecasting enables scalable autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 14 673–14 684
2024
-
[24]
Forging spatial intelligence: A roadmap of multi-modal data pre- training for autonomous systems,
S. Wang, L. Kong, X. Liu, H. Shi, W. Li, J. Zhu, and S. C. H. Hoi, “Forging spatial intelligence: A roadmap of multi-modal data pre- training for autonomous systems,”arXiv preprint arXiv:2512.24385, 2025
arXiv 2025
-
[25]
Masked autoencoder for self-supervised pre-training on lidar point clouds,
G. Hess, J. Jaxing, E. Svensson, D. Hagerman, C. Petersson, and L. Svensson, “Masked autoencoder for self-supervised pre-training on lidar point clouds,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023, pp. 350–359
2023
-
[26]
Bev-mae: Bird’s eye view masked autoencoders for point cloud pre-training in autonomous driving scenarios,
Z. Lin, Y . Wang, S. Qi, N. Dong, and M.-H. Yang, “Bev-mae: Bird’s eye view masked autoencoders for point cloud pre-training in autonomous driving scenarios,” inProceedings of the AAAI Conference on Artificial Intelligence (AAAI), vol. 38, no. 4, 2024, pp. 3531–3539
2024
-
[27]
Is pseudo- lidar needed for monocular 3d object detection?
D. Park, R. Ambrus, V . Guizilini, J. Li, and A. Gaidon, “Is pseudo- lidar needed for monocular 3d object detection?” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 3142–3152
2021
-
[28]
Pimae: Point cloud and image interactive masked autoencoders for 3d object detection,
A. Chen, K. Zhang, R. Zhang, Z. Wang, Y . Lu, Y . Guo, and S. Zhang, “Pimae: Point cloud and image interactive masked autoencoders for 3d object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 5291– 5301
2023
-
[29]
Point cloud forecasting as a proxy for 4d occupancy forecasting,
T. Khurana, P. Hu, D. Held, and D. Ramanan, “Point cloud forecasting as a proxy for 4d occupancy forecasting,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 1116–1124
2023
-
[30]
Uniworld: Au- tonomous driving pre-training via world models,
C. Min, D. Zhao, L. Xiao, Y . Nie, and B. Dai, “Uniworld: Au- tonomous driving pre-training via world models,”arXiv preprint arXiv:2308.07234, 2023
Pith/arXiv arXiv 2023
-
[31]
Driveworld: 4d pre- trained scene understanding via world models for autonomous driving,
C. Min, D. Zhao, L. Xiao, J. Zhao, X. Xu, Z. Zhu, L. Jin, J. Li, Y . Guo, J. Xing, L. Jing, Y . Nie, and B. Dai, “Driveworld: 4d pre- trained scene understanding via world models for autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 15 522–15 533
2024
-
[32]
Crt-fusion: Camera, radar, temporal fusion using motion information for 3d object detection,
J. Kim, M. Seong, and J. W. Choi, “Crt-fusion: Camera, radar, temporal fusion using motion information for 3d object detection,” inAdvances in Neural Information Processing Systems (NeurIPS), vol. 37, 2024, pp. 108 625–108 648
2024
-
[33]
Cross-modality knowledge distillation network for monocular 3d object detection,
Y . Hong, H. Dai, and Y . Ding, “Cross-modality knowledge distillation network for monocular 3d object detection,” inProceedings of the European Conference on Computer Vision (ECCV). Springer, 2022, pp. 87–104
2022
-
[34]
X-align: Cross-modal cross-view alignment for bird’s-eye-view segmentation,
S. Borse, M. Klingner, V . R. Kumar, H. Cai, A. Almuzairee, S. Yoga- mani, and F. Porikli, “X-align: Cross-modal cross-view alignment for bird’s-eye-view segmentation,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023, pp. 3287–3297
2023
-
[35]
Unifying voxel-based representation with transformer for 3d object detection,
Y . Li, Y . Chen, X. Qi, Z. Li, J. Sun, and J. Jia, “Unifying voxel-based representation with transformer for 3d object detection,” inAdvances in Neural Information Processing Systems (NeurIPS), vol. 35, 2022, pp. 18 442–18 455
2022
-
[36]
Masked au- toencoders are scalable vision learners,
K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked au- toencoders are scalable vision learners,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 16 000–16 009
2022
-
[37]
Point-bert: Pre-training 3d point cloud transformers with masked point modeling,
X. Yu, L. Tang, Y . Rao, T. Huang, J. Zhou, and J. Lu, “Point-bert: Pre-training 3d point cloud transformers with masked point modeling,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 19 313–19 322
2022
-
[38]
Masked autoencoders for point cloud self-supervised learning,
Y . Pang, W. Wang, F. E. Tay, W. Liu, Y . Tian, and L. Yuan, “Masked autoencoders for point cloud self-supervised learning,” inProceedings of the European Conference on Computer Vision (ECCV). Springer, 2022, pp. 604–621
2022
-
[39]
Exploring geometry-aware contrast and clustering harmonization for self-supervised 3d object detection,
H. Liang, C. Jiang, D. Feng, X. Chen, H. Xu, X. Liang, W. Zhang, Z. Li, and L. Van Gool, “Exploring geometry-aware contrast and clustering harmonization for self-supervised 3d object detection,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 3293–3302
2021
-
[40]
Visionpad: A vision-centric pre-training paradigm for autonomous driving,
H. Zhang, W. Zhou, Y . Zhu, X. Yan, J. Gao, D. Bai, Y . Cai, B. Liu, S. Cui, and Z. Li, “Visionpad: A vision-centric pre-training paradigm for autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2025, pp. 17 165–17 175
2025
-
[41]
Bootstrapping autonomous driving radars with self-supervised learn- ing,
Y . Hao, S. Madani, J. Guan, M. Alloulah, S. Gupta, and H. Hassanieh, “Bootstrapping autonomous driving radars with self-supervised learn- ing,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 15 012–15 023
2024
-
[42]
Self-supervised sparse sensor fusion for long range percep- tion,
E. Palladin, S. Brucker, F. Ghilotti, P. Narayanan, M. Bijelic, and F. Heide, “Self-supervised sparse sensor fusion for long range percep- tion,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025
2025
-
[43]
Multi-sensor fusion in automated driving: A survey,
Z. Wang, Y . Wu, and Q. Niu, “Multi-sensor fusion in automated driving: A survey,”IEEE Access, vol. 8, pp. 2847–2868, 2019
2019
-
[44]
Radar-camera fusion for object detection and semantic segmentation in autonomous driving: A comprehensive review,
S. Yao, R. Guan, X. Huang, Z. Li, X. Sha, Y . Yue, E. G. Lim, H. Seo, K. L. Man, X. Zhu, and Y . Yue, “Radar-camera fusion for object detection and semantic segmentation in autonomous driving: A comprehensive review,”IEEE Transactions on Intelligent Vehicles, vol. 9, no. 1, pp. 2094–2128, 2024
2094
-
[45]
Cramnet: Camera-radar fusion with ray-constrained cross-attention for robust 3d object detection,
J.-J. Hwang, H. Kretzschmar, J. Manela, S. Rafferty, N. Armstrong- Crews, T. Chen, and D. Anguelov, “Cramnet: Camera-radar fusion with ray-constrained cross-attention for robust 3d object detection,” in Proceedings of the European Conference on Computer Vision (ECCV). Springer, 2022
2022
-
[46]
Mvfusion: Multi-view 3d object detection with semantic-aligned radar and camera fusion,
Z. Wu, G. Chen, Y . Gan, L. Wang, and J. Pu, “Mvfusion: Multi-view 3d object detection with semantic-aligned radar and camera fusion,” in2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 2766–2773
2023
-
[47]
Ea-lss: Edge-aware lift-splat-shot framework for 3d bev object detection,
H. Hu, F. Wang, J. Su, Y . Wang, L. Hu, W. Fang, J. Xu, and Z. Zhang, “Ea-lss: Edge-aware lift-splat-shot framework for 3d bev object detection,”arXiv preprint arXiv:2303.17895, 2023
Pith/arXiv arXiv 2023
-
[48]
Rcbevdet: Radar-camera fusion in bird’s eye view for 3d object detection,
Z. Lin, Z. Liu, Z. Xia, X. Wang, Y . Wang, S. Qi, Y . Dong, N. Dong, L. Zhang, and C. Zhu, “Rcbevdet: Radar-camera fusion in bird’s eye view for 3d object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 14 928–14 937
2024
-
[49]
Sparc-ad: A baseline for radar-camera fusion in end-to-end autonomous driving,
P. Wolters, J. Gilg, T. Teepe, and G. Rigoll, “Sparc-ad: A baseline for radar-camera fusion in end-to-end autonomous driving,” inProceedings of the IEEE/CVF International Conference on Computer Vision Work- shops (ICCVW), October 2025, pp. 1831–1841
2025
-
[50]
Hermes: A unified self-driving world model for simultaneous 3d scene understanding and generation,
X. Zhou, D. Liang, S. Tu, X. Chen, Y . Ding, D. Zhang, F. Tan, H. Zhao, and X. Bai, “Hermes: A unified self-driving world model for simultaneous 3d scene understanding and generation,”arXiv preprint arXiv:2501.14729, 2025
Pith/arXiv arXiv 2025
-
[51]
Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free,
Z. Qiu, Z. Wang, B. Zheng, Z. Huang, K. Wen, S. Yang, R. Men, L. Yu, F. Huang, S. Huanget al., “Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free,” inAdvances in Neural Information Processing Systems (NeurIPS)
-
[52]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778
2016
-
[53]
Fcos3d: Fully convolutional one- stage monocular 3d object detection,
T. Wang, X. Zhu, J. Pang, and D. Lin, “Fcos3d: Fully convolutional one- stage monocular 3d object detection,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 913– 922
2021
-
[54]
Feature pyramid networks for object detection,
T.-Y . Lin, P. Doll´ar, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2117–2125
2017
-
[55]
Pointpillars: Fast encoders for object detection from point clouds,
A. H. Lang, S. V ora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom, “Pointpillars: Fast encoders for object detection from point clouds,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019
2019
-
[56]
Bevformer: learning bird’s-eye-view representation from lidar-camera JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 17 via spatiotemporal transformers,
Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Q. Yu, and J. Dai, “Bevformer: learning bird’s-eye-view representation from lidar-camera JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 17 via spatiotemporal transformers,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 3, pp. 2020–2036, 2024
2021
-
[57]
nuplan: A closed-loop ml- based planning benchmark for autonomous vehicles,
H. Caesar, J. Kabzan, K. S. Tan, W. K. Fong, E. Wolff, A. Lang, L. Fletcher, O. Beijbom, and S. Omari, “nuplan: A closed-loop ml- based planning benchmark for autonomous vehicles,”arXiv preprint arXiv:2106.11810, 2021
Pith/arXiv arXiv 2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.