Pith. sign in

Paper Citation Record · LEDGER

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining

As of 10 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.04541.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04541 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T17:41:06.399906Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 062fba49-2154-4252-83a3-581a32873d73 · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining nuscenes: A multi- modal dataset for autonomous driving,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:928518ecf92703ec8b8623ed8b45aefbdc554f591a91749939b5406e0fda51a7

Observation 7ecf7c28-bd24-49b1-8165-d411f97e777a · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Scalability in perception for autonomous driving: Waymo open dataset,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:5e857f6697645c002bb0088fe0a7bd50f8b3e19556b64e6e074e319a213db9bc

Observation e751352a-7adf-40a6-8a8e-0edc9b70dfb5 · outbound

This paper cites Planning-oriented autonomous driving,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Planning-oriented autonomous driving,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:38a7b68335448914587b88890203a764f2fc58d085db870d20b5fd9eef48e65b

Observation 1564b81e-4c0c-400d-9309-016e581a3306 · outbound

This paper cites Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:52533ac8b556a05f2fe2636f6051e6677b96d8ff733123b0aa711be6d74e4a97

Observation cb2a1b19-dbb0-4122-b56e-3ea5ffb96dc5 · outbound

This paper cites Deep learning-based perception systems for autonomous driving: A comprehensive survey,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Deep learning-based perception systems for autonomous driving: A comprehensive survey,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:7ac7e57af1816a1cd26e37b9cb9f154fb334dd113e38ae00c136c0f5e78935aa

Observation 3e2c7ddb-e386-4aa1-a736-d739c4624cf5 · outbound

This paper cites Towards deep radar perception for autonomous driving: Datasets, methods, and challenges,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Towards deep radar perception for autonomous driving: Datasets, methods, and challenges,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:dfcd6254a3ba43baec282e9ae4798c9247d69442e2cfd5590d897e4553fd0fb3

Observation 4af08331-9973-44cd-afc5-64b636825662 · outbound

This paper cites Crkd: Enhanced camera-radar object detection with cross-modality knowledge distillation,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Crkd: Enhanced camera-radar object detection with cross-modality knowledge distillation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:e7233dc4dc696470ea3d9e4ed4917ebc1b1947920f35c4a578d1e4b8411d5e95

Observation fcf1dbb4-ef65-4bee-a573-a9f8fa714553 · outbound

This paper cites Multi-modal 3d object detection in autonomous driving: a survey,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Multi-modal 3d object detection in autonomous driving: a survey,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:ed3afb3bdc927f53e2bf90a8cf35d619ece768b459e8d398ff90d846ca6784ed

Observation 934cff3c-d8f9-4df0-9063-6e9481f8daee · outbound

This paper cites Vision meets robotics: The kitti dataset,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Vision meets robotics: The kitti dataset,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:b582e97a8457f9c51155e44f538b16d8ac1874520f03349860fb193e6f6a6818

Observation 7fc45cef-bd46-4528-a107-612ff3885da5 · outbound

This paper cites Unifying voxel-based representation with transformer for 3d object detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Unifying voxel-based representation with transformer for 3d object detection,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:c6f2b1c0272c67b33484bd926cad586ccf9c444d8e984c434c9a3be6d4e66daa

Observation 0a0b34e7-1b7f-4baa-877f-4f9515664ecf · outbound

This paper cites Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:5274646075238ae3ac66df3f81f3f57a96981573941a503350710dd57d4e168d

Observation ac29c394-c801-4c6a-9e0d-fa7f5075a35d · outbound

This paper cites Futr3d: A unified sensor fusion framework for 3d detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Futr3d: A unified sensor fusion framework for 3d detection,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:ad6c6306cbd62d0df6fddadebbb8cb7854c3a4d5353cf010f473488c693a16c2

Observation e0d4e467-6f96-413a-8252-193e487885c0 · outbound

This paper cites Radar and Camera Fusion for Object Detection and Tracking: A Comprehensive Survey.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Radar and Camera Fusion for Object Detection and Tracking: A Comprehensive Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:b566c24d094ff48a957d43d72bd0d9aaa0659466094f3e0a7d1bc77f4e4cd43e

Observation cebbb4c0-49ea-43e3-b4e2-1bff689c94d5 · outbound

This paper cites Centerfusion: Center-based radar and camera fusion for 3d object detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Centerfusion: Center-based radar and camera fusion for 3d object detection,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:81d590403b03e22b26078ae98cd2b9bfe11cfaac6b3a5f2958239d3153770bcd

Observation 91ed6553-2435-4bef-b91c-60a6528c5093 · outbound

This paper cites BEV-Guided Multi-Modality Fusion for Driving Perception,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining BEV-Guided Multi-Modality Fusion for Driving Perception,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:4b9844b00a21d50d1724317eea1ac26e62c92abc728625697282d92435370013

Observation 05fc51d7-13d8-450b-8862-2b020a608998 · outbound

This paper cites Crn: Camera radar net for accurate, robust, efficient 3d perception,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Crn: Camera radar net for accurate, robust, efficient 3d perception,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:ad4688978625984494e665f0ee77351a0f48c01ad064e3134b60d96deee1efc8

Observation 8eb52c3c-7d2f-4fa3-b1d9-d7514f081be3 · outbound

This paper cites Bevcar: Camera-radar fusion for bev map and object segmentation,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Bevcar: Camera-radar fusion for bev map and object segmentation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:11000973ba13eee84b902c54128924c1894cbd3309c0792e0278e2c3faaed9d4

Observation 00c05714-2f14-47c4-ae52-4f2d43ce85ff · outbound

This paper cites Licrocc: Teach radar for accurate semantic occupancy prediction using lidar and camera,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Licrocc: Teach radar for accurate semantic occupancy prediction using lidar and camera,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:901796344597337529396e357ad7bd0408fa7c914e5d15a7f5e74bf193c61a49

Observation 0800e28f-1b25-4abc-81db-b4d3728d8380 · outbound

This paper cites Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:ce9b409139a74ce0027336708c7edd83788739d44b32d34bf0d5eb97ab69911a

Observation 304d7ce1-644c-4537-a69a-4da10459b167 · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:6d18f5ad90f3279af6c47344eaeac1baaef8d655cc5d317b45fe3fe643944ad2

Observation 322f5d86-c684-4336-8869-f8b90eacf7b6 · outbound

This paper cites Gd-mae: generative decoder for mae pre-training on lidar point clouds,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Gd-mae: generative decoder for mae pre-training on lidar point clouds,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:f4c0b297aa4ce8786ce2080d331ba10646d76a389fefaa55655aeb12255dc8ed

Observation bba2c985-3fbe-442e-9e91-0468e4c85a60 · outbound

This paper cites Unipad: A universal pre-training paradigm for autonomous driving,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Unipad: A universal pre-training paradigm for autonomous driving,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:0108faf6191d70a50c3ca8adbb7c5ced2d7ad7228c7ce49170080735380e0403

Observation 4d58cf56-2753-4e0f-baef-b3eb90f5f726 · outbound

This paper cites Visual point cloud forecasting enables scalable autonomous driving,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Visual point cloud forecasting enables scalable autonomous driving,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:55fb434dad2ba0afba59ab089063a9dc5f58b13b8bd3a8b7f364444947f0c0ff

Observation 99b02e2e-4e30-4fe4-b256-2aa84fdd1ce2 · outbound

This paper cites Forging spatial intelligence: A roadmap of multi-modal data pre- training for autonomous systems,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Forging spatial intelligence: A roadmap of multi-modal data pre- training for autonomous systems,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:dd81a40189614693ea4033607cc604a14fda46cdbb1b1adf416bd290e239d602

Observation c6898dc8-6f24-495a-a8b0-c1e20098a62c · outbound

This paper cites Masked autoencoder for self-supervised pre-training on lidar point clouds,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Masked autoencoder for self-supervised pre-training on lidar point clouds,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:814a0d0a1e04fd9e289ea2ee968170bed93ed381b5a087c9ceb925c24bf126d2

Observation 0006f65d-dda1-4cf4-8d34-f402a52157d9 · outbound

This paper cites Bev-mae: Bird’s eye view masked autoencoders for point cloud pre-training in autonomous driving scenarios,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Bev-mae: Bird’s eye view masked autoencoders for point cloud pre-training in autonomous driving scenarios,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:f4758995523736671cb9cef80b1886f5a9aac24800330ef9a09cc130ad73e36e

Observation cd5233cc-fe29-4734-a119-42b6ef0e4f4d · outbound

This paper cites Is pseudo- lidar needed for monocular 3d object detection?.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Is pseudo- lidar needed for monocular 3d object detection?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:91fa89c0f7254c2bf6b6d588bb7c92c7672f421bdf926f4354d52c460b9d5606

Observation 39df89d9-ffd3-4f4f-ad55-a2d23c765896 · outbound

This paper cites Pimae: Point cloud and image interactive masked autoencoders for 3d object detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Pimae: Point cloud and image interactive masked autoencoders for 3d object detection,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:7279ef54bf0eee9845b3ba8a0ec564919e2bb2fc4216e7bd7e2f6e33a0b0e842

Observation 2324b04e-cca9-4898-ac0a-ecbfbe070659 · outbound

This paper cites Point cloud forecasting as a proxy for 4d occupancy forecasting,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Point cloud forecasting as a proxy for 4d occupancy forecasting,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:753fee5ed7b560e5b3a6154a7d26c1f09bd2dbbb35fa00d2a4f5e31957501d0e

Observation db276bf3-a01a-4839-a03c-3a972a2ae897 · outbound

This paper cites UniWorld: Autonomous Driving Pre-training via World Models.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining UniWorld: Autonomous Driving Pre-training via World Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:6659be5af00bf04e0a8bed5a092f09262b0d2a7294e0078d7fd8bdc49e8fad49

Observation 964ab74a-477c-4f8b-90ce-6a789f4a4b70 · outbound

This paper cites Driveworld: 4d pre- trained scene understanding via world models for autonomous driving,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Driveworld: 4d pre- trained scene understanding via world models for autonomous driving,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:cf54b62b119e859f42f2f09e7060f49ce540943718f62517616ead04c71b8763

Observation 72006e77-9a10-44b5-93c6-7e6a66738a99 · outbound

This paper cites Crt-fusion: Camera, radar, temporal fusion using motion information for 3d object detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Crt-fusion: Camera, radar, temporal fusion using motion information for 3d object detection,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:ef6e1f59022a78638e7bf9f78b9f970ef924b3736805eb446b19113510227d83

Observation 5879092a-c1fe-49cc-bba2-38b39952cabe · outbound

This paper cites Cross-modality knowledge distillation network for monocular 3d object detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Cross-modality knowledge distillation network for monocular 3d object detection,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:ef3e28ca5862ca32f2876c2d42defbf7347f04da4a02a741cf2bd40264657547

Observation 00c5814f-4724-4f08-80c1-d6b99386c498 · outbound

This paper cites X-align: Cross-modal cross-view alignment for bird’s-eye-view segmentation,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining X-align: Cross-modal cross-view alignment for bird’s-eye-view segmentation,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:efa16c231ac0d7802a584af2d5cb9209a988bf97f66c0615aadca2c51b177545

Observation e81bfd8c-db19-4058-84e6-365b16beb067 · outbound

This paper cites Unifying voxel-based representation with transformer for 3d object detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Unifying voxel-based representation with transformer for 3d object detection,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:68c3c0c038bbbf04c45d0f69a39e93f6c0834fc95496e7aecbcc533b4de04958

Observation e85eebda-ed7d-4620-9866-de419687b7aa · outbound

This paper cites Masked au- toencoders are scalable vision learners,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Masked au- toencoders are scalable vision learners,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:7511e8cfcedd3d86085b6020817e5e15873e42f096ab28804efa04b78189e3c5

Observation bf3b9b43-17d9-4164-a122-d06db045b272 · outbound

This paper cites Point-bert: Pre-training 3d point cloud transformers with masked point modeling,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Point-bert: Pre-training 3d point cloud transformers with masked point modeling,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:e18695afbaaafd0ed08000814d7a260d578904b4360c256e0d91ab2531b70c7f

Observation 858884ce-d266-417c-b1e8-613a285df5d9 · outbound

This paper cites Masked autoencoders for point cloud self-supervised learning,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Masked autoencoders for point cloud self-supervised learning,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:1bfb8bcd03e4a71cdb66a48ed27be0855b475195a6576e1057e0e1d198a6a059

Observation 335f2fe5-99c6-4d37-ba35-5f8495595ff1 · outbound

This paper cites Exploring geometry-aware contrast and clustering harmonization for self-supervised 3d object detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Exploring geometry-aware contrast and clustering harmonization for self-supervised 3d object detection,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:faa3c2cfafb72c8908217ac2cd074204628ff1b48ad648cf02dc6a84a626c70a

Observation 2e94a117-8695-4af8-ab03-802e1e9a4554 · outbound

This paper cites Visionpad: A vision-centric pre-training paradigm for autonomous driving,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Visionpad: A vision-centric pre-training paradigm for autonomous driving,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:efd5b0ba0464804ccee53bd569c2add0513c3fe4ff7241e634138f38a7d1dd20

Observation 28933d85-e0ad-426d-bae8-c298d908c191 · outbound

This paper cites Bootstrapping autonomous driving radars with self-supervised learn- ing,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Bootstrapping autonomous driving radars with self-supervised learn- ing,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:fd4d5bee75611d8e5419f1ce11c961174def208c536d12378e72748d44e4669e

Observation 807615f1-767c-4ceb-8282-76bb6c1003ea · outbound

This paper cites Self-supervised sparse sensor fusion for long range percep- tion,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Self-supervised sparse sensor fusion for long range percep- tion,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:ebb69ab6a30d7185bbfa4d8d9a10135ffea02a00a1e2bba6d38e344b90b16141

Observation d7015175-c733-4c9a-9618-126bff329aed · outbound

This paper cites Multi-sensor fusion in automated driving: A survey,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Multi-sensor fusion in automated driving: A survey,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:6145f65857151d0e25b0461d0ac3bebd5c05ebffe8eb3c2f0f2925feb3257c52

Observation 16b50faf-2f58-4de9-914b-88ca657d651c · outbound

This paper cites Radar-camera fusion for object detection and semantic segmentation in autonomous driving: A comprehensive review,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Radar-camera fusion for object detection and semantic segmentation in autonomous driving: A comprehensive review,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:9da59360b47ad7003b15ae9638442c06ace089a52a17c7977ef293e0bfb938f9

Observation 738ddddb-499d-4da1-b27d-e84f6200fbef · outbound

This paper cites Cramnet: Camera-radar fusion with ray-constrained cross-attention for robust 3d object detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Cramnet: Camera-radar fusion with ray-constrained cross-attention for robust 3d object detection,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:214c9daa577ecaf2142f0bc44f81429587cb2e598e914092db2eee73206362e9

Observation d961c668-6e69-4e39-98f3-7095d73fff3f · outbound

This paper cites Mvfusion: Multi-view 3d object detection with semantic-aligned radar and camera fusion,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Mvfusion: Multi-view 3d object detection with semantic-aligned radar and camera fusion,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:9a7939d0193c11bc019abab7eab06e49de961067b27970685a980fa12d935b8b

Observation d29ed390-1250-4158-97ab-2fb5bea753df · outbound

This paper cites EA-LSS: Edge-aware Lift-splat-shot Framework for 3D BEV Object Detection.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining EA-LSS: Edge-aware Lift-splat-shot Framework for 3D BEV Object Detection

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:68bf013e647b4a6cd4bd50412cd7c0c0c34ea550200fabe43103c5e541125be7

Observation c4536b1d-f098-4986-8599-9a7a7c43c7ce · outbound

This paper cites Rcbevdet: Radar-camera fusion in bird’s eye view for 3d object detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Rcbevdet: Radar-camera fusion in bird’s eye view for 3d object detection,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:6b8f1a24f5ebf25d9f6013d8326e78d298e60b3e679b31192a5d514588bb4245

Observation a09c2691-f83a-41ed-aad7-ab5317dbfa52 · outbound

This paper cites Sparc-ad: A baseline for radar-camera fusion in end-to-end autonomous driving,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Sparc-ad: A baseline for radar-camera fusion in end-to-end autonomous driving,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:95a595770822482045e27428ab32130ee93e03b6d313b42818b8b770d8a1c712

Observation adf4a1c5-6851-4433-b54b-8de23a192288 · outbound

This paper cites HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:1d6fde9d9b3459ab77a14967e1db864b900c4fc6f3660ba5a93d9201762b3497

Observation 69a3f444-99c5-4e16-83bc-76bdc7d437b3 · outbound

This paper cites Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:9c9191095971e8c77a54e85413149f03da4e0490b435a5a4614d4def9b4890cd

Observation cc6822de-a09a-403d-bcdd-e6c34f84ac4b · outbound

This paper cites Deep residual learning for image recognition,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Deep residual learning for image recognition,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:da0f6f1ec4003486d0f44aeec93cf2cb5a654a6a9e0066f1c64b19fbcdaacea7

Observation 01de08d9-d58c-4beb-b7ed-b19f3705b91f · outbound

This paper cites Fcos3d: Fully convolutional one- stage monocular 3d object detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Fcos3d: Fully convolutional one- stage monocular 3d object detection,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:0516ca4220f099930ced93e698ec04d4cb1672f7347217bbb10c63e28e7a7621

Observation db2f2d82-e78a-4425-bcee-6e01fdc81182 · outbound

This paper cites Feature pyramid networks for object detection,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Feature pyramid networks for object detection,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:c8dd847976794293281ce36d677e7a6899b9200cf9fdb5ec2791ccc704cb4cf1

Observation b9176a1c-c29d-4909-bd7f-019569d38310 · outbound

This paper cites Pointpillars: Fast encoders for object detection from point clouds,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Pointpillars: Fast encoders for object detection from point clouds,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:606890a90ec9b530bc1b456578cadd9101b833c2aed4d53d2789f0d72b6806a9

Observation 6d3aa147-eb36-428a-b164-6ce086c0082b · outbound

This paper cites Bevformer: learning bird’s-eye-view representation from lidar-camera JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 17 via spatiotemporal transformers,.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining Bevformer: learning bird’s-eye-view representation from lidar-camera JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 17 via spatiotemporal transformers,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:51d309a4b1acf25e9257453e8e193ce22a9649f7694753dd1f2120dc5dc7d324

Observation d98dfb6f-0704-49f5-928d-b13aab3f06ca · outbound

This paper cites NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles.

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-11T17:41:06.399906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:41:06.399906Z digest=sha256:5a8ddb037bea4dbf79f35b1a41d3f44e75ae53098c20c0ba4af0fb0ce835a17c

Pith citing papers

No inbound Pith citation observations are available.