Pith. sign in

Paper Citation Record · LEDGER

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

As of 19 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 12 inbound Pith citation observations for arXiv:2506.07497.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07497 v4

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:37:53.713132Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:57:08.231458Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:28:52.737750Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a48e3d34-be12-4c25-8957-eee046ed32f6 · outbound

This paper cites Qwen Technical Report.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.480480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.480480Z digest=sha256:a892efb47103c6320e4f932d78727b2bfaf71163e5a63473dce628e6e15cd67d

Observation 5cf9c107-4208-4985-8261-389c8dbe4bda · outbound

This paper cites Transfusion: Robust lidar-camera fusion for 3d object detection with transformers.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Transfusion: Robust lidar-camera fusion for 3d object detection with transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.501574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.486277Z digest=sha256:aff72cc6e8a883d650f7c97143e3b3b464a220beb681dea4c52bca02cca1966d

Observation 951a427e-9b2f-4954-aef0-924c0dafc292 · outbound

This paper cites Camera-lidar inte- gration: Probabilistic sensor fusion for semantic mapping.IEEE Transactions on Intelligent Transportation Systems, 23(7):7637–7652, 2021.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Camera-lidar inte- gration: Probabilistic sensor fusion for semantic mapping.IEEE Transactions on Intelligent Transportation Systems, 23(7):7637–7652, 2021

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.487360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.490987Z digest=sha256:e3565cb1bd2dd6e547667463de3ee3aaf488b38040c6663da106ddf2cf82f33a

Observation 4a22c19d-e3b2-4e1d-825a-02fd06bffdac · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency nuscenes: A multimodal dataset for autonomous driving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.495865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.495865Z digest=sha256:d78b8b1da78138baf7c1c9de1ecdd8d4141863f12863793c61d6bd5090d28b4c

Observation cda1960b-60af-4262-b5bc-26c8006656af · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.500583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.500583Z digest=sha256:a042604c5ae74e74c4788050b279446f14a17955a50870f0e72a81056f8cfd32

Observation a21c1d5d-5742-4009-8023-4dd70a137bf5 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.505322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.505322Z digest=sha256:c687a77360877bdb6aeda13633f884d822c541ed406e6451b236f296c99eb273

Observation 10daedaf-dc4b-4d75-8b65-b0d607bf7e60 · outbound

This paper cites Trafficgen: Learning to generate diverse and realistic traffic scenarios.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Trafficgen: Learning to generate diverse and realistic traffic scenarios

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.445835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.510195Z digest=sha256:6bd77ee10d7540a90ea430f36b660edca5ec517199b2f67450a2a20f5b4369ed

Observation da919632-7ffd-434f-b87f-79a35b2122bb · outbound

This paper cites MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.514590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.514590Z digest=sha256:4ddad40463272acbd0a8e4444a750a32389fbfe1a29a180c69a74d52bc71a44a

Observation b2be098a-fc3e-4f7b-be2c-1cabcc112221 · outbound

This paper cites MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.519590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.519590Z digest=sha256:3b965f18dce03da44dbd36ab73ece2675117ea0595574ed5b27ff9edb753ab8f

Observation e36da3c9-0cb7-4ca7-adad-16d9513f481a · outbound

This paper cites MagicDrive: Street View Generation with Diverse 3D Geometry Control.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency MagicDrive: Street View Generation with Diverse 3D Geometry Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.524532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.524532Z digest=sha256:112bad6a57c8a37c910324cd13d2bd8191c4067702b7389eda2fe11496031114

Observation 973f4d24-146c-4556-84f1-cc72adaa1088 · outbound

This paper cites Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.529341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.529341Z digest=sha256:e28bdef8c5ce690bd1fec0973fb01e8892159ffe7ccac672bf25f0e9c47c30e0

Observation c2f7d725-15d0-43c3-b676-7d7fdc604d78 · outbound

This paper cites Vision meets robotics: The kitti dataset.The international journal of robotics research, 32(11):1231–1237, 2013.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Vision meets robotics: The kitti dataset.The international journal of robotics research, 32(11):1231–1237, 2013

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.534200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.534200Z digest=sha256:f55088dae5109fa520ac0ba7b68392a040c4e610e4d428de9a7103f10dc3e88e

Observation 8de2b146-c2f5-4380-bf0a-8604a0b1059a · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.538438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.538438Z digest=sha256:806d5a8901999a0d1a6eff1879f11e88cad22fe37d81c4f46d2bfdf23896ff2d

Observation 8b6be0d0-9fd9-4173-ba29-d76ec2c4ef50 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.542834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.542834Z digest=sha256:0cf45420f19313cea6a863832c34e59923480067089f73f74f99c1bfaf47284e

Observation 913b628b-250c-4159-8935-e1ae5e0e3e78 · outbound

This paper cites CoGen: 3D Consistent Video Generation via Adaptive Conditioning for Autonomous Driving.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency CoGen: 3D Consistent Video Generation via Adaptive Conditioning for Autonomous Driving

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.547745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.547745Z digest=sha256:676dc04cedff34ab963316c9dcf8953abd14a845baff51121a9e398c91ce0311

Observation 02560e80-bbaa-447a-bb00-5ed5cdf6d0cd · outbound

This paper cites DiVE: DiT-based Video Generation with Enhanced Control.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency DiVE: DiT-based Video Generation with Enhanced Control

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.552383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.552383Z digest=sha256:c844e02dff6bff2f37b589dcdcbfe51fd5e113cb3fc7fad5ee37878f23f24722

Observation 8e9d4c05-0c91-43f8-8ef7-9336d47a01dd · outbound

This paper cites Point cloud forecasting as a proxy for 4d occupancy forecasting.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Point cloud forecasting as a proxy for 4d occupancy forecasting

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.556773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.556773Z digest=sha256:29a271cdd5420de081a04964ca602c70e97869f7386e6706d04bd068ebee08ac

Observation f2d47840-adea-440e-be8f-8bc4740ac7c2 · outbound

This paper cites UniScene: Unified Occupancy-centric Driving Scene Generation.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency UniScene: Unified Occupancy-centric Driving Scene Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.561406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.561406Z digest=sha256:074ddab0c9e861a21d258a3d63ccabb88611c859667f3c7915d0de69b4ee4c12

Observation 3ae218aa-f448-430e-afc0-e5b031fe2903 · outbound

This paper cites Deepfusion: Lidar-camera deep fusion for multi-modal 3d object detection.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Deepfusion: Lidar-camera deep fusion for multi-modal 3d object detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.403234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.565899Z digest=sha256:2d915e0f3d48aac2d39eda7c054360abe5798b2f0844e6615adaa4ce9150ba10

Observation 97817166-7d39-4b11-a189-dc96bde9cd80 · outbound

This paper cites Bevfusion: A simple and robust lidar-camera fusion framework.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Bevfusion: A simple and robust lidar-camera fusion framework

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.570093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.570093Z digest=sha256:e6216862edafe3fe57fd275d3c62c8aefb6419e33aa92a771f30f497564d4a8c

Observation 7bd90cf4-5e54-4f8d-b008-e13244d031a6 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.575430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.575430Z digest=sha256:4594203250c9ddc819a2ce0e55437b8da97f9cc1fa4326cd4d1432c24f71c131

Observation 9737325c-0413-4e95-b419-23730d09000f · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Swin transformer: Hierarchical vision transformer using shifted windows

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.579945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.579945Z digest=sha256:8c54a27d57fbed72e0d3520e0568ed61c349b3a18e8e97158b08a575f4629a01

Observation c857f0af-00ae-45cb-baf7-11f143b1bac9 · outbound

This paper cites Wovogen: World volume-aware diffusion for controllable multi-camera driving scene generation.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Wovogen: World volume-aware diffusion for controllable multi-camera driving scene generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.360663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.584599Z digest=sha256:0924ba0d45ec728ad5b59291ccde76d6b4e19ea8677e029d78f1adfecf3a0477

Observation 2915b128-9a06-49ca-a5d1-64de3241a0e5 · outbound

This paper cites Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.589125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.589125Z digest=sha256:fa4d5ddd8723d45403c42c6979d92c59c8a25b3355ceeb72ec84218145431ff9

Observation 1f7dda2d-9abb-4652-9afd-a3f97aba60be · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.593752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.593752Z digest=sha256:b7b387ecdb12e242472e593b7b345aba55b61927a00adbff8d4e0883eb007000

Observation 21ceec62-ab6d-467b-bedf-2291c9799b41 · outbound

This paper cites Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.337469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.598487Z digest=sha256:5ed743aa984bc27435376638d4b25b850abad4f2a881c8249dabd227105de7af

Observation 64bdebe8-b95b-4f1d-ae90-fe2176caec3c · outbound

This paper cites Scenario diffusion: Controllable driving scenario generation with diffusion.Advances in Neural Information Processing Systems, 36:68873–68894, 2023.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Scenario diffusion: Controllable driving scenario generation with diffusion.Advances in Neural Information Processing Systems, 36:68873–68894, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.322428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.602729Z digest=sha256:77e45b3c7f0d38a59d40ded25f400ef6f2ce36e4050803f248a34eb2ea69e53d

Observation 22009dd1-b790-4965-b85b-7a378d5fc560 · outbound

This paper cites Pointnet: Deep learning on point sets for 3d classification and segmentation.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Pointnet: Deep learning on point sets for 3d classification and segmentation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.607077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.607077Z digest=sha256:988cc877d9ae10297c4a85ac850f46acd733e011ad95d1db5653841a9e908c9f

Observation 7cbb251d-e85a-437c-8743-6dbdd8b3547e · outbound

This paper cites Towards realistic scene generation with lidar diffusion models.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Towards realistic scene generation with lidar diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.298762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.611655Z digest=sha256:87cfab5cdcad4d6e4e8845d76bc50ea6612257ee4e29a807bcf533e5a4cf63f1

Observation a17f0717-87f3-446c-af2d-92414bfc62b4 · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Scalability in perception for autonomous driving: Waymo open dataset

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.615860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.615860Z digest=sha256:c0ca63568a66e3b32ee329921abbd5070955f191a727aa28b328f11a1c27c7c1

Observation 875c1f86-1877-4056-8221-4cb9cd0c7973 · outbound

This paper cites Drivescenegen: Generating diverse and realistic driving scenarios from scratch.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Drivescenegen: Generating diverse and realistic driving scenarios from scratch

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.275109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.620049Z digest=sha256:dab714a8ae2012f564f2176b182ed742992fed402d1607c13e66bdddbfe73147

Observation 2edd54ca-b8e8-4128-b6af-f50d2f8d4774 · outbound

This paper cites Street-view image generation from a bird’s-eye view layout.IEEE Robotics and Automation Letters, 2024.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Street-view image generation from a bird’s-eye view layout.IEEE Robotics and Automation Letters, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.259721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.628534Z digest=sha256:ed00501ba4cf251b7cdddbcc3aa2a5709a5af3ee9c972551a15dd637f96e63f6

Observation 9391a233-0bf3-48f2-a4ce-b5a7b6ebffa3 · outbound

This paper cites Fvd: A new metric for video generation.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Fvd: A new metric for video generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.633174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.633174Z digest=sha256:eca35ece673652f0bb38801eaa94abeb9555296ea37e203f831ca3f27bf428e0

Observation cf12d541-a77e-45d6-b49d-d0f521364c61 · outbound

This paper cites MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.638574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.638574Z digest=sha256:0e93d8a4964e983d2fcb2ea04629f9339ebb2f8fad1957c3c65caca430126429

Observation 4b1ba8b9-7c0f-4cd7-a855-94eeba645fe7 · outbound

This paper cites Drive- dreamer: Towards real-world-drive world models for autonomous driving.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Drive- dreamer: Towards real-world-drive world models for autonomous driving

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.644426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.644426Z digest=sha256:9200ef0f36731f332dc07c876af97cfb76b69d0f65851056c9292f2a27fd9839

Observation 8225b22e-44d9-4c4d-8b80-d6b4a97f9960 · outbound

This paper cites Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.649011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.649011Z digest=sha256:2bff76ecfa51f2a4626d68d7213cc2e83452d6722286b0c45022da91f732d6a5

Observation 0a9488ad-af98-47a8-9014-7efb31e322e6 · outbound

This paper cites Lidenerf: Neural radiance field reconstruction with depth prior provided by lidar point cloud.ISPRS Journal of Photogrammetry and Remote Sensing, 208:296–307, 2024.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Lidenerf: Neural radiance field reconstruction with depth prior provided by lidar point cloud.ISPRS Journal of Photogrammetry and Remote Sensing, 208:296–307, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.217316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.653249Z digest=sha256:957ab93a7ed8539a3b2eb717297a9f1b1416ed4ee5be7ee330c646b7672325fa

Observation 2746e56b-f07c-4aa8-80ce-5d50bf531a20 · outbound

This paper cites Panacea: Panoramic and controllable video generation for autonomous driving.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Panacea: Panoramic and controllable video generation for autonomous driving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.657642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.657642Z digest=sha256:c07e1d1d6bb7dee8a515739d1614e1411b480512b0c22e8d67c0646f3d66acee

Observation f7c8a747-501a-4309-9bef-d5326deff41e · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.662034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.662034Z digest=sha256:100fac41cad61e316f34091b8100baf3ec2eb58037515c96ebabc357fc45236b

Observation 041e5865-e3c7-46e1-b6ff-e3b6f39493a9 · outbound

This paper cites Driver lane change intention recognition based on attention enhanced residual-mbi-lstm network.IEEE Access, 10:58050– 58061, 2022.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Driver lane change intention recognition based on attention enhanced residual-mbi-lstm network.IEEE Access, 10:58050– 58061, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.184278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.666347Z digest=sha256:0fb16a82e522b22289424e53e9e59e1885f8ca9e8177dc721fcfbe52782499cd

Observation 0b07300e-7184-40d3-8449-2acb762681fe · outbound

This paper cites Point-nerf: Point-based neural radiance fields.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Point-nerf: Point-based neural radiance fields

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.671416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.671416Z digest=sha256:85de58f94f498debdc48cd6c52d4fca263bc12f86b25970ff56b3a359d919fdb

Observation 9f18b92f-e605-4921-9927-d8ac85237541 · outbound

This paper cites BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency BEVControl: Accurately Controlling Street-view Elements with Multi-perspective Consistency via BEV Sketch Layout

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.675805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.675805Z digest=sha256:b08689478f1301b034824012eda3a02871417aaedad5c2ea00f55eeff04b3ca2

Observation 1639931f-30c7-4fe4-aa52-8d6268b93b7e · outbound

This paper cites Visual point cloud forecasting enables scalable autonomous driving.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Visual point cloud forecasting enables scalable autonomous driving

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.680568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.680568Z digest=sha256:ff44329a54d2fa323e19e20294c04cfcf342c61ab11ee96b7256ec2c16cb7c89

Observation 4f38f51f-333a-43ae-b05d-84c017338b22 · outbound

This paper cites Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.685162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.685162Z digest=sha256:6a91f27883475841639d1941d1027ac357e58e328f7095d8496c4af899816bcf

Observation 99a95cf3-3760-4543-bbcf-64ba67862d16 · outbound

This paper cites BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.689840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.689840Z digest=sha256:a635ba2688e029eeb719bfa8849219c53e1bd118fa3e26b422b6b00fbff6cedc

Observation 469a0893-7342-48a4-803a-2ce7a1c2b7b0 · outbound

This paper cites Drivedreamer-2: Llm-enhanced world models for diverse driving video gen- eration.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Drivedreamer-2: Llm-enhanced world models for diverse driving video gen- eration

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.151020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.694355Z digest=sha256:6ca3ee810c829d7015e7de33ba1b91e30bc904d94ce0b9b5d3b71bdba5ef10f5

Observation aa96d767-5e76-4214-afd4-77db0d641785 · outbound

This paper cites HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.698728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.698728Z digest=sha256:274264d881979c084af5ff48c1b04b5c4b97d3024f9d80d5a95bb0dd3d865959

Observation 580001f9-4d2b-467d-aaa4-6e5c54160ed1 · outbound

This paper cites V oxelnet: End-to-end learning for point cloud based 3d object detection.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency V oxelnet: End-to-end learning for point cloud based 3d object detection

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.703572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.703572Z digest=sha256:ff7b1ff458ed80ac9f4565e332e5bfbc1917cbdc8a767321e2716b97fc7ffe8e

Observation b4964914-1372-4e3f-a0b4-e9c6788d3f3b · outbound

This paper cites Lidardm: Generative lidar simulation in a generated world.arXiv preprint arXiv:2404.02903, 2024.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Lidardm: Generative lidar simulation in a generated world.arXiv preprint arXiv:2404.02903, 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:53.708796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:53.708796Z digest=sha256:763adc30d07512c775d1342b6581f232888b31d0faba8c077676c4fd565bbbd1

Observation 883a735b-0381-49e1-9e5b-1cb44fe15c5b · outbound

This paper cites Daytime",.

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency Daytime",

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:54.125711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:37:53.713132Z digest=sha256:5c8880a822cff05a9a250682a81bc28710fa50ddfd2ad11e4e8c1c89b7cac6d5

Pith citing papers

Observation 1f4e4ccf-3a5d-48c7-bd29-26a10afca9d9 · inbound

OmniNWM: Omniscient Driving Navigation World Models cites this paper.

OmniNWM: Omniscient Driving Navigation World Models Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T08:57:08.231458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:57:08.231458Z digest=sha256:eb620c78ef1870137d42d41346798dff0e7770da9144733af1af685bfc43609b

Observation bc0aecb1-ed81-4793-ae62-f216b225084a · inbound

A Survey on the Applications of Generative Artificial Intelligence in Automated Driving Systems Test Scenario Generation Methods cites this paper.

A Survey on the Applications of Generative Artificial Intelligence in Automated Driving Systems Test Scenario Generation Methods Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T15:54:39.583328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:54:39.583328Z digest=sha256:9daad9e47a0d4496b45cd618ffbe62c5e6ef6ad17382c0c1c364d4fe5d85a84c

Observation 37eb6db4-a748-4454-8e71-c49ba98fabb7 · inbound

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World cites this paper.

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:38:21.112262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T19:34:39.518649Z digest=sha256:56a8bc10716b4415e6e9db51e73b249f1687e3d68c95dde6e99922bfd8eb4a04

Observation 41783d20-7bf4-4541-b45e-20fb32be0b90 · inbound

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models cites this paper.

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T20:36:09.847931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:36:09.847931Z digest=sha256:2dd78f87d8569a3d6001a07e2bcf34e6450fc990b905bed5c1f08d25d33b43c3

Observation 6510c82b-a3cd-4d81-946b-6fd8ced5044a · inbound

From Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation cites this paper.

From Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:01.778224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T08:45:46.944379Z digest=sha256:71912954d5abc847f7eadcadc94c96fabf89a018b6a7e110403ac21349cf3d44

Observation ca8be82c-36cf-427f-82a8-afccb44d388f · inbound

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving cites this paper.

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:28.605367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T03:48:36.717026Z digest=sha256:7ec4136ee9c920281a8417028c9773bd319929262658d5defecd138aeeece343

Observation cabd6262-800a-4b4f-b89c-cab10128566b · inbound

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving cites this paper.

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.560018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T21:36:52.396245Z digest=sha256:84c4c57f8b90a1af9e7f65f3a4db8b9bc5e0f65cb1923eca1a330bc0c12571a5

Observation 3086f0b7-76e7-4d65-9840-fb635026f0cd · inbound

Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving cites this paper.

Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:01:09.043469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:56:45.641080Z digest=sha256:27739b91f60bb21cc01dae6bafad3d65a5545fccadf0795bbb720f737dae48b3

Observation 25f7670b-e538-47fb-8b7c-c1bd06464768 · inbound

Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving cites this paper.

Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:55.995325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T16:40:24.858861Z digest=sha256:504a625d352fe4e692f2ec70ffdf854738a3959dd6cb792b5139469fd1af9551

Observation 0bfa4449-05bd-4707-a3ea-fe2515e5471e · inbound

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation cites this paper.

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:52.739321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T01:46:14.430539Z digest=sha256:c958bd68bac0b79c27382eb542e001712cbf1e7cb5600c22127690a060646fe0

Observation 89427131-e1a6-4272-901f-3792c54795eb · inbound

ReWorld: Learning Better Representations for World Action Models cites this paper.

ReWorld: Learning Better Representations for World Action Models Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:58.424941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T01:58:46.435886Z digest=sha256:14643ff465bce21008569519c8b8b2da21d67f0ee17b1b99a44bf66ac8a70cb3

Observation dceccfac-7cfb-4f7f-9062-8aa62f9a58a0 · inbound

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation cites this paper.

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T08:19:04.131379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:19:04.131379Z digest=sha256:a4942b657aea25a4ab0887da8adb9e39d549f766a37ffdac295086b02efaecae