Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:48.142792Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2505.24139.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:48.142792Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 31c9008d-33aa-44b6-af9a-559fbd3d502c · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed61a3a-8073-4315-877b-5a3a05bd4045 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beddcfe4-0057-4cd1-8d22-de14e5b055cc · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66aac8d1-ebcc-4f26-8130-3fbccb7b9677 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation PaliGemma: A versatile 3B VLM for transfer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed03f34-f29f-4830-a37a-767fcaf7edfc · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation nuscenes: A multi- modal dataset for autonomous driving
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb07947d-7a04-40f8-b8fd-1dce74a61729 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Mp3: A unified model to map, perceive, predict and plan
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74c6998d-0bb9-484f-9753-cf4ef7dc9242 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61e4c13f-0271-4849-8a80-d10972dd5686 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Driving with llms: Fusing object-level vec- tor modality for explainable autonomous driving
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc1d5806-5d52-4b51-87e6-86b48484e907 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Pali: A jointly- scaled multilingual language-image model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ef6ce2e-1f01-4d44-aac5-6835a73a4daa · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c1db7ac-7239-46d4-a0be-41b8f3b2a0d3 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c0425df-174d-4689-8fb1-2c0ecc62c389 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe1abe47-c89e-4e09-b562-e7577a5048b2 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation End-to-end driving via conditional imitation learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f760794-0c1f-42ac-8532-55d501c2ad73 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Instructblip: Towards general- purpose vision-language models with instruction tuning,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96922c51-2048-4dd1-9fd2-44256feccb66 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Holistic autonomous driving un- derstanding by bird’s-eye-view injected multi-modal large models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb7e1ec2-84e9-4a6b-8a38-e17b054a1e1c · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Carla: An open urban driv- ing simulator
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0013014d-7d77-4c1e-baee-dae7af983233 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8049e6c-3ac0-4c77-801f-7df1528b2882 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Large scale interactive motion forecasting for autonomous driving: The waymo open mo- tion dataset
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f954c2c1-d40a-4d9f-953e-175b80a289ba · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Eva: Exploring the limits of masked visual representa- tion learning at scale
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55954c71-019d-45fb-9b36-c3e7d0a0adc3 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Simple-bev: What really mat- ters for multi-sensor bev perception? In2023 IEEE Inter- national Conference on Robotics and Automation (ICRA), pages 2759–2765
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae08d836-f8f3-4380-a4f6-34f825239a14 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation The Curious Case of Neural Text Degeneration
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15c7e1e6-df82-4de6-8f25-392ef1cecf0e · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b59e1dfc-2918-4289-b8cc-7f1207360a04 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation LoRA: Low-Rank Adaptation of Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f18ee96d-3712-4bea-909d-5cd3c557066a · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d21fce-3f9f-4d2d-a6d9-a1cfab86d2eb · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Planning-oriented autonomous driving
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa993d85-ab48-4e48-849a-7611211b010f · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Emma: End-to-end multimodal model for autonomous driving, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6754c6eb-422a-492d-8fc9-68165356f283 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sym- phony: Learning realistic and diverse agents for autonomous driving simulation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 664aad22-6a91-4e60-bfb3-8104a111ae04 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Vad: Vectorized scene representa- tion for efficient autonomous driving
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2be857d5-4059-46d5-9d42-28e4b1e3daf9 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Textual explanations for self-driving ve- hicles
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c205576-dc4f-4a24-9360-742215e8f925 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bf21733-b6ee-4834-8b13-17e90502d7aa · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9c72479-8522-4d78-bbcc-03ffe838c2e5 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bda4cc1-3aed-4cc9-9c3b-47b8eb5bca8c · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Maptr: Structured modeling and learning for online vectorized hd map construction
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 228d6bf1-2abc-471a-88f1-f6bdc6a814bd · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d01682a3-6782-4c96-85f1-e44173bc6431 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2ddc1c8-05c5-4301-99e0-ce15d4d69c00 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models.arXiv preprint arXiv:2405.10255, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5409517-a9f0-48ae-a5f2-665bc24bc018 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Dolphins: Multimodal Language Model for Driving
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bfa5ad5-60c5-40f4-9171-974bef9222c6 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GPT-Driver: Learning to Drive with GPT
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd6c8412-f4d4-4fd7-afe4-ec11081536a1 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation A Language Agent for Autonomous Driving
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e39d9be-e59f-403f-976d-c0213474736c · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f61e82be-4f81-4ba0-8913-d1fad4f12ba0 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Wayformer: Motion forecasting via simple & efficient attention networks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e67a1496-3d2d-4c88-ae80-1e58be36b576 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Gpt-4v(ision) system card, 2023
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f28f4a2c-ab29-4a21-8347-cb967f1b32f6 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Alvinn: An autonomous land vehicle in a neural network.Advances in neural information processing systems, 1, 1988
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 164843cb-a90f-4556-ac38-254ef3a26b16 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Qi, Yin Zhou, Mahyar Najibi, Pei Sun, Khoa T
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e711174-492a-42da-b0b0-db22b9d36e76 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Learning transferable visual models from natural language supervi- sion
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3e9b324-3513-4cc7-93ea-a422d1914fe1 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Motionlm: Multi-agent motion forecast- ing as language modeling
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d89266e2-3fe4-4b3b-8013-48662518bcc4 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c5eda30-d8a7-4be8-8e5c-70e790622df1 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Lmdrive: Closed-loop end-to-end driving with large language models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c18b1de5-5d85-418d-9a19-45888623b327 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Drivelm: Driving with graph visual question answering
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f5b6328-35fb-47f7-b67c-3ccee3ab4dc0 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation UL2: Unifying Language Learning Paradigms
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879eb1ed-e2e9-4b61-84ef-e2af356c1312 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Gemini: A Family of Highly Capable Multimodal Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7888535f-c224-4d0a-9828-58511b917915 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Tokenize the world into object-level knowledge to address long-tail events in autonomous driving
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b28db14-2292-474c-ae4f-7fa7680c4b4b · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a03b6eb-83bf-4bae-9457-f30b89c83a66 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Multipath++: Efficient information fu- sion and trajectory aggregation for behavior prediction
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f776f205-7415-4d28-abd2-ed2c8eda2da1 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a79353-1439-4b73-8101-767c0ba465eb · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d694b7e-af38-40e2-9fb7-028f7c06024d · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Para-drive: Parallelized architecture for real- time autonomous driving
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ca55e1d-e88e-442e-a77a-233eda41c84c · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong base- line.Advances in Neural Information Processing Systems, 35:6119–6132, 2022
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba4c6815-35bc-404c-9a2a-03768f953b19 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Grok-1.5 vision preview, 2024
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82a398da-2cd3-41a4-8527-1dae377c99d3 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sparsefusion: Fusing multi-modal sparse rep- resentations for multi-sensor 3d object detection
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 590f44fb-28c3-4cdc-86c2-d8339a9b5484 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Drivegpt4: Interpretable end-to-end autonomous driving via large language model.IEEE Robotics and Automation Let- ters, 2024
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2c482d2-a73b-4bfe-9cf6-f7ee97d7fef2 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4cfb3a3-b7f1-4f2b-a5d0-80d6e43ded02 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sigmoid loss for language image pre-training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a57a1a17-8e8e-485d-8ff8-4aafe132ef36 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation P Xing, Hao Zhang, Joseph E
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 201df00b-e06b-4070-8860-a2f6dee2fa7f · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eab5a7d9-eafd-418e-bcb6-80cc2d23eb9e · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GenAD: Generative End-to-End Autonomous Driving
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73dfb289-b232-4559-bb1a-f4a08fca04c0 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation go straight forward
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 496eae92-decf-4a09-8c35-75b441213256 · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 7, we report theADE@5smetric of S4-Driver for each ego-vehicle behavior on WOMD-Planning-ADE benchmark separately
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c2c5b16-bcd8-4a11-a3dd-437c8fde576a · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 10, we visualize more planning results on WOMD-Planning-ADE
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2256c5ff-8d20-4b75-b615-7ac5f6e3322e · outbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Camera configuration.We apply different configurations of camera sensors in Tab
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.