Pith. sign in

Paper Citation Record · LEDGER

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation

As of 8 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2505.24139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24139 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:48.142792Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved37
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31c9008d-33aa-44b6-af9a-559fbd3d502c · outbound

This paper cites GPT-4 Technical Report.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.676305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.676305Z digest=sha256:cd1aa1e890cfe37a3c65cb2a6297e6831257acceef6f30efa622f467123882f9

Observation 2ed61a3a-8073-4315-877b-5a3a05bd4045 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.773765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.773765Z digest=sha256:cdfa3403bb83e9abf02548101d6de1b3fda7c6948b7b50464d3587452aaff21c

Observation beddcfe4-0057-4cd1-8d22-de14e5b055cc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:38.885701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:38.885701Z digest=sha256:56a0b2686d92a7d68dfa6ca77e60b340c4f1d51779941dd9b5da49dd63013c0e

Observation 66aac8d1-ebcc-4f26-8130-3fbccb7b9677 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.010726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.010726Z digest=sha256:6ad35722d6b0e22b5a4177b43c156d29b345432d42d7fbdf7fee2b97784772fc

Observation bed03f34-f29f-4830-a37a-767fcaf7edfc · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation nuscenes: A multi- modal dataset for autonomous driving

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.115143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.115143Z digest=sha256:6b7620c847004059fb59c9e3828ef0cd1e57c544b310fda281f348ed3a325d2e

Observation cb07947d-7a04-40f8-b8fd-1dce74a61729 · outbound

This paper cites Mp3: A unified model to map, perceive, predict and plan.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Mp3: A unified model to map, perceive, predict and plan

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.952192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:39.247829Z digest=sha256:96e65254f8e3131119afb672c73dd14f53a2f30c42ae5360975d8e48d1914ff3

Observation 74c6998d-0bb9-484f-9753-cf4ef7dc9242 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.414992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.414992Z digest=sha256:b63b6d604c3af128b96bf07a8cff2a641fe75bd7acd468e7634716d4a892d96a

Observation 61e4c13f-0271-4849-8a80-d10972dd5686 · outbound

This paper cites Driving with llms: Fusing object-level vec- tor modality for explainable autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Driving with llms: Fusing object-level vec- tor modality for explainable autonomous driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.689795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:39.514849Z digest=sha256:652e2bd800f5d6da6b7bb0c60340d571f008d58564b9e15d946dbdcacc0da475

Observation cc1d5806-5d52-4b51-87e6-86b48484e907 · outbound

This paper cites Pali: A jointly- scaled multilingual language-image model.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Pali: A jointly- scaled multilingual language-image model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.421405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:39.632649Z digest=sha256:a9403df183d7a0450289c2b1338ad397ce49f50ef19c47c0e18169368bcb2e91

Observation 7ef6ce2e-1f01-4d44-aac5-6835a73a4daa · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.786935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.786935Z digest=sha256:a3ba39574f86da2bfdd8349658ea58b03d1cb42824b37e3636340dccff72b64c

Observation 6c1db7ac-7239-46d4-a0be-41b8f3b2a0d3 · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.933165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.933165Z digest=sha256:79e2ddb1b551da72796a969b64b900a60540b8162b88a07d0ec8702870ae34a7

Observation 1c0425df-174d-4689-8fb1-2c0ecc62c389 · outbound

This paper cites Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.065086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.065086Z digest=sha256:2b9db390b391453309b1c6596ac2f165ee141ed8cf0eebb3037bebbdd0b59de3

Observation fe1abe47-c89e-4e09-b562-e7577a5048b2 · outbound

This paper cites End-to-end driving via conditional imitation learning.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation End-to-end driving via conditional imitation learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.196150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.196150Z digest=sha256:294b74319a59640c09e3ddb858277716fab51b4c0ca04c12d240007093135977

Observation 9f760794-0c1f-42ac-8532-55d501c2ad73 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.368657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.368657Z digest=sha256:cd5579da47800ccd0b9ffd551cd9c2e1b0c55c3d796b21ac3bc6ffd3ebfa72d1

Observation 96922c51-2048-4dd1-9fd2-44256feccb66 · outbound

This paper cites Holistic autonomous driving un- derstanding by bird’s-eye-view injected multi-modal large models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Holistic autonomous driving un- derstanding by bird’s-eye-view injected multi-modal large models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:57.166691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:40.524852Z digest=sha256:b86deaa652c506f69f071608632982375c8b6b12b975478aab41e9a55ea75e70

Observation cb7e1ec2-84e9-4a6b-8a38-e17b054a1e1c · outbound

This paper cites Carla: An open urban driv- ing simulator.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Carla: An open urban driv- ing simulator

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.663443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.663443Z digest=sha256:090914754cf895cf1963206176e840da510722c00b2cd828e47a722a3076e6a9

Observation 0013014d-7d77-4c1e-baee-dae7af983233 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.779008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.779008Z digest=sha256:9dad5194314f8664b77f23dda1cedc9e2afb16de7fc90a536eb2dc291229d3ba

Observation f8049e6c-3ac0-4c77-801f-7df1528b2882 · outbound

This paper cites Large scale interactive motion forecasting for autonomous driving: The waymo open mo- tion dataset.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Large scale interactive motion forecasting for autonomous driving: The waymo open mo- tion dataset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.921078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:40.900656Z digest=sha256:38f89bf812ff59b9fd4dd7a6c23238f699ad111ddf7484b4e127486fe557613a

Observation f954c2c1-d40a-4d9f-953e-175b80a289ba · outbound

This paper cites Eva: Exploring the limits of masked visual representa- tion learning at scale.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Eva: Exploring the limits of masked visual representa- tion learning at scale

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.646465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:41.060670Z digest=sha256:0c284e06b28d3def9971b4ba433a7528e82bc53b8e338f18728d99ca88e237b7

Observation 55954c71-019d-45fb-9b36-c3e7d0a0adc3 · outbound

This paper cites Simple-bev: What really mat- ters for multi-sensor bev perception? In2023 IEEE Inter- national Conference on Robotics and Automation (ICRA), pages 2759–2765.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Simple-bev: What really mat- ters for multi-sensor bev perception? In2023 IEEE Inter- national Conference on Robotics and Automation (ICRA), pages 2759–2765

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.341971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:41.217284Z digest=sha256:2a73c8ef371566ca21bf9e404456c4155ecf93df5ea5fedf11953905238dae05

Observation ae08d836-f8f3-4380-a4f6-34f825239a14 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation The Curious Case of Neural Text Degeneration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.350155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.350155Z digest=sha256:816d345ef7a61d1390bd2099d7ea8b37ce57e05c7b56bc5e73ec687f65f8fc70

Observation 15c7e1e6-df82-4de6-8f25-392ef1cecf0e · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.475718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.475718Z digest=sha256:b30cbab2e66a0351d48eeb138df16eeb94920e09aa77d084681dcdd645fc0fb6

Observation b59e1dfc-2918-4289-b8cc-7f1207360a04 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation LoRA: Low-Rank Adaptation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.597287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.597287Z digest=sha256:5d6f15740f0ad5ae026baf78e22b5cd18cfbf564e2dc4ffa5a64d295d5a703e9

Observation f18ee96d-3712-4bea-909d-5cd3c557066a · outbound

This paper cites St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation St-p3: End-to-end vision-based au- tonomous driving via spatial-temporal feature learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:41.723680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:41.723680Z digest=sha256:4ad5f4c42ab65688c46ccec741c69c33b31cbfb74b98bc7907edfb3dac21b66e

Observation 35d21fce-3f9f-4d2d-a6d9-a1cfab86d2eb · outbound

This paper cites Planning-oriented autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Planning-oriented autonomous driving

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:56.030218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:41.903017Z digest=sha256:2e656cdae2298f035eb737e4dd92086fb80aeee2e7df25b71f3a992e51931e91

Observation aa993d85-ab48-4e48-849a-7611211b010f · outbound

This paper cites Emma: End-to-end multimodal model for autonomous driving, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Emma: End-to-end multimodal model for autonomous driving, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:55.747632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:42.052634Z digest=sha256:d652dc2fb5b02844c843b956b6462504d84df3d59a07621b3ea4e3fedf0db503

Observation 6754c6eb-422a-492d-8fc9-68165356f283 · outbound

This paper cites Sym- phony: Learning realistic and diverse agents for autonomous driving simulation.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sym- phony: Learning realistic and diverse agents for autonomous driving simulation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:55.497172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:42.191165Z digest=sha256:0c0a361a408e9267f68a8e0b35277f5f9e529981a07abd1fb9bd0aee3cb1db23

Observation 664aad22-6a91-4e60-bfb3-8104a111ae04 · outbound

This paper cites Vad: Vectorized scene representa- tion for efficient autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Vad: Vectorized scene representa- tion for efficient autonomous driving

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:55.250573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:42.281582Z digest=sha256:686affa0acf10d4cba58eddd5f698831042cbbbb8ca587331e80a67ada942d76

Observation 2be857d5-4059-46d5-9d42-28e4b1e3daf9 · outbound

This paper cites Textual explanations for self-driving ve- hicles.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Textual explanations for self-driving ve- hicles

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:54.976491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:42.410239Z digest=sha256:f0ca69c5349e84d0fe766b6098aedc0cee1e6428f0ea8bcc54d393d0a4970ace

Observation 3c205576-dc4f-4a24-9360-742215e8f925 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:42.471180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:42.471180Z digest=sha256:e91c369a54b29a934b8a9e643063fe814af89ed0118d816387df49f33b865da4

Observation 5bf21733-b6ee-4834-8b13-17e90502d7aa · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:54.676964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:42.638491Z digest=sha256:9c9c787d3237121b0b9844572396288c667c999c0b66fa1a6eb92e6d19bf2955

Observation e9c72479-8522-4d78-bbcc-03ffe838c2e5 · outbound

This paper cites an unresolved cited work.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:38:54.395892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:42.760321Z digest=sha256:71dac946fdd7cbb316e42f5a2db08740d3e77999b87de818724426c4ed019899

Observation 4bda4cc1-3aed-4cc9-9c3b-47b8eb5bca8c · outbound

This paper cites Maptr: Structured modeling and learning for online vectorized hd map construction.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Maptr: Structured modeling and learning for online vectorized hd map construction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:54.075130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:42.821141Z digest=sha256:bfd0889a46c3ac377da2c23fa100b3558ea8641a41e920a5cded889695451c9a

Observation 228d6bf1-2abc-471a-88f1-f6bdc6a814bd · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:42.932843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:42.932843Z digest=sha256:e76def263419458218180a347476fa96b40b4cad454f5e34ab500556da32775b

Observation d01682a3-6782-4c96-85f1-e44173bc6431 · outbound

This paper cites Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Bevfusion: Multi- task multi-sensor fusion with unified bird’s-eye view repre- sentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:53.771640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:43.036990Z digest=sha256:ec7f36893199995361b523152e8db04c379e96b64a90c6375659c06a077c65b8

Observation c2ddc1c8-05c5-4301-99e0-ce15d4d69c00 · outbound

This paper cites When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models.arXiv preprint arXiv:2405.10255, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models.arXiv preprint arXiv:2405.10255, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.162693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.162693Z digest=sha256:15adec85b6e780be6b5279b1e7c60892dace9e2eedf3131fc5bae3ebafc3b095

Observation a5409517-a9f0-48ae-a5f2-665bc24bc018 · outbound

This paper cites Dolphins: Multimodal Language Model for Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Dolphins: Multimodal Language Model for Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.273602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.273602Z digest=sha256:b78766e53c6090c513b32a863a0f2070320b3c20ed83db9ec650aa8d0e176498

Observation 2bfa5ad5-60c5-40f4-9171-974bef9222c6 · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GPT-Driver: Learning to Drive with GPT

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.401500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.401500Z digest=sha256:0dab115b9207e4a7028ff92ba361cd3432fe8182a4920d9d7e542dccda9ec41f

Observation bd6c8412-f4d4-4fd7-afe4-ec11081536a1 · outbound

This paper cites A Language Agent for Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation A Language Agent for Autonomous Driving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.517282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.517282Z digest=sha256:69cc35084c7b25fc4f8ec6fdc6136572b9d1ef3b51271cfd24f866f842cfdf63

Observation 9e39d9be-e59f-403f-976d-c0213474736c · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.703061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.703061Z digest=sha256:b7471898ffacb1ee71d4f5f253fa0e5339ec1e747cbf3fcdfd63a66659f03cd9

Observation f61e82be-4f81-4ba0-8913-d1fad4f12ba0 · outbound

This paper cites Wayformer: Motion forecasting via simple & efficient attention networks.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Wayformer: Motion forecasting via simple & efficient attention networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:53.486190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:43.832722Z digest=sha256:0bd1cc24357b01f93c8c15a1c1c01bd21dd9bd521f30cd11cf509ddca66adc68

Observation e67a1496-3d2d-4c88-ae80-1e58be36b576 · outbound

This paper cites Gpt-4v(ision) system card, 2023.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Gpt-4v(ision) system card, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.955298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.955298Z digest=sha256:800425fb4076ef1f157392fb91ecea3983457f6a45fe998dee5a9bcd05d5d319

Observation f28f4a2c-ab29-4a21-8347-cb967f1b32f6 · outbound

This paper cites Alvinn: An autonomous land vehicle in a neural network.Advances in neural information processing systems, 1, 1988.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Alvinn: An autonomous land vehicle in a neural network.Advances in neural information processing systems, 1, 1988

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:53.211246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:44.059307Z digest=sha256:e42ccb6ee21cc2125196af1503c6534634f248dd68182431e8d4efbd9cd1c20c

Observation 164843cb-a90f-4556-ac38-254ef3a26b16 · outbound

This paper cites Qi, Yin Zhou, Mahyar Najibi, Pei Sun, Khoa T.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Qi, Yin Zhou, Mahyar Najibi, Pei Sun, Khoa T

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.982600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:44.176812Z digest=sha256:b6b92062db3307b25f4c5eb54a91e9bec3c9d992ed3fd6a2a17e54f4b0bc5e41

Observation 2e711174-492a-42da-b0b0-db22b9d36e76 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Learning transferable visual models from natural language supervi- sion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:44.272624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:44.272624Z digest=sha256:3e9ac8cb467a9b19b3933128689a2a980998e154dad8eea1f66bf98b215ee4ca

Observation b3e9b324-3513-4cc7-93ea-a422d1914fe1 · outbound

This paper cites Motionlm: Multi-agent motion forecast- ing as language modeling.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Motionlm: Multi-agent motion forecast- ing as language modeling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.748592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:44.450630Z digest=sha256:5ecf03608d40d6d9cdf34f98da57c45eb6ceed4e0808aec603608e5fdaf80dfa

Observation d89266e2-3fe4-4b3b-8013-48662518bcc4 · outbound

This paper cites LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:44.598445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:44.598445Z digest=sha256:e2a6f854944d13b0f704a726296cc0aaa9dd6869f35221647c8dd0b4f00c1dae

Observation 2c5eda30-d8a7-4be8-8e5c-70e790622df1 · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Lmdrive: Closed-loop end-to-end driving with large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.504616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:44.772508Z digest=sha256:b5e3244363589f87f3fa1307d0b1384eb1a889146446c1909d1a5e89d3cd7bab

Observation c18b1de5-5d85-418d-9a19-45888623b327 · outbound

This paper cites Drivelm: Driving with graph visual question answering.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Drivelm: Driving with graph visual question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.276600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:44.920320Z digest=sha256:fc7ad98915bb7ad1346a54251d41d4f4bb09e589a30b9fe5e18f34ae3d8f67bd

Observation 9f5b6328-35fb-47f7-b67c-3ccee3ab4dc0 · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation UL2: Unifying Language Learning Paradigms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.070721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.070721Z digest=sha256:59d2fecdf5ce6bd88d2c5b2f8e316d5bbd71635904037e29597cf8f293a71c0d

Observation 879eb1ed-e2e9-4b61-84ef-e2af356c1312 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.214604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.214604Z digest=sha256:b3200ba2ab0f8d922f5876eabb2af1936065e93b71dbbef7d9c8934b879ea236

Observation 7888535f-c224-4d0a-9828-58511b917915 · outbound

This paper cites Tokenize the world into object-level knowledge to address long-tail events in autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Tokenize the world into object-level knowledge to address long-tail events in autonomous driving

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:52.003873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:45.377764Z digest=sha256:b1887e73187a71faa93f4de4918bb19386219805a9b3ad0fa386d1d58a0c7e96

Observation 9b28db14-2292-474c-ae4f-7fa7680c4b4b · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.549532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.549532Z digest=sha256:c99673a97f8c8b1b27089088bb26e91d8c34ee3829e1df8b8cc0154d40b986ef

Observation 0a03b6eb-83bf-4bae-9457-f30b89c83a66 · outbound

This paper cites Multipath++: Efficient information fu- sion and trajectory aggregation for behavior prediction.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Multipath++: Efficient information fu- sion and trajectory aggregation for behavior prediction

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.699027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.699027Z digest=sha256:fff3531551c7bfca717a762667c033c159cee651bf4e8cbf4f007f0e4c1f059c

Observation f776f205-7415-4d28-abd2-ed2c8eda2da1 · outbound

This paper cites OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.850206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.850206Z digest=sha256:99895d8bbda1f9bde50820c38a060bf7e6ec352755d5d7f2ad18363aa0c3e19a

Observation 26a79353-1439-4b73-8101-767c0ba465eb · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:45.988947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:45.988947Z digest=sha256:650a8c30d985bc4c6a77a200522109aff0fd5b0c4cca9486b7fe8490afb39396

Observation 8d694b7e-af38-40e2-9fb7-028f7c06024d · outbound

This paper cites Para-drive: Parallelized architecture for real- time autonomous driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Para-drive: Parallelized architecture for real- time autonomous driving

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:51.716475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:46.128576Z digest=sha256:cbb3980b791dc19d1a96ef92a3ac9fec8a0fd4d289ef057ff07af7d3d775decf

Observation 2ca55e1d-e88e-442e-a77a-233eda41c84c · outbound

This paper cites Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong base- line.Advances in Neural Information Processing Systems, 35:6119–6132, 2022.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong base- line.Advances in Neural Information Processing Systems, 35:6119–6132, 2022

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:51.402951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:46.345104Z digest=sha256:ed7fb1876e5c4d956d9a256a1d8c38c8f3c7c11e4f8ae792233c80bc836d420e

Observation ba4c6815-35bc-404c-9a2a-03768f953b19 · outbound

This paper cites Grok-1.5 vision preview, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Grok-1.5 vision preview, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:51.146156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:46.496905Z digest=sha256:51905f22fe35abd28db6ad9b11fb92f3c991689a91fa1bcc8c6eb85b699adbae

Observation 82a398da-2cd3-41a4-8527-1dae377c99d3 · outbound

This paper cites Sparsefusion: Fusing multi-modal sparse rep- resentations for multi-sensor 3d object detection.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sparsefusion: Fusing multi-modal sparse rep- resentations for multi-sensor 3d object detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.858663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:46.704458Z digest=sha256:c5fee29d03e96d9ebb40e11afd60c106e067e5ddb5ed8a2ca3abd7c06eae9ba7

Observation 590f44fb-28c3-4cdc-86c2-d8339a9b5484 · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model.IEEE Robotics and Automation Let- ters, 2024.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Drivegpt4: Interpretable end-to-end autonomous driving via large language model.IEEE Robotics and Automation Let- ters, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.578166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:46.855253Z digest=sha256:0fcb09b3e0f2e2b617f0e8b471c98f01b36e4fc68828f4f2bdc93ed07461fe9c

Observation d2c482d2-a73b-4bfe-9cf6-f7ee97d7fef2 · outbound

This paper cites Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:47.007757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:47.007757Z digest=sha256:e2deb24d47c11cd0b77713b200bb8a7907da671a62efe8b27763ab6bff61ad99

Observation c4cfb3a3-b7f1-4f2b-a5d0-80d6e43ded02 · outbound

This paper cites Sigmoid loss for language image pre-training.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Sigmoid loss for language image pre-training

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.381230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:47.131659Z digest=sha256:fc978047964f9a01a4981ed904f64f03bb8b57a9acfc4227e949a2752c0c5973

Observation a57a1a17-8e8e-485d-8ff8-4aafe132ef36 · outbound

This paper cites P Xing, Hao Zhang, Joseph E.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation P Xing, Hao Zhang, Joseph E

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:50.071638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:47.261925Z digest=sha256:0bf2cd7c813b57c4945318c287817fc3b1458c7215af1cc9da82c91933e5161d

Observation 201df00b-e06b-4070-8860-a2f6dee2fa7f · outbound

This paper cites OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:47.398500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:47.398500Z digest=sha256:d21c6c1760314805d6fff97ae04c38689b3c4dd61a0249b30185ff0e95903450

Observation eab5a7d9-eafd-418e-bcb6-80cc2d23eb9e · outbound

This paper cites GenAD: Generative End-to-End Autonomous Driving.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation GenAD: Generative End-to-End Autonomous Driving

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:47.557681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:47.557681Z digest=sha256:c9d9f4315fd5ea900f1e0eed1a6c2e2c380de0ef0bda0a7b5108b644efc163db

Observation 73dfb289-b232-4559-bb1a-f4a08fca04c0 · outbound

This paper cites go straight forward.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation go straight forward

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:38:49.760921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:47.711590Z digest=sha256:0fe2e5a883be0888ce1fe29799870e7f0680c0db5308ba9c4d9c2fb72e6bb6bc

Observation 496eae92-decf-4a09-8c35-75b441213256 · outbound

This paper cites 7, we report theADE@5smetric of S4-Driver for each ego-vehicle behavior on WOMD-Planning-ADE benchmark separately.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 7, we report theADE@5smetric of S4-Driver for each ego-vehicle behavior on WOMD-Planning-ADE benchmark separately

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:49.444570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:47.858650Z digest=sha256:350b02d04cd792feff465ce18be95636e57b700c5ae1c38c0cd2db663fe04d7b

Observation 8c2c5b16-bcd8-4a11-a3dd-437c8fde576a · outbound

This paper cites 10, we visualize more planning results on WOMD-Planning-ADE.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation 10, we visualize more planning results on WOMD-Planning-ADE

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:49.104731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:48.010515Z digest=sha256:857ab51f0ef00a774d5bbdfd76b36f4379d876d1841a6829c28973efea32240d

Observation 2256c5ff-8d20-4b75-b615-7ac5f6e3322e · outbound

This paper cites Camera configuration.We apply different configurations of camera sensors in Tab.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Camera configuration.We apply different configurations of camera sensors in Tab

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:38:48.812486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:38:48.142792Z digest=sha256:ade8e921e35ca05b3634ba26bd7aae1efcad0d108842f03d50ef8a486b529ca1

Pith citing papers

No inbound Pith citation observations are available.