Pith. sign in

Paper Citation Record · LEDGER

Spatial-aware Vision Language Model for Autonomous Driving

As of 7 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 5 inbound Pith citation observations for arXiv:2512.24331.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.24331 v2

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T13:26:11.454206Z

measured 93 of 93 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-05T11:39:05.686584Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T11:41:02.600232Z

Reference resolution

88 of 88 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved88
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d8f00c4-279f-45fb-831d-5f469199e314 · outbound

This paper cites Qwen2.5-VL Technical Report.

Spatial-aware Vision Language Model for Autonomous Driving Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:04.102974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:04.102974Z digest=sha256:d7b8962033cf21082240727ec97140f9f576c8323a36aa90867f3c0997215b9c

Observation 9fdc826d-f5ab-48d7-99a6-3f6e4ae5eb0f · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gian- carlo Baldan, and Oscar Beijbom.

Spatial-aware Vision Language Model for Autonomous Driving Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gian- carlo Baldan, and Oscar Beijbom

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:04.229161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:04.229161Z digest=sha256:8d2837b879e3d536f8e1f35e6f6ff4012114ece13cc0971ac68945492d0d3521

Observation 795edb89-fa86-4894-a991-a98283a9fa88 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Spatial-aware Vision Language Model for Autonomous Driving Emerg- ing properties in self-supervised vision transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:04.348454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:04.348454Z digest=sha256:02ae753bef3698f9c2bf5e63eb6203fc2aa474470b50164135f96a4bcbc4e116

Observation 4b9e61bb-647e-4da8-8cd5-8ed246e6e960 · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

Spatial-aware Vision Language Model for Autonomous Driving SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:04.402166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:04.402166Z digest=sha256:4a582d5224bdf99bde50bc44e2d929522723c65437c8bfab472a13e57457a6fc

Observation b053fe16-3499-4b08-aa81-0a3a95788435 · outbound

This paper cites Persformer: 3d lane detection via perspective transformer and the openlane benchmark.

Spatial-aware Vision Language Model for Autonomous Driving Persformer: 3d lane detection via perspective transformer and the openlane benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:04.443067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:04.443067Z digest=sha256:370e3cd033367f5d79d5335e2f95567e0ced07c0fcdab5e11bfd7098d2946cbd

Observation 7a4fee93-bbb5-4103-9179-29832c60f2db · outbound

This paper cites An empirical study of training self-supervised vision transformers.

Spatial-aware Vision Language Model for Autonomous Driving An empirical study of training self-supervised vision transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:04.485438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:04.485438Z digest=sha256:45337cf5e84539b06c6d3f6098be0d8da4273eff5c3b724c4eeaf3faf7077924

Observation 45e76b3b-cec0-4b5c-ac05-7c646360a21d · outbound

This paper cites Spa- tialRGPT: Grounded Spatial Reasoning in Vision Language Models.

Spatial-aware Vision Language Model for Autonomous Driving Spa- tialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:04.624641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:04.624641Z digest=sha256:171a6fbb2da589203287d026d0aef8d5b6a5b8e28981d65ea60636011062bdc5

Observation 092c3ca5-c092-4414-8f2c-5a9acd3c04f9 · outbound

This paper cites Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving.

Spatial-aware Vision Language Model for Autonomous Driving Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:04.766321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:04.766321Z digest=sha256:2fbb74e3ddd105d76b8eb4e09e546edaf95679a23bb2f30ae56fec3f261b8b35

Observation 49457181-fbc1-4e90-8e5b-28f57d5c7ffd · outbound

This paper cites Talk2car: Taking control of your self-driving car.

Spatial-aware Vision Language Model for Autonomous Driving Talk2car: Taking control of your self-driving car

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:04.887908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:04.887908Z digest=sha256:70c8d6087c15f1894fc9ab661e3ca0edeecdbf3a1a5cc2e60898400715112319

Observation 8e454b60-2b99-451b-b44d-1e55aee83e79 · outbound

This paper cites Holistic autonomous driving un- derstanding by bird’s-eye-view injected multi-modal large models.

Spatial-aware Vision Language Model for Autonomous Driving Holistic autonomous driving un- derstanding by bird’s-eye-view injected multi-modal large models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.002937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.002937Z digest=sha256:817bf250cfdf86e4df3c45297d968db417378837ca406537d8864ad9725f8c33

Observation 7c2f64fe-55ad-4dfb-b443-d9721013371d · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Spatial-aware Vision Language Model for Autonomous Driving An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.138241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.138241Z digest=sha256:e2345c345484894c3f2e7a2bc6b591283ecef2db269272bcd1aa16a74b3f54e6

Observation 1ec40d66-0471-44ff-8617-a5e16bdf73b8 · outbound

This paper cites Fsd v2: Improving fully sparse 3d object detection with vir- tual voxels.IEEE Transactions on Pattern Analysis and Ma- chine Intelligence, 2024.

Spatial-aware Vision Language Model for Autonomous Driving Fsd v2: Improving fully sparse 3d object detection with vir- tual voxels.IEEE Transactions on Pattern Analysis and Ma- chine Intelligence, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.220639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.220639Z digest=sha256:1441e37fb3e48cfa35dc03f82fafce929871b72c14b14ddec04072a33f1dad49

Observation 4a0c3f46-6c8a-461d-9195-b5ba888a384d · outbound

This paper cites Eva-02: A visual representation for neon genesis.Image and Vision Computing, page 105171,.

Spatial-aware Vision Language Model for Autonomous Driving Eva-02: A visual representation for neon genesis.Image and Vision Computing, page 105171,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.259987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.259987Z digest=sha256:d1c6cb9402f384932754795c00cb0fcef0b74d2b2f74820bfa438c3c1c61a040

Observation 7e8d08e6-fd34-4cff-b520-5ebe9b252c84 · outbound

This paper cites ORION: A Holistic End-to- End Autonomous Driving Framework by Vision-Language Instructed Action Generation.

Spatial-aware Vision Language Model for Autonomous Driving ORION: A Holistic End-to- End Autonomous Driving Framework by Vision-Language Instructed Action Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.349383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.349383Z digest=sha256:26855c5eec923c991fe8b09928ca25af078fe77b4db50543a7e23d7d383836dd

Observation 52925170-7639-4839-acbb-1d697b65e85c · outbound

This paper cites Multi-Frame, Lightweight & Efficient Vision-Language Mod- els for Question Answering in Autonomous Driving.

Spatial-aware Vision Language Model for Autonomous Driving Multi-Frame, Lightweight & Efficient Vision-Language Mod- els for Question Answering in Autonomous Driving

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.415969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.415969Z digest=sha256:84aed7e5531d7b7f3bf8d5ac32b47a72aec9c3ee7d5f3b7b31ca172417b31107

Observation 2587d52f-f32f-4747-bd68-d9d167d755cb · outbound

This paper cites Mask r-cnn.

Spatial-aware Vision Language Model for Autonomous Driving Mask r-cnn

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.480439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.480439Z digest=sha256:9b31a3735d76d43c695998ed841aade1042a33b62048c74cdd833d2934692a73

Observation d166557d-1bf2-4559-a3af-ec191b8a4ccb · outbound

This paper cites Masked autoencoders are scalable vision learners.

Spatial-aware Vision Language Model for Autonomous Driving Masked autoencoders are scalable vision learners

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.546368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.546368Z digest=sha256:f1478798a760f900384e7b98dceaba502cc3bd8b8d162b487af6395f2cbd9f0d

Observation 77f7ba63-0f3a-47bf-878b-2e737e5a923a · outbound

This paper cites Planning-oriented autonomous driving.

Spatial-aware Vision Language Model for Autonomous Driving Planning-oriented autonomous driving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.625367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.625367Z digest=sha256:6799fcf00a8232fd7eed67d3c110acd2178902ec93ae38d0f0a96040a43beec5

Observation f0760355-46f3-42a5-adab-2b3854494847 · outbound

This paper cites Drivlme: Enhancing llm-based autonomous driving agents with embodied and social experiences.

Spatial-aware Vision Language Model for Autonomous Driving Drivlme: Enhancing llm-based autonomous driving agents with embodied and social experiences

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.714982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.714982Z digest=sha256:177b7d5fdbdcbe5fff6e7c7c7654a9094fe612c4c5254c37a674ad48666d8c63

Observation 842926fd-15f1-4c48-a826-105b18c6454a · outbound

This paper cites GPT-4o System Card.

Spatial-aware Vision Language Model for Autonomous Driving GPT-4o System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.774332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.774332Z digest=sha256:bcd5110330a95bf59a2f9a02e2e884beb925510312d21627efcd292c804bca2e

Observation 0b735342-7dce-41c9-af12-379676b54f8d · outbound

This paper cites Omnispa- tial: Towards comprehensive spatial reasoning benchmark for vision language models.arXiv preprint arXiv:2506.03135,.

Spatial-aware Vision Language Model for Autonomous Driving Omnispa- tial: Towards comprehensive spatial reasoning benchmark for vision language models.arXiv preprint arXiv:2506.03135,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.879351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.879351Z digest=sha256:e4396420bb6e4c4a6b8a403503e8a2990a88ba6afd74590c4a3ce2fce60704dd

Observation 77d74363-c8ce-4378-be6e-50b4a59919f3 · outbound

This paper cites Vad: Vectorized scene representation for ef- ficient autonomous driving.

Spatial-aware Vision Language Model for Autonomous Driving Vad: Vectorized scene representation for ef- ficient autonomous driving

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:05.957461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:05.957461Z digest=sha256:6c4b6cc73ec6ab61bf1af57e1ab1c6aea5dee908bac86284897e78b6966b0530

Observation 1da77f5e-8146-475e-8dff-332fd3128df5 · outbound

This paper cites Textual explanations for self-driving vehicles.

Spatial-aware Vision Language Model for Autonomous Driving Textual explanations for self-driving vehicles

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.050138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.050138Z digest=sha256:05165d94cdc1261436cb1fd091767da267ad1695ee7c9da25181082cde43c347

Observation ef09e4f7-468b-4fb4-9af3-4bb86ae688aa · outbound

This paper cites Driving everywhere with large language model policy adaptation.

Spatial-aware Vision Language Model for Autonomous Driving Driving everywhere with large language model policy adaptation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.119663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.119663Z digest=sha256:41e6cc22bad49b098c678bc5f9ecbcce4565d25d66cf8c1848618735afd6cd50

Observation a14c2991-c5f2-4fe4-ab1b-a0711dcbc05a · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

Spatial-aware Vision Language Model for Autonomous Driving Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.174798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.174798Z digest=sha256:df95835a339f9d8b2dd06e8ea0240dc7ed32381ebf9dad4358c18a51d72f0efc

Observation 198adf4a-2e26-4a85-86d6-99fbbd2b008b · outbound

This paper cites Blip- 2: bootstrapping language-image pre-training with frozen image encoders and large language models.

Spatial-aware Vision Language Model for Autonomous Driving Blip- 2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.278735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.278735Z digest=sha256:abff664ffdc108f969a2f6215ed7900ab3b8577afc212b9ad80d534896ad018b

Observation 86056dfb-039f-4938-8f1b-f4c6aa88f119 · outbound

This paper cites an unresolved cited work.

Spatial-aware Vision Language Model for Autonomous Driving Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.335962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.335962Z digest=sha256:f004a2db3e4e5b01ac6e489eb84330a5dd6d04264c2295c304e1b62e69a99ea8

Observation 215e248c-9a15-4237-876a-4fed81754d8a · outbound

This paper cites Bevfusion: A simple and robust lidar-camera fusion framework.

Spatial-aware Vision Language Model for Autonomous Driving Bevfusion: A simple and robust lidar-camera fusion framework

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.408093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.408093Z digest=sha256:7018fb60e1bedbffb1b51abbe0b3c977c492c823a2db3ce566b2339fdc78ae47

Observation eee31b8f-1193-47a9-993a-5a1a7218503e · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Spatial-aware Vision Language Model for Autonomous Driving Rouge: A package for automatic evaluation of summaries

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.505518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.505518Z digest=sha256:e4eb57f21ea7133afc668ad66f7ca2749a12e3d8e67fbc4909be9b966d487cc8

Observation 303d5d29-3b75-481f-aff7-7ed5fe817837 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Spatial-aware Vision Language Model for Autonomous Driving Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.547793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.547793Z digest=sha256:5018318f111bccbd3ee1d128594909671c9f44680614f8c00543c5dc80c01722

Observation 2554cd72-db52-4710-af52-5e621ed9888e · outbound

This paper cites Improved baselines with visual instruction tuning.

Spatial-aware Vision Language Model for Autonomous Driving Improved baselines with visual instruction tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.646904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.646904Z digest=sha256:16cf717c75d71d1f8868792bf407abddb251225eb04ce9951b3471d3fd7c5db4

Observation 09564187-549e-4c2d-84e4-0eef8ac830b1 · outbound

This paper cites Can Multimodal Large Language Models Understand Spatial Relations?.

Spatial-aware Vision Language Model for Autonomous Driving Can Multimodal Large Language Models Understand Spatial Relations?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.713886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.713886Z digest=sha256:aba29e8e04038e6dbc002e7794b260c7d9b5e8cda83dd972326d4f807ca331c5

Observation cc31a511-0726-4128-8295-55a78bb5a3ad · outbound

This paper cites CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving.

Spatial-aware Vision Language Model for Autonomous Driving CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.814034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.814034Z digest=sha256:bc16a5df042d6087de6bef7224475d90cd0aa7d635d439ff6e2d5b418b5adabe

Observation c3e9bf31-27d0-47de-acea-87c5b2e3e085 · outbound

This paper cites 3dsrbench: A comprehensive 3d spatial reasoning benchmark.

Spatial-aware Vision Language Model for Autonomous Driving 3dsrbench: A comprehensive 3d spatial reasoning benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.900202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.900202Z digest=sha256:9886830bf59c1751903c776ef2390944748aca4c9be366dfddaf733858947da3

Observation 7a72f88e-63ac-4761-8a60-498734d01265 · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

Spatial-aware Vision Language Model for Autonomous Driving GPT-Driver: Learning to Drive with GPT

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:06.977588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:06.977588Z digest=sha256:98fa0f3a3d007df88b7ef1a996952204303a388adafa95ba0493db0b80cd5911

Observation 032b3462-e5e3-469d-b9df-5f37cbdae398 · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving.

Spatial-aware Vision Language Model for Autonomous Driving Lingoqa: Visual question answering for autonomous driving

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:07.060840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:07.060840Z digest=sha256:697c882bc88199ada0c97f1edbaad1d026807845faec2732433794b456bfeca2

Observation 66aea150-ff76-4a06-bb33-93393a7b36f3 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744,.

Spatial-aware Vision Language Model for Autonomous Driving Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:07.137220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:07.137220Z digest=sha256:d852aeeb00a873bc3df6c8d806369f16a8f1ac6933680ce12ee280c959ed2878

Observation dc236a92-bc36-4f61-b29d-cf85259ab088 · outbound

This paper cites Vlp: Vision language planning for autonomous driving.

Spatial-aware Vision Language Model for Autonomous Driving Vlp: Vision language planning for autonomous driving

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:07.224029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:07.224029Z digest=sha256:4050b507493f661493c8090d67f52cd9caa739584a426dbb784511e1341a02e1

Observation 28285f16-f827-4531-b0a0-b2b05b9f7898 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Spatial-aware Vision Language Model for Autonomous Driving Bleu: a method for automatic evaluation of machine translation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:07.311902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:07.311902Z digest=sha256:7e7b4338e9aa3c4a397f7698a098aed42ffbd69e353ec1c9ed515d6576531328

Observation 2b846917-d0cb-4730-8cfb-202726c24275 · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario.

Spatial-aware Vision Language Model for Autonomous Driving Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:07.413343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:07.413343Z digest=sha256:f17cb2f5c9b09544932c7323c1ed2b8d38467e4cc8d0f42fd6dcd1d449ecaaa7

Observation 201f7912-e90a-4629-9f2d-e4a0d9dd052c · outbound

This paper cites Learning transferable visual models from natural language supervision.

Spatial-aware Vision Language Model for Autonomous Driving Learning transferable visual models from natural language supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:07.566009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:07.566009Z digest=sha256:2820ce1eca5f72f627359de71f01889e281dc44528d3d7fb5ff8ffe2f792ed61

Observation ce2f30e5-7a99-4e9a-94c5-3817d1a6cdd6 · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models.

Spatial-aware Vision Language Model for Autonomous Driving Lmdrive: Closed-loop end-to-end driving with large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:07.681334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:07.681334Z digest=sha256:84b1a2c16e70316c7f346d9e3d17520b32670c76b84e1a872a55083ae6a2d2f2

Observation eacd45d3-8d04-448f-a6c6-f81e1cea767c · outbound

This paper cites Drivelm: Driving with graph visual question answering.

Spatial-aware Vision Language Model for Autonomous Driving Drivelm: Driving with graph visual question answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:07.791241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:07.791241Z digest=sha256:09be3537ada7874d976f3a83304421f0ead6008c7a6bc77b67318a89e93a7b2c

Observation ad45164d-61ac-49c2-ac7b-df7de0ca19fd · outbound

This paper cites LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving.

Spatial-aware Vision Language Model for Autonomous Driving LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:07.904116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:07.904116Z digest=sha256:04cb56b1d1092ddc9201fe1333382b6c04d416dd2f576ab42d692836edd3ae94

Observation 4eb3b491-ba19-4ed5-971e-cc96d40ccfdf · outbound

This paper cites Bev-tsr: Text-scene retrieval in bev space for autonomous driving.

Spatial-aware Vision Language Model for Autonomous Driving Bev-tsr: Text-scene retrieval in bev space for autonomous driving

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.026106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.026106Z digest=sha256:211dc0a043a2f4f18c439934d0a79fbcedfe752b1e46bb6b79d22f1722852057

Observation 3eb9bfbf-1a90-419c-9d21-c705c4a95f34 · outbound

This paper cites NuScenes-SpatialQA: A Spa- tial Understanding and Reasoning Benchmark for Vision- Language Models in Autonomous Driving.

Spatial-aware Vision Language Model for Autonomous Driving NuScenes-SpatialQA: A Spa- tial Understanding and Reasoning Benchmark for Vision- Language Models in Autonomous Driving

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.097736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.097736Z digest=sha256:68a485c597cf494809aeb4ba173aa1291841c013e80b652f602acf8270631cc9

Observation 5c2147f5-72a2-44bf-ba6e-1e3ff9db0637 · outbound

This paper cites Drivevlm: The convergence of autonomous driving and large vision-language models.

Spatial-aware Vision Language Model for Autonomous Driving Drivevlm: The convergence of autonomous driving and large vision-language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.225979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.225979Z digest=sha256:27216a416b36d01e35caea7ab8cc24edac866abfcb4003c33d333316389e087c

Observation 86167633-8eff-40cf-a7e3-bd43e9c8d441 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Spatial-aware Vision Language Model for Autonomous Driving LLaMA: Open and Efficient Foundation Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.319483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.319483Z digest=sha256:a17f869c38331c573d4bd4b8f9a3ef8dec3d4eeab04e8c7474111ceede5a551f

Observation 52552ab9-9b24-4336-a3a1-b6f07b32dee8 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Spatial-aware Vision Language Model for Autonomous Driving Lawrence Zitnick, and Devi Parikh

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.409639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.409639Z digest=sha256:7e133b331120623b8db156574baa0913a1259e6e87fbb6f71ac3b05c573b84d2

Observation afc0a3d0-6c70-4c78-a501-0dc3712a4ffe · outbound

This paper cites Om- nidrive: A holistic vision-language dataset for autonomous driving with counterfactual reasoning.

Spatial-aware Vision Language Model for Autonomous Driving Om- nidrive: A holistic vision-language dataset for autonomous driving with counterfactual reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.503851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.503851Z digest=sha256:992dfabb7d2288756f22d2b159ef75c0c9253ab90d095e6636b76b192cdf1e21

Observation 708dff54-b1a7-48f7-ba96-38b98f94b9a9 · outbound

This paper cites Mv2dfusion: Leveraging modality-specific object semantics for multi-modal 3d detection.

Spatial-aware Vision Language Model for Autonomous Driving Mv2dfusion: Leveraging modality-specific object semantics for multi-modal 3d detection

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.571138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.571138Z digest=sha256:ae0314b56fd36c609822a266f8a3c75965ae75ce358b9490b7a08e86ef715641

Observation 52338f5f-7cba-4683-ae4e-57c4ce76ae67 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Spatial-aware Vision Language Model for Autonomous Driving DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.699183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.699183Z digest=sha256:26788c0aef5662027eecd28cda1c9f566e03f25413fa90af43cd12e47a53feb3

Observation 44684c65-3826-4ad2-8ece-63d363117a4e · outbound

This paper cites Segformer: Simple and efficient design for semantic segmentation with transformers.

Spatial-aware Vision Language Model for Autonomous Driving Segformer: Simple and efficient design for semantic segmentation with transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.749857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.749857Z digest=sha256:737f97ae8ffd9e6476b8b8a5ed7f0c7d0e27b8916c9295350cee24e1132e4e3a

Observation e4faeea0-69d6-4dd8-8689-841e7ee79514 · outbound

This paper cites ChatBEV: A Visual Language Model that Understands BEV Maps.

Spatial-aware Vision Language Model for Autonomous Driving ChatBEV: A Visual Language Model that Understands BEV Maps

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.821672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.821672Z digest=sha256:920f37631355f191b664f2a9591b222a5b5e26105cf944ea8f6942f0bd0d4ec4

Observation 068e2522-100c-427b-98cb-eb76d29dcef9 · outbound

This paper cites Explainable object-induced action decision for autonomous vehicles.

Spatial-aware Vision Language Model for Autonomous Driving Explainable object-induced action decision for autonomous vehicles

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.921762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.921762Z digest=sha256:bf7f617347d71b7f9e7d4d0124186d858780e574cad5b0747c34b42042811f75

Observation 938438ad-e962-4386-a0bd-79499374f4d9 · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model.IEEE Robotics and Automation Letters,.

Spatial-aware Vision Language Model for Autonomous Driving Drivegpt4: Interpretable end-to-end autonomous driving via large language model.IEEE Robotics and Automation Letters,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:08.989622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:08.989622Z digest=sha256:1d642184420f2c93866b2590df27a0153f4ebb63713d89d5164fb39c155978a9

Observation 7c208d22-0ce8-4148-b091-0d01b01530dc · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

Spatial-aware Vision Language Model for Autonomous Driving Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.067884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.067884Z digest=sha256:cbedc25d095e4f6608abbd8fcb0d33c9ff4b24ded12ec9afc882332cb118ce30

Observation e18029f9-59bd-47b4-a680-f1dd31acf25c · outbound

This paper cites Lidar-llm: Exploring the potential of large language models for 3d lidar understanding.

Spatial-aware Vision Language Model for Autonomous Driving Lidar-llm: Exploring the potential of large language models for 3d lidar understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.152689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.152689Z digest=sha256:8918c6c003fdd4cd4c27494ba1d3bff541aace92da4a3f5b98fca6cd5b6eaf9e

Observation 5f15abbd-6293-4c0a-8219-9b91a60e4dff · outbound

This paper cites V2X-VLM: End-to-End V2X Cooperative Autonomous Driving Through Large Vision-Language Models.

Spatial-aware Vision Language Model for Autonomous Driving V2X-VLM: End-to-End V2X Cooperative Autonomous Driving Through Large Vision-Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.216484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.216484Z digest=sha256:b613b58b5b67b800f98adf7fb390077c41936a211e0206a3093905471a4e0f2e

Observation 272befc2-33ff-476e-88e3-e9949de9cb0b · outbound

This paper cites Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes.

Spatial-aware Vision Language Model for Autonomous Driving Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.305118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.305118Z digest=sha256:a8e6c2c54d6927f93e99cd60648e0f19c7d3b9ea6494d5b4f1e6bc686dd93448

Observation f4d06326-d0be-4bcb-808c-0d733ac72020 · outbound

This paper cites MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving.

Spatial-aware Vision Language Model for Autonomous Driving MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.388873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.388873Z digest=sha256:73eac331ca3917fa59d75f24849d25c33a6eb2d6f5236d8f0435255ddf8d5349

Observation ba3deffa-841f-43bf-ad7d-4480b9656f97 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of language models with Zero-init Attention.

Spatial-aware Vision Language Model for Autonomous Driving LLaMA-Adapter: Efficient Fine-tuning of language models with Zero-init Attention

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.457217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.457217Z digest=sha256:5c5023e5fa54bd2d7d139f4057bf4ac97858822719300318b4e8e303b4ac2371

Observation d794efb9-1426-47d2-9691-07e9fdc0309c · outbound

This paper cites Interndrive: A multimodal large language model for autonomous driving scenario understand- ing.

Spatial-aware Vision Language Model for Autonomous Driving Interndrive: A multimodal large language model for autonomous driving scenario understand- ing

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.545258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.545258Z digest=sha256:4851ffb665824b339e5350fb3fbc943705ba4f5b1f51311c358aac358423800e

Observation 8fece7bc-a640-4013-a865-114a2a02629b · outbound

This paper cites MPDrive: Improving spatial understanding with Marker-Based Prompt Learning for Autonomous Driving.

Spatial-aware Vision Language Model for Autonomous Driving MPDrive: Improving spatial understanding with Marker-Based Prompt Learning for Autonomous Driving

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.611502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.611502Z digest=sha256:5697264aa882b9f5dc08877f34c61e7bf4bc107eeb538215d77596cd78c3af96

Observation b5927311-ed47-412d-a84c-4e2573449a7b · outbound

This paper cites Opendrivevla: Towards end-to-end au- tonomous driving with large vision language action model.

Spatial-aware Vision Language Model for Autonomous Driving Opendrivevla: Towards end-to-end au- tonomous driving with large vision language action model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.703408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.703408Z digest=sha256:76f03a6be3b9bf3f800da670b81591c16cbe4f3d975b0145c99ad80e83e443ef

Observation 4012c0d0-1828-452a-b1ce-f1b2381eb1e2 · outbound

This paper cites Autovla: A vision- language-action model for end-to-end autonomous driving with adaptive reasoning and reinforcement fine-tuning.

Spatial-aware Vision Language Model for Autonomous Driving Autovla: A vision- language-action model for end-to-end autonomous driving with adaptive reasoning and reinforcement fine-tuning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.794400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.794400Z digest=sha256:4bb59b714474e19748608e5180f1edbd06ae6a339d9b9f5794afa5ec235056e7

Observation 04d7b78e-df96-4b02-8d2c-0914a2e56bac · outbound

This paper cites Further information about our pro- posed SA-QA dataset is presented in Sec.

Spatial-aware Vision Language Model for Autonomous Driving Further information about our pro- posed SA-QA dataset is presented in Sec

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.860635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.860635Z digest=sha256:e99d069f4fca982f796ffd1d06343470f553accb9c927090ebf0495d27caf19c

Observation fceeebf0-91a8-4a6f-9caa-c8564f193e0b · outbound

This paper cites The specific QA formats and the step-by-step gen- eration procedure are summarized in Tab.

Spatial-aware Vision Language Model for Autonomous Driving The specific QA formats and the step-by-step gen- eration procedure are summarized in Tab

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:09.979579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:09.979579Z digest=sha256:fd5bbcb36f731d3abf8f8336ef7cc436ef76eb64d386b11f22b9479713e44f4a

Observation de044b08-5e6c-4b27-b56c-8dcc7ffde62d · outbound

This paper cites Is object A closer than object B?.

Spatial-aware Vision Language Model for Autonomous Driving Is object A closer than object B?

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.017450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.017450Z digest=sha256:eb00b668b6c641cce11ac057a73b6e8fa5d0219b6db3d00a27fec8ba346eed57

Observation 4fce30f9-dd82-4f4b-b938-5d2c1fbaf9d5 · outbound

This paper cites For a potential future position at(x, y), is it in a drivable area?.

Spatial-aware Vision Language Model for Autonomous Driving For a potential future position at(x, y), is it in a drivable area?

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.101378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.101378Z digest=sha256:55a5861f7d9c945024d93b73733a69a2a3508b93c4293945723a42beed7f657e

Observation c115f5a0-c622-4e2c-a4f6-10bdd2493c07 · outbound

This paper cites an unresolved cited work.

Spatial-aware Vision Language Model for Autonomous Driving Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.191266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.191266Z digest=sha256:27d17f7e33e0a80362b3878cef8d135f3e88b9baf1cd7acb31a3af3c2611ebfa

Observation bd81bdf5-1734-4a99-8c78-b2ce5d2155ae · outbound

This paper cites Yes” if inside; “No.

Spatial-aware Vision Language Model for Autonomous Driving Yes” if inside; “No

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.255795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.255795Z digest=sha256:8041e8c45cb2797d3d80c0a094db58335b660f413963acd492000ed645fe30cb

Observation 47112303-12dc-4d5a-a247-dd3788dbcdbc · outbound

This paper cites an unresolved cited work.

Spatial-aware Vision Language Model for Autonomous Driving Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.347558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.347558Z digest=sha256:12500d1add5af9fc3c701a246b8ed23876c61d899b518fa59bf36c2cb9fa744d

Observation 8dabb2bf-b93e-4766-85d8-a3437915d02d · outbound

This paper cites an unresolved cited work.

Spatial-aware Vision Language Model for Autonomous Driving Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.422118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.422118Z digest=sha256:8834b60b88bfca2e150b9c068e54de722318e180018c2a155f7427927c49e83f

Observation 2213da90-c987-425b-8d69-07441124e101 · outbound

This paper cites The object is a <category>in the<CAM>, location:(x, y), length:<l>, width:<w>, height:<h>, angles in degree:<yaw>.

Spatial-aware Vision Language Model for Autonomous Driving The object is a <category>in the<CAM>, location:(x, y), length:<l>, width:<w>, height:<h>, angles in degree:<yaw>

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.514978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.514978Z digest=sha256:29fc90e44ca6faed530b99f36616d07bf61af3d94b87b654894e1d05b7087af8

Observation f27e396d-a9fe-4782-a02d-f438d7d293c9 · outbound

This paper cites an unresolved cited work.

Spatial-aware Vision Language Model for Autonomous Driving Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.577345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.577345Z digest=sha256:ff750ebaa27ae11623becdfa2ed7c01a20c1b06250e1f630263a36a2fd8a65f6

Observation 5a4d019b-1a92-4ae3-a70a-b30eeaaed55a · outbound

This paper cites Identify the object in the masked regionand describe its 3D information.

Spatial-aware Vision Language Model for Autonomous Driving Identify the object in the masked regionand describe its 3D information

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.641834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.641834Z digest=sha256:be7d3808f8d022d6aed3b943c0b25220d600901d696f2534ffb9898481860d31

Observation d4cd8fdf-2832-49a1-a543-6801f07af9e0 · outbound

This paper cites an unresolved cited work.

Spatial-aware Vision Language Model for Autonomous Driving Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.729467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.729467Z digest=sha256:e1a793662b981a13222a73a8d968aff3f4e5b5b40077d490a59574fcde165bf8

Observation 7d7e9b52-0204-42c1-84c9-b57c1f4cb600 · outbound

This paper cites What objects are on the lane defined by points (x1, y1),(x 2, y2),(x 3, y3)?.

Spatial-aware Vision Language Model for Autonomous Driving What objects are on the lane defined by points (x1, y1),(x 2, y2),(x 3, y3)?

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.780893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.780893Z digest=sha256:f6b60840449f30c81b36c1e10aeb8e043499c084d6a3afc4808d6a8d7991b0fc

Observation af62b3e4-1e65-484c-89df-7c5a172d96f2 · outbound

This paper cites an unresolved cited work.

Spatial-aware Vision Language Model for Autonomous Driving Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.869386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.869386Z digest=sha256:d074afa031697615e1cd5b54828849844a75a0731de3ba3e4895e3ee5b88905e

Observation ab44230c-96c0-4765-87ce-4d8d92543346 · outbound

This paper cites The object is a <category>, location.

Spatial-aware Vision Language Model for Autonomous Driving The object is a <category>, location

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:10.928164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:10.928164Z digest=sha256:d8bd2ea664e7997b89233d0e521cc6894785947428d90d769de6c94d793bd03d

Observation 35e25544-6373-4a8e-9ce3-d6e47220f55b · outbound

This paper cites an unresolved cited work.

Spatial-aware Vision Language Model for Autonomous Driving Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:11.008343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:11.008343Z digest=sha256:d08d6615afde107ef0f682dd4b836adf23cbc415d406452cf7fc799fdbab9352

Observation 047e199e-0500-48b6-8596-85c1de859690 · outbound

This paper cites an unresolved cited work.

Spatial-aware Vision Language Model for Autonomous Driving Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:11.095005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:11.095005Z digest=sha256:d7db961a95cdc03a205b1691cb58b7ce542a34c097819867207a9a3574c5ee7a

Observation 23cf0bcb-2166-42d5-a5f0-44f0e38b75e6 · outbound

This paper cites Please determine the metric distance (in meters) separating the two indicated objects.

Spatial-aware Vision Language Model for Autonomous Driving Please determine the metric distance (in meters) separating the two indicated objects

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:11.179600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:11.179600Z digest=sha256:43f52b7eae416798f454b71bb244e531113d506df0821404a14f137f0ac3440e

Observation 73b3e2a9-9cfe-4fbd-971f-d19a48dece12 · outbound

This paper cites an unresolved cited work.

Spatial-aware Vision Language Model for Autonomous Driving Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:11.237154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:11.237154Z digest=sha256:567a939055febf52cf45b234bea19d0c6951d8a4a36f915ea6d8c39815dc44a6

Observation ea2e33d8-9e27-420a-9a1b-66e1afae807f · outbound

This paper cites an unresolved cited work.

Spatial-aware Vision Language Model for Autonomous Driving Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:11.276700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:11.276700Z digest=sha256:a6d09cb627f8f792c8f3e8dc57e525c42c6b974a8787bbba4f407c3498d44bef

Observation 5f276cf3-b1b2-4be0-95d7-4169163d0ca8 · outbound

This paper cites D.” (The value is rounded to 0.1 meters). SR-04 “What is the future position of the object at(x, y)afterT second?.

Spatial-aware Vision Language Model for Autonomous Driving D.” (The value is rounded to 0.1 meters). SR-04 “What is the future position of the object at(x, y)afterT second?

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:11.387856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:11.387856Z digest=sha256:755405b8bbfe01912884a01d2083e1e69d7562e46078def794bf62258dd186f8

Observation 726a9024-5ecc-4846-90f9-ab9ee2a94d7e · outbound

This paper cites (x f ut, yf ut).

Spatial-aware Vision Language Model for Autonomous Driving (x f ut, yf ut)

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T13:26:11.454206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:26:11.454206Z digest=sha256:9a03990804abd4a9e69a0cfe48d6eb26df4f2a27d9fc3829eb300a3c7e52ac5f

Pith citing papers

Observation c9a16e7c-c5b9-4780-bb07-e267be93ce5e · inbound

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks cites this paper.

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks Spatial-aware Vision Language Model for Autonomous Driving

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:09.079975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:58:28.202606Z digest=sha256:1e4f11d6d4436c27b7a528e6ccb85015a08a96d62ede2ca1b0c80ba0ab8e055c

Observation bfcf20d8-8e6d-40c8-9902-7eac5ea03e74 · inbound

Steadily moving semi-infinite fracture in plane poroelasticity cites this paper.

Steadily moving semi-infinite fracture in plane poroelasticity Spatial-aware Vision Language Model for Autonomous Driving

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-07-05T11:41:02.601691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-05T11:39:05.686584Z digest=sha256:8c01a6eefe3af24ca8576eeeaa994579f4fd1afab12818107c7d3c5d41e43768

Observation b853a694-4a69-44e8-a82f-8e365b053eaa · inbound

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments cites this paper.

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments Spatial-aware Vision Language Model for Autonomous Driving

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:09.079975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:46:36.865150Z digest=sha256:a308290a8ec8c7c859223b3d73d25160a46486cffa950ac902f229761f60ea24

Observation 46d5025b-8fad-45e6-9c66-1a8faf01e4eb · inbound

From Scene to Object: Text-Guided Dual-Gaze Prediction cites this paper.

From Scene to Object: Text-Guided Dual-Gaze Prediction Spatial-aware Vision Language Model for Autonomous Driving

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-26T03:04:09.079975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:03:02.394571Z digest=sha256:0b8dfb3f1609b860be99474792350dd6519f3a2aa369ba6537a77fa8dbe263af

Observation e2998a1a-a9a4-4dc6-ad2b-6f34c56712a7 · inbound

nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving cites this paper.

nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving Spatial-aware Vision Language Model for Autonomous Driving

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:26:00.313523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:35:04.552901Z digest=sha256:f8e0b6d1a119bf1debaa327ec54e163b4a6463fa6a2da1911f540881f8ebd936