Pith. sign in

Paper Citation Record · LEDGER

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving

As of 19 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 0 inbound Pith citation observations for arXiv:2605.23176.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.23176 v2

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T16:40:22.441025Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 111 outbound references displayed

  • verified exact47
  • verified fuzzy53
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4124a936-85e4-479c-b649-9b33c7e313a0 · outbound

This paper cites Studies in spatial learning.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Studies in spatial learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:39e4bc66a155ba667853c29298533c46d5d88e2c362c7d01410695d076014419

Observation 4908d26a-f4f6-452a-88d4-23c0b5221b75 · outbound

This paper cites Cognitive maps in rats and men.Psychological review, 55(4):189.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Cognitive maps in rats and men.Psychological review, 55(4):189

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.796120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:e7f56d4e8aceb588b6c8716401d90692397665a2fd4283584a9867cde33ab5a9

Observation 46ad2d0d-0406-4c71-b866-80c32161dc59 · outbound

This paper cites The hippocampus and context revisited.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving The hippocampus and context revisited

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.747734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:6baf92a06e25b45341b1f24d6c66ec87bc44d31c1accf555f52d086e60316a52

Observation fdcd4ef5-bed5-4f43-bab2-81c653fdd9e4 · outbound

This paper cites The effect of vehicle navigation systems on the formation of cognitive maps.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving The effect of vehicle navigation systems on the formation of cognitive maps

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.755008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:534ac8011d612b5852bc12aa8e34fa17b1bfaa57cfbd2e88a024adc9458a8584

Observation db2ce18a-30d3-4a8d-a831-f291b7ef0e26 · outbound

This paper cites Omnidrive: A holistic vision-language dataset for autonomous driving with counterfactual reasoning.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Omnidrive: A holistic vision-language dataset for autonomous driving with counterfactual reasoning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.793920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:ff13952db9b6721700d4049f3b974b2399d5b022944c6f5e0ed1fbaf0b387ff6

Observation 2e812598-d752-44f9-bf5f-43d0e654119c · outbound

This paper cites Is ego status all you need for open-loop end-to-end autonomous driving? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14864–14873.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Is ego status all you need for open-loop end-to-end autonomous driving? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14864–14873

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.773907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:8b121f05efbdaad8f8312cdc9d04a808283c215ba322e0b9bb1fa1dfaea81675

Observation 018ae19c-0741-41cc-a8af-aa4579592b00 · outbound

This paper cites Holistic autonomous driving understanding by bird’s-eye-view injected multi-modal large models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Holistic autonomous driving understanding by bird’s-eye-view injected multi-modal large models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.802326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:cae51369e5774c9990fe9b1eed197dcac56a8db2fca4d2ac70156fc811ae99a5

Observation 7e1c8a1d-1d24-46fc-ba36-c8f34e5f375b · outbound

This paper cites Are vlms ready for autonomous driving? an empirical study from the reliability, data and metric perspectives.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Are vlms ready for autonomous driving? an empirical study from the reliability, data and metric perspectives

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.769424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:54ab22a8e851951c27344dcda73f448d3b662bd6e9acfa77592f35918f2c7181

Observation 5f0159d9-0918-494f-a48c-fed2b967c682 · outbound

This paper cites Vlm-ad: End-to-end autonomous driving through vision-language model supervision.Conference on Robot Learning (CoRL).

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Vlm-ad: End-to-end autonomous driving through vision-language model supervision.Conference on Robot Learning (CoRL)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.708689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:968320ace5e3b51c4ddb20f5c9975394b2e3689a123f66ef3aa8f1cf1890ef72

Observation 7d9f6ce7-ba05-4d31-a037-fa46e9da8f87 · outbound

This paper cites Robotron-drive: All-in-one large multimodal model for autonomous driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Robotron-drive: All-in-one large multimodal model for autonomous driving

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.720240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:4bf429cd8bd79bf62da79658a7dcccfab899cb78eaa90dcdb43348f96ad02238

Observation 8057aeb5-bcdc-4243-9d22-c010fc6826e3 · outbound

This paper cites Fastdrivevla: Efficient end-to-end driving via plug-and-play reconstruction- based token pruning.the Association for the Advancement of Artificial Intelligence (AAAI).

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Fastdrivevla: Efficient end-to-end driving via plug-and-play reconstruction- based token pruning.the Association for the Advancement of Artificial Intelligence (AAAI)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.782249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:f641fd4c915855bc6a6d6229c8a62fb590e75d5128e47492c5e1f3a86d6e3483

Observation 8ff366c2-099d-45d1-af50-29ca9c812ee1 · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Drivegpt4: Interpretable end-to-end autonomous driving via large language model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.710479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:0da2c58882caf528aff962795132e1990b833ee58c4f413f689d0e8736f9637e

Observation ee46642e-49ac-4018-a940-0eefe3c6dff2 · outbound

This paper cites Covla: Comprehensive vision-language-action dataset for autonomous driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Covla: Comprehensive vision-language-action dataset for autonomous driving

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.706881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:30bb16682b5b89021807d7a0600f00641a6ce987eae6fdb8f14565b8df5bda69

Observation 95333992-8ce8-431a-98af-e4f85118f85e · outbound

This paper cites Emma: End-to-end multimodal model for autonomous driving.TMLR.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Emma: End-to-end multimodal model for autonomous driving.TMLR

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.759266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:9d040f53a73a532494cdfc967e4cdd579c8e186046f4f16031c7944cf3ed89bb

Observation 7f698644-b9af-4646-aa19-6420fa57ad1d · outbound

This paper cites Drivelm: Driving with graph visual question answering.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Drivelm: Driving with graph visual question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.784193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:7681f8b671c78a94e0546ce20f3ae436df7455e8eef69b4ee362f860ee2290a5

Observation c46774f7-2fba-4f9a-bfec-b3732d1293be · outbound

This paper cites DriveVLM: The convergence of autonomous driving and large vision-language models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving DriveVLM: The convergence of autonomous driving and large vision-language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.790101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:0f40b0b96033169ecab2ed822b9d242ce4a0492d542482468348b5375e889d9c

Observation c98558d8-d110-42c0-bf32-aebc48894778 · outbound

This paper cites Futuresightdrive: Thinking visually with spatio-temporal cot for autonomous driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Futuresightdrive: Thinking visually with spatio-temporal cot for autonomous driving

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.788154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:25affc7717e012efc80cef1a228bead8ca989524d0715caa2aa8219457a728a4

Observation 656104bb-f775-47b5-86da-31557de501c6 · outbound

This paper cites NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.038675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:b9ddb86653c9478616aeee0524ec64ced824bd49b3d0cfdedbb159f7b466de04

Observation 0f352cc1-80c3-42c5-9de3-9c991e2e7446 · outbound

This paper cites Maplm: A real-world large-scale vision-language benchmark for map and traffic scene understanding.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Maplm: A real-world large-scale vision-language benchmark for map and traffic scene understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.765343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:369efce1f8b471ed1b546353dd19e6b7a4e900e6983c3ac436dafed8937e7165

Observation 0083e1b5-e35f-4800-8570-49436f921f03 · outbound

This paper cites Spatial reasoning with vision-language models in ego-centric multi-view scenes.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Spatial reasoning with vision-language models in ego-centric multi-view scenes

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.035595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:7d32a089aded1e9b3d10a504357ccf1fb10f3266f338153487da0f3fe051a08c

Observation 9431909e-311e-4b96-983b-43e5332e90e8 · outbound

This paper cites Surds: Benchmarking spatial understanding and reasoning in driving scenarios with vision language models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Surds: Benchmarking spatial understanding and reasoning in driving scenarios with vision language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.722106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:f36bdb1cdc4e51ab6c965405f9052b20478d220060bd13f8f7791edf5b77730d

Observation 4d30be20-c9bc-4ae0-9817-4dc6fbea3db5 · outbound

This paper cites Automated evaluation of large vision-language models on self-driving corner cases.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Automated evaluation of large vision-language models on self-driving corner cases

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.753145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:3a9216aed8049aacb4ce851e6f787afa2baaf622932f769691189daa870fe740

Observation 1a14d964-bc55-4fe0-8344-ab1fc1d163a0 · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Lingoqa: Visual question answering for autonomous driving

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.699401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:3419f276bcee17ac0e7ac72b6aaaab662c681bb79600286e9bcdacc23b755f6f

Observation fd719db5-702f-4167-a61d-1fba9ea4d46d · outbound

This paper cites Towards physics- informed spatial intelligence with human priors: An autonomous driving pilot study.International Conference on Learning Representations (ICLR).

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Towards physics- informed spatial intelligence with human priors: An autonomous driving pilot study.International Conference on Learning Representations (ICLR)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.771422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:2baa3cb1939f18c487ec0e6ffa7d7c71944165a8f6c8c54e43a40f7e70016983

Observation a979584e-9127-45f6-a598-6709d97ce2f0 · outbound

This paper cites Stsbench: A spatio-temporal scenario benchmark for multi-modal large language models in autonomous driving.Conference and Workshop on Neural Information Processing Systems.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Stsbench: A spatio-temporal scenario benchmark for multi-modal large language models in autonomous driving.Conference and Workshop on Neural Information Processing Systems

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.718326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:8a72c945ce3907b7b7204115347708c95b0f0a344f88c3ad65f24d1ab2a8a7ec

Observation 0bd6d633-47d7-4f8a-912f-e55b4bacd336 · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.729645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:3aeb38ae86fda81cb3d0c5e24ac9609d2416f41264fcd510eebea39f368c6286

Observation 0f86d378-36a6-4ec2-a281-8b4ddeaa8f8f · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving nuscenes: A multimodal dataset for autonomous driving

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.740741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:2286f30b76f2f819e66a4a2accd6ae8bf692989b16aa54ff90607cc2f08fc3ea

Observation dc1a85fc-bb8d-4ac2-b412-d6dc854221aa · outbound

This paper cites Argoverse 2: Next generation datasets for self-driving perception and forecasting.Conference on Neural Information Processing Systems, 202.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Argoverse 2: Next generation datasets for self-driving perception and forecasting.Conference on Neural Information Processing Systems, 202

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.778244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:213d25c603905a8d8fd33a0d575c8652f187175e906143576ab5e1ac4a457596

Observation 8aa9f8b7-4216-468e-bc81-167887f2c9c9 · outbound

This paper cites Man truckscenes: A multimodal dataset for autonomous trucking in diverse conditions.Advances in Neural Information Processing Systems, 37:62062–62082.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Man truckscenes: A multimodal dataset for autonomous trucking in diverse conditions.Advances in Neural Information Processing Systems, 37:62062–62082

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.703324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:a57c1bbf85fc129cc39663f00abefbded9a2d23fdba38f3ff33096f6f394d879

Observation b05017c4-67a5-4206-a04a-9b1b87cd9f62 · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Scalability in perception for autonomous driving: Waymo open dataset

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.786274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:058ffa30f6aeacc67f27baeed7ea76620a5217ee3267dfada0f382d9b16af92f

Observation cec05ce0-68c4-434e-8f79-d9661983e896 · outbound

This paper cites One million scenes for autonomous driving: Once dataset.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving One million scenes for autonomous driving: Once dataset

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.697734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:69d02524803397da131f5407ea4d8190675da914ec857fcf01898f02af6f2e36

Observation bf73d056-a96a-4243-9e60-ef44bd737356 · outbound

This paper cites Vision Language Models in Autonomous Driving: A Survey and Outlook.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Vision Language Models in Autonomous Driving: A Survey and Outlook

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.042065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:86adf632b115a9752990a2ce1b1c29c0da9276a73371ad4bc3792799b4eb63cc

Observation 2a7938e7-7066-4654-946c-825b7f270d77 · outbound

This paper cites A Survey on Multimodal Large Language Models for Autonomous Driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving A Survey on Multimodal Large Language Models for Autonomous Driving

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.048382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:0b4647d496c66a483ed2b429e7446772ced9461c1e4981f2b69311c5b3af3318

Observation d32995ca-0283-4a6c-b037-500b8d4bc502 · outbound

This paper cites Qwen3-VL Technical Report.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Qwen3-VL Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.058909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:19203bdda325117acddfc5a9565f135aa059fa1b7e394954491e81834c00ce3e

Observation 9655550d-5953-49d1-bf53-6d7c3ad10dff · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.071393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:f3df75f6efa668cf181ee820a64b2e42e50b1ca64faade87eacd86f69f8df692

Observation 54ed0b38-7617-458e-aef5-01720f9d8b9b · outbound

This paper cites GPT-4 Technical Report.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving GPT-4 Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.143908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:00aa179c8a50b5f9aca162b1f55db5ef21e90013c9862d93eadbd367a02fdf3e

Observation 3ffac277-ad70-467c-b57f-62b80fb750aa · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.029971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:ffa2697a25b725f02110cabc908c63659470d8b9d1af64184608d0c1b770be6c

Observation a1771ba1-042a-4573-a6a8-ab14f219d7e2 · outbound

This paper cites Kimi-VL Technical Report.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Kimi-VL Technical Report

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.005964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:aa5990fde4bd4cad6398ba659d7cf19b2e848e99316b206803f80fd04cdc98ad

Observation b74caffe-3ff8-41e1-ab94-d67456a15197 · outbound

This paper cites Llava-onevision: Easy visual task transfer.TMLR.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Llava-onevision: Easy visual task transfer.TMLR

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.804237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:174c331bcc55e09d2cf574245644da2719efd9fa0d3fabbf82a67760acfff46f

Observation fcf8f26a-70fb-4723-98a8-9426b2272fe5 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.054097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:6ba8b8d8f789063a16ef379a5dd45356fc3d25a7697f2a5cd9d357c2ff8ade37

Observation fe67a38c-3f1e-4445-8eae-97c8cec5f9ea · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Gemini: A Family of Highly Capable Multimodal Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.101355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:b02685acd59d0dda7c8a70012e58167701eea232e6cf11b92f78470c88975e4c

Observation 85a59b92-6efb-48bc-9707-93e9dee589eb · outbound

This paper cites Henasy: Learning to assemble scene-entities for interpretable egocentric video-language model.Advances in Neural Information Processing Systems, 37:86483–86499.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Henasy: Learning to assemble scene-entities for interpretable egocentric video-language model.Advances in Neural Information Processing Systems, 37:86483–86499

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.776006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:129063b267b90b50edfbf0bc60c9d9396ba95ea064977ed5e7aed3aaf5b97d1e

Observation d249392d-32e4-411a-8869-fa1b7a185066 · outbound

This paper cites Directed- tokens: A robust multi-modality alignment approach to large language-vision models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Directed- tokens: A robust multi-modality alignment approach to large language-vision models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.104225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:51bc0114ea80491705c03c359d7ffa45880c93f6f094f54b3016d316dcf259e6

Observation 163974b4-f9eb-4139-8f3f-77a3e8ea6fb1 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving GAIA-1: A Generative World Model for Autonomous Driving

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.106523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:6849ab40513a47ebc2c2028aa01daa8f093db9f8230ab6afe5ef78374564adc0

Observation 75e28c7f-a03a-491c-8cdb-8b5941800416 · outbound

This paper cites Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.757307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:94bd384d1c4a1b53bf4b359645081f3bc4dca7cb1af315e28c23ce840c9b8ee2

Observation 351df3a5-d6fa-4f0f-b19a-5b3d6d918a8c · outbound

This paper cites DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.117349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:1d481802dafba34d707d56c98afc0cd163757f111346108b55fc4dbc67c378a0

Observation edce7211-b13f-4dcf-8bac-e9b4bcc8dd60 · outbound

This paper cites Dolphins: Multimodal language model for driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Dolphins: Multimodal language model for driving

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.731435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:6ebc5d773be9cc83d088fa9938b0a32f5c60a3124039b2abb923ab3fc234b081

Observation 8f7fd824-9804-4732-93aa-f09306f1fe2f · outbound

This paper cites Rea- son2drive: Towards interpretable and chain-based reasoning for autonomous driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Rea- son2drive: Towards interpretable and chain-based reasoning for autonomous driving

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.705102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:e5c1ff634469a02d7a44353ed879075466434f1f463395c137d5f68b2e70220d

Observation c54f2ed0-211a-4d42-ae35-2ee7899b4b08 · outbound

This paper cites Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.737175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:3b1b6aec54c30251335f834b87f0ec1919ed6003f5d3d059c5e70a589435f8af

Observation 7713b246-3966-4641-94a6-a2f2fbcaf864 · outbound

This paper cites An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.119897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:8c9253dc86089dcfaccb9fb011a71e2e9bce02837f21bb0e64388eff36c15535

Observation 77b10584-e331-485b-a269-77c67ee6e5bb · outbound

This paper cites EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.133446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:601202bee2c10f529f63103ee0156b421797d2c7be599464cf0167c14c42ac0b

Observation ddf145be-575c-4c10-828b-04d397eae72d · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.Advances in Neural Information Processing Systems, 37:135062–135093.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Spatialrgpt: Grounded spatial reasoning in vision-language models.Advances in Neural Information Processing Systems, 37:135062–135093

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.735349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:d9940d0ac75e69ca811c8545089cf20c4a25a0afc1ba333a002419095d748a0c

Observation 061fa39c-3333-4047-b12e-fbe3b6d5bba3 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.690286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:63422695be0cea6c54d9ec92c43ad5aa2da3b1d9f6559288fc2b1805f3d01fc1

Observation 7ff65849-f172-4b62-a8af-23eb11dbd7a0 · outbound

This paper cites Spatialbot: Precise spatial understanding with vision language models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Spatialbot: Precise spatial understanding with vision language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.806450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:a8be0a829ccd2aa5fd7ce8378fc2493d36afaf52aec5ea3b014b00341b483924

Observation fdf4a296-7595-4a16-b8cb-82407f15c720 · outbound

This paper cites Spatialllm: A compound 3d- informed design towards spatially-intelligent large multimodal models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Spatialllm: A compound 3d- informed design towards spatially-intelligent large multimodal models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.714442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:c9791a4bcd188ac816de0e52943beb66eef8964a54aa0b973581391094c74d02

Observation 5f4702fa-09a1-4018-9624-164a1344e50f · outbound

This paper cites Robospa- tial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Robospa- tial: Teaching spatial understanding to 2d and 3d vision-language models for robotics

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.725664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:ffd2ab38ce4c27bf6860226529215da55791ce4aecefce548c7b9d3cd191ec4f

Observation 557905cf-2c09-43d6-912e-8cf1e0467720 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Gemini Robotics: Bringing AI into the Physical World

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.098982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:070e7a2b613f7e53f4eeedddcf6439fd2d8936be349c15e559195c6ac3f3a21b

Observation c178bf64-6215-4858-8932-b616a2f27228 · outbound

This paper cites SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.062222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:dc693148ca8906a249c39d4a71733b82b087c0901ab1d248da6bbfb7aa61f9f7

Observation 2bd05019-9f66-47b6-9271-19ef4ba2874e · outbound

This paper cites Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.068860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:0c150e00623760e2073847877b414110f9d5f0f5a50154353c363aa60302050d

Observation 45a9b7c3-7ce1-40dd-8ffe-f8184af67398 · outbound

This paper cites HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-08-12T02:23:50.971990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:28efe44b607aa4a60f8b77fdcc4455fe46a3d59310160ef25a91f95090084794

Observation 12e0fb8a-c213-4825-8799-72f80972fec6 · outbound

This paper cites Spatial-dise: A unified benchmark for evaluating spatial reasoning in vision-language models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Spatial-dise: A unified benchmark for evaluating spatial reasoning in vision-language models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.051540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:9597ea3c1c3de5e7e9f31e2fc40b3b572b8bada7f93964fae0376caa56649a1d

Observation d18c83e4-783f-4d0f-898b-e86c4781003e · outbound

This paper cites Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.045464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:f689872afc9a0afdc97b536c602087cad9c0f12d6c409c22facf4fa45894a6bf

Observation 1e7cab4a-e67c-4343-9651-9fc8462c1f2d · outbound

This paper cites Mmsi-video-bench: A holistic benchmark for video-based spatial intelligence.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Mmsi-video-bench: A holistic benchmark for video-based spatial intelligence

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.056684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:4ced9f5bde3f662de8ea795f223573167383271994ba5515e775373296ebee9a

Observation 4a26d7b4-17ff-41c5-920d-337d4fbb55bd · outbound

This paper cites Cambrian-S: Towards Spatial Supersensing in Video.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Cambrian-S: Towards Spatial Supersensing in Video

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.085193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:56e16efd30bacd3b472c26919c5ee80ea96a7f1880167e0e89edc65b0c770fc0

Observation db16a66f-347c-4084-b50b-cd9aab044778 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.792010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:ddd2101a6f3e3be17809b7dde9da36ca1a2e8eff43df9e4b0c8575a386c8b50f

Observation 5208b9b5-379e-47b1-a6a9-43a1855d0395 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Next-qa: Next phase of question-answering to explaining temporal actions

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.780191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:a624c5ad773b972b79c739903f6d59542c48f14b2ecab980e794ff8cddedfe23

Observation 103f2c57-2eb0-4448-8b8c-40e713f614d9 · outbound

This paper cites Visuospatial perspective taking in multimodal language models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Visuospatial perspective taking in multimodal language models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.148812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:df8f4c15c49e4eb73485712bc9962b7029fd0de785b4930c19b1e0534423fb19

Observation 1dc48536-9cac-4cc0-9600-e63ee7942e78 · outbound

This paper cites Egocentric Bias in Vision-Language Models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Egocentric Bias in Vision-Language Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-15T02:22:00.923596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:2b824d9b239c0fa0f2b4056e949863d8e496c531ed8eab3483114820a31ef124

Observation 2c136e4f-359a-4d3b-921e-bb845528047c · outbound

This paper cites Allocentric perceiver: Disentangling allocentric reasoning from egocentric visual priors via frame instantiation.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Allocentric perceiver: Disentangling allocentric reasoning from egocentric visual priors via frame instantiation

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.019175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:71049ddfb745e3bce2e2e27b2b905c93103b368c730f137b141f73b261956b53

Observation 9d3240a9-9e3e-41ae-be79-769d0542c6fc · outbound

This paper cites Keep it sympl: Symbolic projective layout for allocentric spatial reasoning in vision-language models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Keep it sympl: Symbolic projective layout for allocentric spatial reasoning in vision-language models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.024748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:541215fdde250695ec312af5d963a322a53f909482fcbf1c1ff11d75f778b78b

Observation e513a6c6-598c-42cc-946a-022c513cc61d · outbound

This paper cites CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.124516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:2ed91eeb60db8d9aed7af4c29689bfbd2f42764f833722dfc0e5fca8303bfa0b

Observation ff714df9-b2d9-4417-b1c4-c8c5723470a8 · outbound

This paper cites Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.027656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:282a0298f2e9b401b237e4432883cf2a04d94baf546958b393a32ce93bfe3557

Observation 1656cbd0-b151-4641-bd10-ddd9d3348516 · outbound

This paper cites Mind over space: Can multimodal large language models mentally navigate?.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Mind over space: Can multimodal large language models mentally navigate?

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.013763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:6a1641dfbce5e98c1c98be72ae5b3e9aee8527558f0e3fcc341ab03298340b17

Observation dccea137-34b3-4506-8ec4-22857003cfff · outbound

This paper cites Video2layout: Recall and reconstruct metric-grounded cognitive map for spatial reasoning.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Video2layout: Recall and reconstruct metric-grounded cognitive map for spatial reasoning

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.011148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:0525c31569074a200bd1de40cb22d8c001be13036a6343c9952ea9db844fc739

Observation d13a77b8-cd58-4f65-b461-274a47f11461 · outbound

This paper cites Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.141303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:b4293c139caa47b2b6ef70708d2b8fecc51824fb3097f253e99327d9e09f930c

Observation 6f4d57fd-d34d-4953-821b-dd0e4492c6de · outbound

This paper cites Talk2car: Taking control of your self-driving car.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Talk2car: Taking control of your self-driving car

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.749432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:8de80788795b7450d149c880590202de51f00d529460b59d3cd32ef308805fcc

Observation 053ee545-ae35-4e97-8647-5ceb13442531 · outbound

This paper cites Textual explanations for self-driving vehicles.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Textual explanations for self-driving vehicles

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.763268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:75dd34f48534237da0db63f194c5c1be9589c5380feffe341766d6bf31c95fe0

Observation 6a2d3aa1-88b4-4256-b559-2dbf53d8b5c5 · outbound

This paper cites NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets using Markup Annotations.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets using Markup Annotations

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.021817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:623dcd3263b9f0d2c805e014f9495243b2711c7c551b2934739680c4ee091bee

Observation a55cd804-8ef1-4680-a12c-5a9d8a4e68ec · outbound

This paper cites Stride-qa: Visual question answering dataset for spatiotemporal reasoning in urban driving scenes.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Stride-qa: Visual question answering dataset for spatiotemporal reasoning in urban driving scenes

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.130871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:d989e3e6cab9091b68e1438d3e8cb5ff4d51d0f2d522c1bfce9a9d958f241206

Observation caf445aa-9f86-451c-8aaa-8f277b2bfc7c · outbound

This paper cites Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.032836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:343fb076bdc690f634dd17ec3064fcc9c17599192ff9c660a89faadcb3c09491

Observation e8b1579c-2cc9-435d-9f7b-6116c8b0ad72 · outbound

This paper cites Road scene graph: A semantic graph-based scene representation dataset for intelligent vehicles.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Road scene graph: A semantic graph-based scene representation dataset for intelligent vehicles

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.701165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:a78901c11a2c560df94829e1c57d385f48d6be48e311ed8f9f76cb2a39c95620

Observation a7d5b92b-9b44-4f34-8a81-b6a08b159d92 · outbound

This paper cites RS2G: Data-Driven Scene-Graph Extraction and Embedding for Robust Autonomous Perception and Scenario Understanding.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving RS2G: Data-Driven Scene-Graph Extraction and Embedding for Robust Autonomous Perception and Scenario Understanding

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.127056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:9b8d0a92034d68b4ed4a4e1580047bd8b2099f9a1c4cbdb3879cf9a1b0f59a01

Observation ea6d30d4-b30e-4386-b7e3-d8cfa84120bf · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.135972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:6f47436ab77c81447f38e9de32ad5c38b5de5375539183dad3cc1bc513733a91

Observation f769e9eb-9f84-4297-9410-d03e38befd3a · outbound

This paper cites Qwen3 Technical Report.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Qwen3 Technical Report

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.138487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:dca6ed500b1808777c696f702f18f56aa24723a77ddb397987c5beeb0b74c475

Observation e8cafc9b-107a-40a8-a8f1-4980cfd7c725 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.108882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:1bdcdb309353943d9bc6ce2b43dbce77b464edf22cecaf6ea41b549098e2d07b

Observation 34a05e40-7952-4d2c-b237-17b401fe220c · outbound

This paper cites GPT-4o System Card.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving GPT-4o System Card

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.114512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:939f3e25a391a95c3534ce5e5e0df328d6e54246a3e2f77b443eba192ab33983

Observation 85e50d5d-3ac4-4ad4-94a2-931b68b53dd7 · outbound

This paper cites OpenAI GPT-5 System Card.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving OpenAI GPT-5 System Card

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.122000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:e761d7f71b1fad90e80ac9cfa1de70ef9e81748477d22cacadd66ecb6d650eb6

Observation 51583bb9-4d59-4c0a-863f-5179e44d18ed · outbound

This paper cites Gemma 3 technical report.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Gemma 3 technical report

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.744293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:0f3beacc098362b59130898f4d2698f5334cbd0474ee3b266e81857db33e1c9c

Observation 4d7619c8-83e9-44ca-ae72-7264adb5a8e4 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.146260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:b252e553414b737f537aac850f3e6853afb10f472df28305efa63facfa3cb76a

Observation dbf5427a-699c-44ac-9909-e24c7be5fac7 · outbound

This paper cites Long Context Transfer from Language to Vision.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Long Context Transfer from Language to Vision

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.090190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:ba4e01baacf8638178de875bef383bcefa32865be60e834763ba1dae8b491416

Observation 3379d742-13d3-4106-be51-1e142470df30 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:44:56.081947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:00d597ff028bf9aab5cca928a8718436e2834fc89217a755b181ecbed9ffe355

Observation f0702a1e-fb16-41ec-b180-3d71a3f1cdb4 · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Scaling spatial intelligence with multimodal foundation models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.008458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:dab3ad0ebc138be08df8560baff4781ffdc26035ef87323f260455e74ce33aa8

Observation 38dfe81d-c142-4e66-b0af-0bf30c46bbe1 · outbound

This paper cites SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-07T01:16:00.407423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:9b92295013db5464e1d0a1478f47856857c9ecb893e8e08cdf2614589ec94323

Observation f33e53a8-abdf-422a-b6a6-fafc3a2d35a6 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.695942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:f7d9999d114d47b5d89a70bbbbb68981fc8fcfd7bc47629dff8424c3baf2898e

Observation 587c41a5-2cd0-4360-a093-042a196a5092 · outbound

This paper cites Measuring massive multitask language understanding.International Conference on Learning Representations (ICLR).

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Measuring massive multitask language understanding.International Conference on Learning Representations (ICLR)

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.738967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:290594055490f466f1348dc3ff725ba98c7f90db9f1e8e3f532b7f19b7566c7d

Observation 05b314d4-1ef3-4fc2-9582-9a2ba799ee5a · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.723905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:f84e79b8e7a71d7e53fe90740074dfd11c7ec1e388805e4049e617e9d1e23a67

Observation 054cba32-fab1-442a-bc10-bc6f296c5a19 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.712450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:04e401dc07829d3cb9d8a08f3845aac7e04a114965db1de4e7c6e24575e91695

Observation d34e47c0-cc7c-4161-843c-21f29473bc91 · outbound

This paper cites A coefficient of agreement for nominal scales.Educational and psychological measure- ment, 20(1):37–46.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving A coefficient of agreement for nominal scales.Educational and psychological measure- ment, 20(1):37–46

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.733472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:ce6ad550078e1a776432d3f520189883e7b35755660bf92d4e5669a3a3cc6101

Observation b14ca13c-2e9b-4417-bc5b-f68011cbd662 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.800274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:e5753a6351a556c6c82111ca306f474a4552754da2b120ea6fd4ba6fc56e6275

Observation 2c90cc1f-c329-4edc-87a4-ed4957bc7ceb · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues?The Twelfth International Conference on Learning Representations.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Swe-bench: Can language models resolve real-world github issues?The Twelfth International Conference on Learning Representations

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T12:44:53.798270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:f0fdd5ec37071daf614e260109aa3bc25c10a121d8f1d17fb47e80374c0e9bc4

Pith citing papers

No inbound Pith citation observations are available.