Pith. sign in

Paper Citation Record · LEDGER

Embodied Scene Understanding for Vision Language Models via MetaVQA

As of 14 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2501.09167.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09167 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:15:04.751881Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:16.561169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:29:35.618998Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact3
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d489a10d-73b8-4a93-985f-b51d0b174dad · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Embodied Scene Understanding for Vision Language Models via MetaVQA RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.519489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.519489Z digest=sha256:e60531b3fbcd003873e57c88082d0f73f72b65e76afd297a1c7604c053b5ae71

Observation b788943c-3b2f-4173-8da4-e5441099f3d2 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Embodied Scene Understanding for Vision Language Models via MetaVQA OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.524899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.524899Z digest=sha256:925fd0e51d0a5550eddbedbbdb7a629a2c33759730c5288434a54baed05b0997

Observation 4a281461-9294-4aed-b7fb-8b7420e487e5 · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

Embodied Scene Understanding for Vision Language Models via MetaVQA DriveLM: Driving with Graph Visual Question Answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.535121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.535121Z digest=sha256:044b30e6d7728f056f27e02b1ac7697ce4d6223f1e45da13087dda483e3708de

Observation 763ff7de-1dcb-4148-9394-60d3cd0bf5ec · outbound

This paper cites Embodied Understanding of Driving Scenarios.

Embodied Scene Understanding for Vision Language Models via MetaVQA Embodied Understanding of Driving Scenarios

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.539684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.539684Z digest=sha256:eceb81252dfac9d8dc0b646b6ca66273b7e2b6635eedad692890db1a9d3b6a44

Observation 683832e1-6c8d-4ce9-873b-90946482e665 · outbound

This paper cites DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model.

Embodied Scene Understanding for Vision Language Models via MetaVQA DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.544857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.544857Z digest=sha256:b439f72f431a54233f985637fea740c198798b1882c4f0cc198b62b4af7ddffb

Observation 3f2c9c9a-b212-4ac2-8dbe-0795c4a7dacc · outbound

This paper cites Talk2car: Taking control of your self-driving car.

Embodied Scene Understanding for Vision Language Models via MetaVQA Talk2car: Taking control of your self-driving car

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.504923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.549818Z digest=sha256:8e7dec7f1b4b043f0efd8bed4e46a31487f53e1ae4b6bfba2cf7b5f988591c66

Observation 5fda9b66-d9ce-4dd8-9264-030c3100f772 · outbound

This paper cites AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving.

Embodied Scene Understanding for Vision Language Models via MetaVQA AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.490944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.554602Z digest=sha256:ad8102ce93654a966889b9b2ed408a85dc32ab1f9e5090fea0be2561bfe08728

Observation 4736ec31-de29-47e3-8185-a092854dd199 · outbound

This paper cites SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models.

Embodied Scene Understanding for Vision Language Models via MetaVQA SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.558811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.558811Z digest=sha256:39a57ac911dac9c3e12fc52e15e3880836540faaced2bc4e80e17b3ad6641add

Observation 1edfb3b8-3238-453f-8215-b3785e49666f · outbound

This paper cites Vista: A generalizable driving world model with high fidelity and versatile controllability.

Embodied Scene Understanding for Vision Language Models via MetaVQA Vista: A generalizable driving world model with high fidelity and versatile controllability

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.477747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.563246Z digest=sha256:2d4caf4e79a2b9d6eadc420e3314c2a6dcb5cb43af979c183b38952600ad342a

Observation 6e31db1e-8972-43d7-a458-52a0e7eb6ddb · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Embodied Scene Understanding for Vision Language Models via MetaVQA Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.567591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.567591Z digest=sha256:e841b088c8e9559ce384751ba62f0109121c09dd0f9a3c678c7dcdf8eab0c2da

Observation 26085164-2ccf-4ae1-a5a9-3efee1e0253b · outbound

This paper cites Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning.

Embodied Scene Understanding for Vision Language Models via MetaVQA Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.464637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.572079Z digest=sha256:82f18160017c9f089bd46491fc6d926bc1f59ba7b38009a92a009b6a793cf966

Observation 6d058d51-207b-4679-9cb5-a9f09920c44e · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yuxin Pan, Giancarlo Baldan, and Oscar Beijbom.

Embodied Scene Understanding for Vision Language Models via MetaVQA Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yuxin Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.448535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.576039Z digest=sha256:65a66181f681f92ee47fa1f06a650c379be4cf37642f7c40c2dbe1c4e008b393

Observation 021d19f9-1e90-4d7d-9db5-4df843e0dfbd · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset.

Embodied Scene Understanding for Vision Language Models via MetaVQA Scalability in perception for autonomous driving: Waymo open dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.434513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.580861Z digest=sha256:2bcbb441e45b26b2a8e28aaee5c2762ba5ed073015018a0058c80244b63ec37b

Observation f72c1596-de40-4631-98bd-cb21d6fdbe7f · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Embodied Scene Understanding for Vision Language Models via MetaVQA Improved baselines with visual instruction tuning, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.419688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.585536Z digest=sha256:4da5350e915e2830bf2b0eb257877e54f2a92a913eb1adc116749119f3c0c95b

Observation 138298ac-2959-49fd-afe0-2437fa49bdf8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Embodied Scene Understanding for Vision Language Models via MetaVQA LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.592076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.592076Z digest=sha256:b818baee7e0a1bd5c18dd4c93c4bdfd7f892a4836afbf3237daf94a856d3c508

Observation fd44d3d9-f800-447c-8ea6-e663f4001fa5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Embodied Scene Understanding for Vision Language Models via MetaVQA Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.597921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.597921Z digest=sha256:2450fe92d79881c88d8fd91d183f1816e50199b7363ee81a2073bbce14f9561e

Observation fcf7c656-65e2-4236-9d0b-1c2fd2a6d147 · outbound

This paper cites Chatgpt-4, 2024.

Embodied Scene Understanding for Vision Language Models via MetaVQA Chatgpt-4, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.407074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.603673Z digest=sha256:fa8c53aca5caac0f753b1ffa3ca582a65acde4c18548000fd5ae0a4b63c96922

Observation 7fe76bf7-530c-4f2e-8baa-bba6ea27d6d0 · outbound

This paper cites an unresolved cited work.

Embodied Scene Understanding for Vision Language Models via MetaVQA Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:15:05.395176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.609358Z digest=sha256:86980d396eed8af641add60fd59a68cc3387f7bbd1954b3d7d30c60737bbc8d0

Observation 9d2a61b7-ea0c-4371-b571-ec73ea003504 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Embodied Scene Understanding for Vision Language Models via MetaVQA InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.614077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.614077Z digest=sha256:dc2a0ca03851ee56b8c19a778b06bbc1bcd8c5d60e579926228cb2d64cc88e4a

Observation ac97deee-3b08-47a9-b67b-f53665dc10bd · outbound

This paper cites Language Prompt for Autonomous Driving.

Embodied Scene Understanding for Vision Language Models via MetaVQA Language Prompt for Autonomous Driving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.619541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.619541Z digest=sha256:29ae20f46673669ee8271316b678bd3aa10d7b80dd4cd6dba9a6194b19bc8b1f

Observation 2d18145c-b355-49cc-ace6-13ae1eec12e2 · outbound

This paper cites Explainable object-induced action decision for autonomous vehicles.

Embodied Scene Understanding for Vision Language Models via MetaVQA Explainable object-induced action decision for autonomous vehicles

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.383207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.623978Z digest=sha256:d4c74b4179eaeaca4e215cda61d41ef56f944b376c5c5adee6ce0bb3536cbffb

Observation 93d19153-179e-4538-88eb-771d698fef6e · outbound

This paper cites Referring Multi- Object tracking.

Embodied Scene Understanding for Vision Language Models via MetaVQA Referring Multi- Object tracking

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.370273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.629000Z digest=sha256:f9106ced149185ab4e9d1e2388e9c822b2fd9d7018514e412e00001596709a99

Observation 19608b59-cdb7-461b-b43c-e7825487810f · outbound

This paper cites Open-sourced data ecosystem in autonomous driving: the present and future, 2024.

Embodied Scene Understanding for Vision Language Models via MetaVQA Open-sourced data ecosystem in autonomous driving: the present and future, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.357920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.633312Z digest=sha256:31c3440787ce94a23bf8a6944347307ac0eb20c02e85b68461b538906a5c6277

Observation bd3cf905-5523-4ca2-9f2d-37145d77f7ca · outbound

This paper cites Grounding human-to-vehicle advice for self-driving vehicles.

Embodied Scene Understanding for Vision Language Models via MetaVQA Grounding human-to-vehicle advice for self-driving vehicles

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.344963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.637207Z digest=sha256:85d5bcace67dc10f497ad51c95fc9a579a625abe6467f9a5e0f3f0ac0c969ba2

Observation 1b94b15b-5209-44b3-a84d-15a37a524053 · outbound

This paper cites Can you text what is happening? Integrating pre-trained language encoders into trajectory prediction models for autonomous driving.

Embodied Scene Understanding for Vision Language Models via MetaVQA Can you text what is happening? Integrating pre-trained language encoders into trajectory prediction models for autonomous driving

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:15:04.970413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.641767Z digest=sha256:20d3f1e65bb664d85fcc89aa1c9dd0bee3811a16232e8b6c9bfbf0b87458f00d

Observation 966fcf28-0e93-4953-8b19-03d660adcb53 · outbound

This paper cites Drama: Joint risk localization and captioning in driving.

Embodied Scene Understanding for Vision Language Models via MetaVQA Drama: Joint risk localization and captioning in driving

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.331836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.647422Z digest=sha256:739bbb9cef5510bfd0a02aca49424c9d601c4cc7764dc9d14e7e1e320b398f3b

Observation 489f0890-5bbf-408e-b307-bfb1aa00e251 · outbound

This paper cites Driving through the Concept Gridlock: Unraveling Explainability Bottlenecks in Automated Driving.

Embodied Scene Understanding for Vision Language Models via MetaVQA Driving through the Concept Gridlock: Unraveling Explainability Bottlenecks in Automated Driving

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:15:04.946475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.652141Z digest=sha256:e937a3981fcd000923c3b194339594b9e14a1ffdb76ba802b9a09dbde9b31254

Observation 8d47fadd-530c-4454-b21e-1ced9eeb0264 · outbound

This paper cites HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving.

Embodied Scene Understanding for Vision Language Models via MetaVQA HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.656528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.656528Z digest=sha256:de9dfb9a472759d42cea6612cd1a17ba1f7d025e879bdaba55c56aff9d660e75

Observation a9174a74-db69-41de-8a78-33c1b981730b · outbound

This paper cites Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving.

Embodied Scene Understanding for Vision Language Models via MetaVQA Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.660381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.660381Z digest=sha256:820a9378530741ab608085e83975540d8a9a8dab3e0da485ed53f7fa2f4f91ff

Observation b1017120-30ea-4961-827e-10d4c064449a · outbound

This paper cites ADAPT: Action-aware Driving Caption Transformer.

Embodied Scene Understanding for Vision Language Models via MetaVQA ADAPT: Action-aware Driving Caption Transformer

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:15:04.890239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.664367Z digest=sha256:11fdd71ba9888223f6fb65c334917c4a63fa8af1ceb0abae5a678cb28dbd34bd

Observation 8909ee23-24e9-446b-89a6-e5faad423f4e · outbound

This paper cites Textual explanations for self-driving ve- hicles.

Embodied Scene Understanding for Vision Language Models via MetaVQA Textual explanations for self-driving ve- hicles

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.317037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.668192Z digest=sha256:11f9db80eb9e8cb8a1d2dbe2e9a32fe30ef1141f96257c11de439d1d1003abd2

Observation ff67f547-3bcc-48de-a6a8-5048ecd3c7f0 · outbound

This paper cites Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning.

Embodied Scene Understanding for Vision Language Models via MetaVQA Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.303179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.671802Z digest=sha256:18474e1d38b8935b8c4052e6ca880fccf7c2df0381d79359189dc13b3d3e35d9

Observation e354465f-33dd-4bde-a30b-76a686faf8fb · outbound

This paper cites NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario.

Embodied Scene Understanding for Vision Language Models via MetaVQA NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.675610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.675610Z digest=sha256:06058b2589e3e754330cae3bba3459135e588cc092138ea51c5b55a08be1191a

Observation 1db2c9ce-32d0-485e-af35-c2dd0cddc1da · outbound

This paper cites Semantic anomaly detection with large language models, 2023.

Embodied Scene Understanding for Vision Language Models via MetaVQA Semantic anomaly detection with large language models, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.289383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.680792Z digest=sha256:7f7f5b45789002f4da3f41d6924e54fde5fc3769abb1d5db83194af8f7603692

Observation 635f0923-e8ae-4538-9bc4-c39c089fba48 · outbound

This paper cites MotionLM: Multi-agent motion forecasting as language modeling.

Embodied Scene Understanding for Vision Language Models via MetaVQA MotionLM: Multi-agent motion forecasting as language modeling

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.275572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.686773Z digest=sha256:e9c71cc210c63ffbec49970d01fd9ea98c1e650432bbbef5f6964427bbfd1492

Observation 37d36429-7164-4288-8309-c21289629a19 · outbound

This paper cites LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving.

Embodied Scene Understanding for Vision Language Models via MetaVQA LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.690975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.690975Z digest=sha256:73881a94b7650475eab2be00fd73fb5cca3eda86d2150a8990911571cab26b05

Observation 3e307456-a02e-49cd-9a5b-8515d3de43d2 · outbound

This paper cites an unresolved cited work.

Embodied Scene Understanding for Vision Language Models via MetaVQA Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:15:05.262051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.697816Z digest=sha256:3df5278d610b6bab7a7af72588ad5688a1bf1a9465a72c67b984da1eb95eee7f

Observation 2c1a1695-60dd-40f5-973b-c0328fb79fdc · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

Embodied Scene Understanding for Vision Language Models via MetaVQA GPT-Driver: Learning to Drive with GPT

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.702849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.702849Z digest=sha256:1392e1e76a764a1743f44de6b062041a684899196d62901f58051e6c7ba0d60b

Observation 3d1a7015-cb02-4ae2-aa13-f107fb64958a · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

Embodied Scene Understanding for Vision Language Models via MetaVQA DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.707819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.707819Z digest=sha256:39d17ba272726388bf3157740f9abb2372dc9c305e94fbf8c80e5244605ea630

Observation 5387fbd0-a305-4c15-b479-6a412f37ae51 · outbound

This paper cites LimSim++: A Closed-Loop Platform for Deploying Multimodal LLMs in Autonomous Driving.

Embodied Scene Understanding for Vision Language Models via MetaVQA LimSim++: A Closed-Loop Platform for Deploying Multimodal LLMs in Autonomous Driving

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.714227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.714227Z digest=sha256:85eda740e80f4c87bfe34f32e4504b418fb671c2f93e9e84d0c85fb2bb3fd15a

Observation 4a49e77c-5dcc-4ba1-9da6-9a447cb72358 · outbound

This paper cites Carla: An open urban driv- ing simulator.

Embodied Scene Understanding for Vision Language Models via MetaVQA Carla: An open urban driv- ing simulator

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.247857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.718646Z digest=sha256:e93a9f9ea45df768643be5f2ed1ccc93e3ffb020876830d8f8bb5e1b07105704

Observation 5cec8735-1f1b-43f2-a7cd-a50ae4af1593 · outbound

This paper cites NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles.

Embodied Scene Understanding for Vision Language Models via MetaVQA NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.723676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.723676Z digest=sha256:2bf33c7a1f07d56f8898611bbca53caf62755776f9b2f84d29e80eb8ec4e9a49

Observation c86f3af8-ffdf-482b-b628-713c8539add2 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elemen- tary visual reasoning.

Embodied Scene Understanding for Vision Language Models via MetaVQA Clevr: A diagnostic dataset for compositional language and elemen- tary visual reasoning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.234400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.728496Z digest=sha256:d64c5d9e8dc49570fe9724e7e2679bd8b4c3b64b61027c591139b888be1a6cc0

Observation 57cffc86-d4bf-40db-b75f-ae742dd16eb1 · outbound

This paper cites Large scale interactive motion forecasting for autonomous driving: The waymo open mo- tion dataset.

Embodied Scene Understanding for Vision Language Models via MetaVQA Large scale interactive motion forecasting for autonomous driving: The waymo open mo- tion dataset

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.220777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.731810Z digest=sha256:f066331c660d1ad5f9f16099e4b7c1993cda1e9543968f278c684745aee42dca

Observation 2f5fb141-65cc-4a59-95b8-6f71eeaf43f6 · outbound

This paper cites Scenarionet: Open-source platform for large-scale traffic scenario simu- lation and modeling.

Embodied Scene Understanding for Vision Language Models via MetaVQA Scenarionet: Open-source platform for large-scale traffic scenario simu- lation and modeling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.204629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.734987Z digest=sha256:210aff59b7dd2cf81e27f4eb647a36869157689d14a781abc45bfbcb3e8c42c3

Observation 60e58ac0-5a51-4bdf-bc9e-da1ba6a929e6 · outbound

This paper cites Yolov3: An incremental improvement.

Embodied Scene Understanding for Vision Language Models via MetaVQA Yolov3: An incremental improvement

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.187965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.738213Z digest=sha256:2fe92aaeead293cb06cda478743969479d29a768ea985538dc06d06d26eb2397

Observation 43049e48-d581-4154-9301-ef5891f58417 · outbound

This paper cites Segment Anything.

Embodied Scene Understanding for Vision Language Models via MetaVQA Segment Anything

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.741154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.741154Z digest=sha256:ecec37281b5e1b525e675da09796924f9d35851735d7af7dc51b944ddb4aa5be

Observation 657b7f8f-4b35-40d7-896d-fea2b99a1967 · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

Embodied Scene Understanding for Vision Language Models via MetaVQA Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:04.744522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:04.744522Z digest=sha256:1dd0d529ddafa4f0955c300e20f0dd2e3104608d042dd5f8bfb36f8cabfd91ce

Observation 69d2d939-85c8-4bce-a445-0805d9aa3540 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Embodied Scene Understanding for Vision Language Models via MetaVQA Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.174468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.748164Z digest=sha256:d6aa0624a17232154c065476ab7bfa686b3f9180e7179de73ac73def0152331e

Observation 37d1f5ed-5aa4-4095-a337-55223922ccb3 · outbound

This paper cites 3” level of visibility or is scanned by less than five rays of Lidar, then it is considered.

Embodied Scene Understanding for Vision Language Models via MetaVQA 3” level of visibility or is scanned by less than five rays of Lidar, then it is considered

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:15:05.160156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:15:04.751881Z digest=sha256:7aedb157f48ba2e1caaf157484481ed7924b82a9d047df6aabfacdb842fe1705

Pith citing papers

Observation 6540e05e-b877-4a08-8c87-cb52f2737317 · inbound

Dreamland: Controllable World Creation with Simulator and Generative Models cites this paper.

Dreamland: Controllable World Creation with Simulator and Generative Models Embodied Scene Understanding for Vision Language Models via MetaVQA

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:16.561169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:16.561169Z digest=sha256:675d9016d40234a4c7e3851b8e5301c3653c09980d245b310a3a563679cb74fa

Observation 5d27a2dc-8b8b-44a5-a981-d70e7e2e29f2 · inbound

Vesta: A Generalist Embodied Reasoning Model cites this paper.

Vesta: A Generalist Embodied Reasoning Model Embodied Scene Understanding for Vision Language Models via MetaVQA

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:29:35.620749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T16:55:12.518255Z digest=sha256:b59dd2a0a69d260a65db9967539351c8017665e490b294f70cd68cc75da0ed44

Observation fd61dd70-9e2c-45f7-ae68-c961c5ad0e27 · inbound

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers cites this paper.

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers Embodied Scene Understanding for Vision Language Models via MetaVQA

Reference 143

Resolution
unresolved
no resolver link, observed 2026-08-01T07:02:52.563375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:02:52.563375Z digest=sha256:e86f246c782d329dcd97a0f78fc535256a6131e0338d58949c2beb6c17055459