Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:15:04.751881Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2501.09167.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:15:04.751881Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:16.561169Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T04:29:35.618998Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d489a10d-73b8-4a93-985f-b51d0b174dad · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b788943c-3b2f-4173-8da4-e5441099f3d2 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA OpenVLA: An Open-Source Vision-Language-Action Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a281461-9294-4aed-b7fb-8b7420e487e5 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA DriveLM: Driving with Graph Visual Question Answering
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 763ff7de-1dcb-4148-9394-60d3cd0bf5ec · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Embodied Understanding of Driving Scenarios
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683832e1-6c8d-4ce9-873b-90946482e665 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f2c9c9a-b212-4ac2-8dbe-0795c4a7dacc · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Talk2car: Taking control of your self-driving car
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5fda9b66-d9ce-4dd8-9264-030c3100f772 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4736ec31-de29-47e3-8185-a092854dd199 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1edfb3b8-3238-453f-8215-b3785e49666f · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Vista: A generalizable driving world model with high fidelity and versatile controllability
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6e31db1e-8972-43d7-a458-52a0e7eb6ddb · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26085164-2ccf-4ae1-a5a9-3efee1e0253b · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6d058d51-207b-4679-9cb5-a9f09920c44e · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yuxin Pan, Giancarlo Baldan, and Oscar Beijbom
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 021d19f9-1e90-4d7d-9db5-4df843e0dfbd · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Scalability in perception for autonomous driving: Waymo open dataset
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f72c1596-de40-4631-98bd-cb21d6fdbe7f · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Improved baselines with visual instruction tuning, 2023
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 138298ac-2959-49fd-afe0-2437fa49bdf8 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA LLaVA-OneVision: Easy Visual Task Transfer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd44d3d9-f800-447c-8ea6-e663f4001fa5 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcf7c656-65e2-4236-9d0b-1c2fd2a6d147 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Chatgpt-4, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7fe76bf7-530c-4f2e-8baa-bba6ea27d6d0 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9d2a61b7-ea0c-4371-b571-ec73ea003504 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac97deee-3b08-47a9-b67b-f53665dc10bd · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Language Prompt for Autonomous Driving
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d18145c-b355-49cc-ace6-13ae1eec12e2 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Explainable object-induced action decision for autonomous vehicles
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 93d19153-179e-4538-88eb-771d698fef6e · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Referring Multi- Object tracking
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 19608b59-cdb7-461b-b43c-e7825487810f · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Open-sourced data ecosystem in autonomous driving: the present and future, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bd3cf905-5523-4ca2-9f2d-37145d77f7ca · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Grounding human-to-vehicle advice for self-driving vehicles
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1b94b15b-5209-44b3-a84d-15a37a524053 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Can you text what is happening? Integrating pre-trained language encoders into trajectory prediction models for autonomous driving
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 966fcf28-0e93-4953-8b19-03d660adcb53 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Drama: Joint risk localization and captioning in driving
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 489f0890-5bbf-408e-b307-bfb1aa00e251 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Driving through the Concept Gridlock: Unraveling Explainability Bottlenecks in Automated Driving
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8d47fadd-530c-4454-b21e-1ced9eeb0264 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9174a74-db69-41de-8a78-33c1b981730b · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1017120-30ea-4961-827e-10d4c064449a · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA ADAPT: Action-aware Driving Caption Transformer
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8909ee23-24e9-446b-89a6-e5faad423f4e · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Textual explanations for self-driving ve- hicles
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ff67f547-3bcc-48de-a6a8-5048ecd3c7f0 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e354465f-33dd-4bde-a30b-76a686faf8fb · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db2c9ce-32d0-485e-af35-c2dd0cddc1da · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Semantic anomaly detection with large language models, 2023
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 635f0923-e8ae-4538-9bc4-c39c089fba48 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA MotionLM: Multi-agent motion forecasting as language modeling
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 37d36429-7164-4288-8309-c21289629a19 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e307456-a02e-49cd-9a5b-8515d3de43d2 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2c1a1695-60dd-40f5-973b-c0328fb79fdc · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA GPT-Driver: Learning to Drive with GPT
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d1a7015-cb02-4ae2-aa13-f107fb64958a · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5387fbd0-a305-4c15-b479-6a412f37ae51 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA LimSim++: A Closed-Loop Platform for Deploying Multimodal LLMs in Autonomous Driving
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a49e77c-5dcc-4ba1-9da6-9a447cb72358 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Carla: An open urban driv- ing simulator
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5cec8735-1f1b-43f2-a7cd-a50ae4af1593 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c86f3af8-ffdf-482b-b628-713c8539add2 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Clevr: A diagnostic dataset for compositional language and elemen- tary visual reasoning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 57cffc86-d4bf-40db-b75f-ae742dd16eb1 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Large scale interactive motion forecasting for autonomous driving: The waymo open mo- tion dataset
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2f5fb141-65cc-4a59-95b8-6f71eeaf43f6 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Scenarionet: Open-source platform for large-scale traffic scenario simu- lation and modeling
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 60e58ac0-5a51-4bdf-bc9e-da1ba6a929e6 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Yolov3: An incremental improvement
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 43049e48-d581-4154-9301-ef5891f58417 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Segment Anything
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 657b7f8f-4b35-40d7-896d-fea2b99a1967 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d2d939-85c8-4bce-a445-0805d9aa3540 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 37d1f5ed-5aa4-4095-a337-55223922ccb3 · outbound
Embodied Scene Understanding for Vision Language Models via MetaVQA 3” level of visibility or is scanned by less than five rays of Lidar, then it is considered
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6540e05e-b877-4a08-8c87-cb52f2737317 · inbound
Dreamland: Controllable World Creation with Simulator and Generative Models Embodied Scene Understanding for Vision Language Models via MetaVQA
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d27a2dc-8b8b-44a5-a981-d70e7e2e29f2 · inbound
Vesta: A Generalist Embodied Reasoning Model Embodied Scene Understanding for Vision Language Models via MetaVQA
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fd61dd70-9e2c-45f7-ae68-c961c5ad0e27 · inbound
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers Embodied Scene Understanding for Vision Language Models via MetaVQA
Reference 143
Source-reported events for the cited work
Unavailable: canonical work link unavailable.