Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:45:47.406893Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2411.13211.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:45:47.406893Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fcc889a0-0f25-4e85-9f8f-3a7b480f2842 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Mas- tering the game of go with deep neural networks and tree search
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 534fdd08-0d02-4d4d-bc0a-d5ea9d2d8de7 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7042915-8ccb-4735-8104-11b6b7592550 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Specification gaming: the flip side of ai ingenuity
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e93a6f76-1f30-4f21-868b-f11ca058f76c · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Defining and characterizing reward gaming
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ae624301-8431-4ccc-9a5a-72ae93f62973 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? On the importance of hyperparameter optimization for model-based reinforcement learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ee53c557-2a5a-4e63-a00a-7aada7a3959c · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Variational inverse control with events: A general framework for data-driven reward definition
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b13ec86d-3df7-422f-bb34-42b565521e38 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Deep reinforcement learning from human preferences
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c616136-d84a-46b6-ad8c-e02bdf1af33c · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae1fa5bf-97f7-4b05-8bdb-1c3086adaa46 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? RoboCLIP: One Demonstration is Enough to Learn Robot Policies
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa9aa0c6-3229-493d-a173-15357e8c13cb · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Task Success is not Enough: Investigating the Use of Video-Language Models as Behavior Critics for Catching Undesirable Agent Behaviors
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0169577-c9f9-4b4f-aace-dcf08d21b2bf · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6718bc60-6efa-4b08-99ee-8934da57ee5c · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? The unsur- prising effectiveness of pre-trained vision models for control
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e8819983-5ecb-496f-bbce-4df5d85468f5 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Reward Design with Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed1eaabf-97a8-4d83-ba2c-94c173cec3a2 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? RewardBench: Evaluating Reward Models for Language Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d89699d-a082-48fc-9303-2e05cfd4779b · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Reward learning from narrated demonstrations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 710e80fd-f955-423f-8ebe-7018aaa971bc · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Zero-shot reward specification via grounded natural language
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc57d9c5-0e98-4c8f-95e1-cf4e4564c3f0 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e93e662b-5681-4cfe-b64a-ed702c29e75b · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd236a1c-d166-4044-b0fc-81fa4e9e125e · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Paxion: Patching Action Knowledge in Video-Language Foundation Models,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 12d3fae1-9b50-49d4-b082-a0c025bed755 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5874afd9-ede9-4663-bd7a-ac3b4119de37 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? VirtualHome: Simulating Household Activities via Programs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71db451c-40d2-45f9-8b25-7e66452edb7b · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Habitat: A Platform for Embodied AI Research, 2019
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bb1bd04c-aecc-4cc9-89e0-8037a97f1019 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2279356c-1ac0-415b-9d63-798035e8fa5d · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0770c95-1ffe-4f66-bf58-fe6bfe25b17b · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? A Short Note on the Kinetics-700 Human Action Dataset
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d0ccb6-c9d5-4759-9588-a02d4e5e6e0b · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? BEDD: the minerl BASALT evaluation and demonstrations dataset for training and benchmarking agents that solve fuzzy tasks
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7826661d-af97-436e-821c-bab9a859cffa · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2683055-6a7f-4111-9dd9-1f7db5471581 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Learning transferable visual models from natural language supervision
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e028cb0c-5dc6-43e0-a6a0-b2a019949f23 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Reproducible scaling laws for contrastive language-image learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb16373-0c7a-453d-b7b7-d9eb3609c871 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1c1dc7f-dd84-4c26-a1f5-56965d91caf5 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5599288-1fb0-4b3a-a851-82adc3c08e5a · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? GPT-4o System Card
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b56b9f6a-7f3e-411b-b0e0-2f5d1da7985b · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0d463477-81d5-4b03-80d1-0e4d632067d4 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? likely does not describe the video
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 30a37c82-3666-4da5-88af-79c089e76955 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Player holds an oak fence block in their hands
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 70e1214b-42bd-4e42-9f27-ab6895ef9a60 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation df5360ce-2569-4aa7-b5ad-0b5ad313d235 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a87f0140-0eb4-4f33-af8e-bb5b875d2568 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 57397734-49ee-4851-95fa-7ba559af7cbb · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1514cef4-3eb7-405d-aa8d-230ac9932723 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 17260624-6b44-4002-beb6-dc5def08f35c · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6b87587f-48bf-46ac-b72a-5ea90ab7105e · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 23868eb5-550a-4cdc-aaaf-bc4eb6984541 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 89dd84ec-b538-41e0-aba9-af172bd47aad · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Figure 10: The prompt used to obtain frame-by-frame descriptions for Minecraft videos
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 70ef0f7b-7277-46a0-89bb-879759f4bd46 · outbound
ViSTa Dataset: Do vision-language models understand sequential tasks? Paxion: Patching Action Knowledge in Video-Language Foundation Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.