Pith. sign in

Paper Citation Record · LEDGER

ViSTa Dataset: Do vision-language models understand sequential tasks?

As of 13 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2411.13211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13211 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:45:47.406893Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcc889a0-0f25-4e85-9f8f-3a7b480f2842 · outbound

This paper cites Mas- tering the game of go with deep neural networks and tree search.

ViSTa Dataset: Do vision-language models understand sequential tasks? Mas- tering the game of go with deep neural networks and tree search

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.161355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.161355Z digest=sha256:29d0e0244c1ace0f2ec7c90f12364adf9ff41c9af840127eb7052438ce71d047

Observation 534fdd08-0d02-4d4d-bc0a-d5ea9d2d8de7 · outbound

This paper cites Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes.

ViSTa Dataset: Do vision-language models understand sequential tasks? Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.168907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.168907Z digest=sha256:1e03efeb25e49b8efe170a9503bf4de5ee441c128eb47e17ba7da1cea05b5b42

Observation e7042915-8ccb-4735-8104-11b6b7592550 · outbound

This paper cites Specification gaming: the flip side of ai ingenuity.

ViSTa Dataset: Do vision-language models understand sequential tasks? Specification gaming: the flip side of ai ingenuity

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.155246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.176949Z digest=sha256:2914bb3116aa3f5f7db52b6d439f34393fc3d37ad7f69ca0690f3264e0460452

Observation e93a6f76-1f30-4f21-868b-f11ca058f76c · outbound

This paper cites Defining and characterizing reward gaming.

ViSTa Dataset: Do vision-language models understand sequential tasks? Defining and characterizing reward gaming

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.138168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.184902Z digest=sha256:39d75d660cf1073760e4bdac701bea6d0c1138a9adb9b99fb155903457615d3b

Observation ae624301-8431-4ccc-9a5a-72ae93f62973 · outbound

This paper cites On the importance of hyperparameter optimization for model-based reinforcement learning.

ViSTa Dataset: Do vision-language models understand sequential tasks? On the importance of hyperparameter optimization for model-based reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.120939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.191282Z digest=sha256:6793ac13fef5eb698b615c1b2dc077b2c75aac6f5650dce05072e0e988004a7f

Observation ee53c557-2a5a-4e63-a00a-7aada7a3959c · outbound

This paper cites Variational inverse control with events: A general framework for data-driven reward definition.

ViSTa Dataset: Do vision-language models understand sequential tasks? Variational inverse control with events: A general framework for data-driven reward definition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.101686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.198072Z digest=sha256:023c96a7c1dbd6eb36655bdf88f5bb8413cbc0c7f387d40f62e56f6778152bef

Observation b13ec86d-3df7-422f-bb34-42b565521e38 · outbound

This paper cites Deep reinforcement learning from human preferences.

ViSTa Dataset: Do vision-language models understand sequential tasks? Deep reinforcement learning from human preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.204429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.204429Z digest=sha256:10d3e8bdf2959a53fbb0ab29335ef4db0b80f66dbe0223cdd54266d55f07e586

Observation 1c616136-d84a-46b6-ad8c-e02bdf1af33c · outbound

This paper cites Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning.

ViSTa Dataset: Do vision-language models understand sequential tasks? Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.211087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.211087Z digest=sha256:65cf84f9efe37d8efc91181822f2f2a504f5b31d43137d9295d55a69555c3755

Observation ae1fa5bf-97f7-4b05-8bdb-1c3086adaa46 · outbound

This paper cites RoboCLIP: One Demonstration is Enough to Learn Robot Policies.

ViSTa Dataset: Do vision-language models understand sequential tasks? RoboCLIP: One Demonstration is Enough to Learn Robot Policies

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.217346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.217346Z digest=sha256:64070f0dad3659621f07d3b24583c52d3afb1fd3a696ab66b828c54a98764f70

Observation aa9aa0c6-3229-493d-a173-15357e8c13cb · outbound

This paper cites Task Success is not Enough: Investigating the Use of Video-Language Models as Behavior Critics for Catching Undesirable Agent Behaviors.

ViSTa Dataset: Do vision-language models understand sequential tasks? Task Success is not Enough: Investigating the Use of Video-Language Models as Behavior Critics for Catching Undesirable Agent Behaviors

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.222909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.222909Z digest=sha256:38c4d61adea121cf3ea8a1ab7bbfc81885124fa006a075226f9732bc1a5e887d

Observation a0169577-c9f9-4b4f-aace-dcf08d21b2bf · outbound

This paper cites Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos.

ViSTa Dataset: Do vision-language models understand sequential tasks? Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.228810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.228810Z digest=sha256:d1da99b5d1bdd1340420db225483342654fdeedb5552e13e671024342593830f

Observation 6718bc60-6efa-4b08-99ee-8934da57ee5c · outbound

This paper cites The unsur- prising effectiveness of pre-trained vision models for control.

ViSTa Dataset: Do vision-language models understand sequential tasks? The unsur- prising effectiveness of pre-trained vision models for control

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.070961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.233963Z digest=sha256:db8c209310064cbc659c0d8fe43b6ef79d9727cad99676b4f31bab08ba0a98fd

Observation e8819983-5ecb-496f-bbce-4df5d85468f5 · outbound

This paper cites Reward Design with Language Models.

ViSTa Dataset: Do vision-language models understand sequential tasks? Reward Design with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.239421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.239421Z digest=sha256:676427cc0312fd66b7cd679f76a9d389a9b7d30fc313951a2d2d356c2d1f85e2

Observation ed1eaabf-97a8-4d83-ba2c-94c173cec3a2 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

ViSTa Dataset: Do vision-language models understand sequential tasks? RewardBench: Evaluating Reward Models for Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.244847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.244847Z digest=sha256:ea9573f402d02a35e30e5489b9cd65d5f3c9b0e85e68e826a7015d16cbdb7c08

Observation 2d89699d-a082-48fc-9303-2e05cfd4779b · outbound

This paper cites Reward learning from narrated demonstrations.

ViSTa Dataset: Do vision-language models understand sequential tasks? Reward learning from narrated demonstrations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.053066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.251372Z digest=sha256:40424729bd954b7b93a81330dea17ffcd56901d7e14e2d0a5f3cc8ace38d502d

Observation 710e80fd-f955-423f-8ebe-7018aaa971bc · outbound

This paper cites Zero-shot reward specification via grounded natural language.

ViSTa Dataset: Do vision-language models understand sequential tasks? Zero-shot reward specification via grounded natural language

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.035370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.256190Z digest=sha256:6903557272cc352134bbdb54cd27619a2ab66694e9f5c97a6b9c10b0347e1c4c

Observation cc57d9c5-0e98-4c8f-95e1-cf4e4564c3f0 · outbound

This paper cites RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback.

ViSTa Dataset: Do vision-language models understand sequential tasks? RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.260990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.260990Z digest=sha256:003111388115991478d7b44ec8113890b682076a108379b0fd6e4c2c983ffd1e

Observation e93e662b-5681-4cfe-b64a-ed702c29e75b · outbound

This paper cites PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain.

ViSTa Dataset: Do vision-language models understand sequential tasks? PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.265925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.265925Z digest=sha256:fe797d3d23910922db4746ca9e7e0eb31ea39c5503724f1be1e80e99430e1a50

Observation bd236a1c-d166-4044-b0fc-81fa4e9e125e · outbound

This paper cites Paxion: Patching Action Knowledge in Video-Language Foundation Models,.

ViSTa Dataset: Do vision-language models understand sequential tasks? Paxion: Patching Action Knowledge in Video-Language Foundation Models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.018311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.271210Z digest=sha256:9e445fd0b130a03eea1ad0a8f6a379d5a86a26bd86544eb21cc4202847254a3c

Observation 12d3fae1-9b50-49d4-b082-a0c025bed755 · outbound

This paper cites BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.

ViSTa Dataset: Do vision-language models understand sequential tasks? BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.281865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.281865Z digest=sha256:561b140e39257a7c5de8100de737b1edcfadb1a1adee97b98a6c6939134ceeed

Observation 5874afd9-ede9-4663-bd7a-ac3b4119de37 · outbound

This paper cites VirtualHome: Simulating Household Activities via Programs.

ViSTa Dataset: Do vision-language models understand sequential tasks? VirtualHome: Simulating Household Activities via Programs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.286943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.286943Z digest=sha256:c25a8470192054629b18b9058bb63f3e3530a06ea3dcabf648fc2d98be2016cc

Observation 71db451c-40d2-45f9-8b25-7e66452edb7b · outbound

This paper cites Habitat: A Platform for Embodied AI Research, 2019.

ViSTa Dataset: Do vision-language models understand sequential tasks? Habitat: A Platform for Embodied AI Research, 2019

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.002322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.292786Z digest=sha256:76b6d6c476ab533d8e3effe3d6e3dbc9cf6ac41856dba0009b212f73d8eee985

Observation bb1bd04c-aecc-4cc9-89e0-8037a97f1019 · outbound

This paper cites Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots.

ViSTa Dataset: Do vision-language models understand sequential tasks? Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.297556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.297556Z digest=sha256:f0bb145679c3bc08f4f7d616ed5dcbf25fd234f4bb2762976c0baa4e467fa558

Observation 2279356c-1ac0-415b-9d63-798035e8fa5d · outbound

This paper cites ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks.

ViSTa Dataset: Do vision-language models understand sequential tasks? ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.303205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.303205Z digest=sha256:71b4fcb2cdd3872a3131baa82895d1da2cdbdebd816a91dfdd5cf7c7da3dc842

Observation f0770c95-1ffe-4f66-bf58-fe6bfe25b17b · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

ViSTa Dataset: Do vision-language models understand sequential tasks? A Short Note on the Kinetics-700 Human Action Dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.308527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.308527Z digest=sha256:26ecbf45879de227f33d036c559bc70adef9b0368cca47a3187e08ea5677a09d

Observation c5d0ccb6-c9d5-4759-9588-a02d4e5e6e0b · outbound

This paper cites BEDD: the minerl BASALT evaluation and demonstrations dataset for training and benchmarking agents that solve fuzzy tasks.

ViSTa Dataset: Do vision-language models understand sequential tasks? BEDD: the minerl BASALT evaluation and demonstrations dataset for training and benchmarking agents that solve fuzzy tasks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.986056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.314036Z digest=sha256:563dfa422353a319b828e5c8f776001d9befb977d3f2d9c1806aaa73ac3bd81e

Observation 7826661d-af97-436e-821c-bab9a859cffa · outbound

This paper cites Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks.

ViSTa Dataset: Do vision-language models understand sequential tasks? Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.319151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.319151Z digest=sha256:1321586a02ad47955266a870a41d805e96a50966713879c2b4503dadcc244790

Observation d2683055-6a7f-4111-9dd9-1f7db5471581 · outbound

This paper cites Learning transferable visual models from natural language supervision.

ViSTa Dataset: Do vision-language models understand sequential tasks? Learning transferable visual models from natural language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.324717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.324717Z digest=sha256:ccd52f1158c409cfe6f97cf2310875b1aee6bf704694c02bee0e8f890d41b1ee

Observation e028cb0c-5dc6-43e0-a6a0-b2a019949f23 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

ViSTa Dataset: Do vision-language models understand sequential tasks? Reproducible scaling laws for contrastive language-image learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.329565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.329565Z digest=sha256:7ae98c829fe0eada3de857a4174da14d8c6d5c3f880ddb438c967de90be55b5d

Observation 7cb16373-0c7a-453d-b7b7-d9eb3609c871 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

ViSTa Dataset: Do vision-language models understand sequential tasks? LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.334727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.334727Z digest=sha256:879614efa9ec5428b1b9fed62534acd5cf08c372941efeac950cabf05ee2b201

Observation b1c1dc7f-dd84-4c26-a1f5-56965d91caf5 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

ViSTa Dataset: Do vision-language models understand sequential tasks? InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.339786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.339786Z digest=sha256:910e538c456fef05f6168413d17a5d446d433a98e2d19db939bbc25f33b1144e

Observation a5599288-1fb0-4b3a-a851-82adc3c08e5a · outbound

This paper cites GPT-4o System Card.

ViSTa Dataset: Do vision-language models understand sequential tasks? GPT-4o System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.344793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.344793Z digest=sha256:b772ad417d3cf792571f1c19f124e6a97b975b04ab2fda2caeadf798e5a48c33

Observation b56b9f6a-7f3e-411b-b0e0-2f5d1da7985b · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.959247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.350936Z digest=sha256:ff3108fd6b56540b8c99fb281b68ff69de4542f693ae926f9ac8025d468bcb5e

Observation 0d463477-81d5-4b03-80d1-0e4d632067d4 · outbound

This paper cites likely does not describe the video.

ViSTa Dataset: Do vision-language models understand sequential tasks? likely does not describe the video

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.941465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.356076Z digest=sha256:ef358d3a00e1b62df937eb26eb13b74cf9e6f37eeb5398db6ea6f3eaa9e474e6

Observation 30a37c82-3666-4da5-88af-79c089e76955 · outbound

This paper cites Player holds an oak fence block in their hands.

ViSTa Dataset: Do vision-language models understand sequential tasks? Player holds an oak fence block in their hands

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.924572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.361398Z digest=sha256:f621579d1f312a8d94d8f681ef1cc9bb49d6766948b88e8bf634007dfdf4a080

Observation 70e1214b-42bd-4e42-9f27-ab6895ef9a60 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.907865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.366264Z digest=sha256:35a9c661e28aa7638165161d5e5d702b49eb206c11bbdb40e6a3acf0a4439654

Observation df5360ce-2569-4aa7-b5ad-0b5ad313d235 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.891764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.371881Z digest=sha256:99825ba994cc70095fd0b6345b7d56d2351f38df41cd25317969458588fbac43

Observation a87f0140-0eb4-4f33-af8e-bb5b875d2568 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.876042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.377065Z digest=sha256:ae9b7e3f4f585b9df793f2cfd1ffd4c8af1788b2ffd9540bdf853335f7f131e3

Observation 57397734-49ee-4851-95fa-7ba559af7cbb · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.858713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.382130Z digest=sha256:04e0711506f4fee267dfb60f48c9c247f2eb5d8da6eec0c377c69faa892a5509

Observation 1514cef4-3eb7-405d-aa8d-230ac9932723 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.840209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.387037Z digest=sha256:36211e0650ba66bce2ce2c667eaa7f82087c1249f8aeabbf5c36a6e65e2795e0

Observation 17260624-6b44-4002-beb6-dc5def08f35c · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.823881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.392457Z digest=sha256:59ad469c636880fd872ff6ffaa228d40bf6566cd9a98265718f8fc2d3b7d1279

Observation 6b87587f-48bf-46ac-b72a-5ea90ab7105e · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.807962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.397470Z digest=sha256:e6bbf6c5905f26f2618951b756526acbe98c9800699967acc81d2acf5f5e5b47

Observation 23868eb5-550a-4cdc-aaaf-bc4eb6984541 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.792703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.402130Z digest=sha256:a863fb25001dc17b0d38014869fb595c3327e641d8202b702d99f09e1810d9b1

Observation 89dd84ec-b538-41e0-aba9-af172bd47aad · outbound

This paper cites Figure 10: The prompt used to obtain frame-by-frame descriptions for Minecraft videos.

ViSTa Dataset: Do vision-language models understand sequential tasks? Figure 10: The prompt used to obtain frame-by-frame descriptions for Minecraft videos

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.776841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:47.406893Z digest=sha256:0a4fe9aa14fdfcf58754f779d5fd592c0577173921d80a6705d5d1dd463bfd1a

Observation 70ef0f7b-7277-46a0-89bb-879759f4bd46 · outbound

This paper cites Paxion: Patching Action Knowledge in Video-Language Foundation Models.

ViSTa Dataset: Do vision-language models understand sequential tasks? Paxion: Patching Action Knowledge in Video-Language Foundation Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.276058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.276058Z digest=sha256:ab389fb1bb261d55905709e2874faa1d806656e17e9aaa9b914d27feefe74e67

Pith citing papers

No inbound Pith citation observations are available.