Pith. sign in

Paper Citation Record · LEDGER

ViSTa Dataset: Do vision-language models understand sequential tasks?

As of 13 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2411.13211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13211 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:45:47.406893Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcc889a0-0f25-4e85-9f8f-3a7b480f2842 · outbound

This paper cites Mas- tering the game of go with deep neural networks and tree search.

ViSTa Dataset: Do vision-language models understand sequential tasks? Mas- tering the game of go with deep neural networks and tree search

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.161355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.161355Z digest=sha256:29d0e0244c1ace0f2ec7c90f12364adf9ff41c9af840127eb7052438ce71d047

Observation 534fdd08-0d02-4d4d-bc0a-d5ea9d2d8de7 · outbound

This paper cites Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes.

ViSTa Dataset: Do vision-language models understand sequential tasks? Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.168907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.168907Z digest=sha256:1e03efeb25e49b8efe170a9503bf4de5ee441c128eb47e17ba7da1cea05b5b42

Observation e7042915-8ccb-4735-8104-11b6b7592550 · outbound

This paper cites Specification gaming: the flip side of ai ingenuity.

ViSTa Dataset: Do vision-language models understand sequential tasks? Specification gaming: the flip side of ai ingenuity

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.155246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.176949Z digest=sha256:db649cde164f8c78d98cacd7a98d0f41795c6ea70ca38a4792f1c41fe8ae7550

Observation e93a6f76-1f30-4f21-868b-f11ca058f76c · outbound

This paper cites Defining and characterizing reward gaming.

ViSTa Dataset: Do vision-language models understand sequential tasks? Defining and characterizing reward gaming

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.138168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.184902Z digest=sha256:a1bb0debd66df38ca225555f5bb8bf11c5ff58fccf72ab1bbfe7fe95fea9be68

Observation ae624301-8431-4ccc-9a5a-72ae93f62973 · outbound

This paper cites On the importance of hyperparameter optimization for model-based reinforcement learning.

ViSTa Dataset: Do vision-language models understand sequential tasks? On the importance of hyperparameter optimization for model-based reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.120939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.191282Z digest=sha256:898b3c29126d6a41d24bf26a2eac7df27c186ff6152946c18e2761d8ac089f75

Observation ee53c557-2a5a-4e63-a00a-7aada7a3959c · outbound

This paper cites Variational inverse control with events: A general framework for data-driven reward definition.

ViSTa Dataset: Do vision-language models understand sequential tasks? Variational inverse control with events: A general framework for data-driven reward definition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.101686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.198072Z digest=sha256:57671d04358d2996d6ac2d10945a5a8f3002817a9d2cb081fc30529f19e9be09

Observation b13ec86d-3df7-422f-bb34-42b565521e38 · outbound

This paper cites Deep reinforcement learning from human preferences.

ViSTa Dataset: Do vision-language models understand sequential tasks? Deep reinforcement learning from human preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.204429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.204429Z digest=sha256:10d3e8bdf2959a53fbb0ab29335ef4db0b80f66dbe0223cdd54266d55f07e586

Observation 1c616136-d84a-46b6-ad8c-e02bdf1af33c · outbound

This paper cites Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning.

ViSTa Dataset: Do vision-language models understand sequential tasks? Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.211087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.211087Z digest=sha256:4b2fe7cdf6fe2734980ea2dc9f52ecec81b73e9e3c7d1dce4a26d83ea38a211a

Observation ae1fa5bf-97f7-4b05-8bdb-1c3086adaa46 · outbound

This paper cites RoboCLIP: One Demonstration is Enough to Learn Robot Policies.

ViSTa Dataset: Do vision-language models understand sequential tasks? RoboCLIP: One Demonstration is Enough to Learn Robot Policies

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.217346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.217346Z digest=sha256:61837a1cefd06b62299bc0f122dc13ab0c678b42739d3296ea00f0a437a6c6e7

Observation aa9aa0c6-3229-493d-a173-15357e8c13cb · outbound

This paper cites Task Success is not Enough: Investigating the Use of Video-Language Models as Behavior Critics for Catching Undesirable Agent Behaviors.

ViSTa Dataset: Do vision-language models understand sequential tasks? Task Success is not Enough: Investigating the Use of Video-Language Models as Behavior Critics for Catching Undesirable Agent Behaviors

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.222909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.222909Z digest=sha256:dbf86634fc5f67d4a50b971bb573b10e0649d54a5d654f8fd69080b61323caee

Observation a0169577-c9f9-4b4f-aace-dcf08d21b2bf · outbound

This paper cites Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos.

ViSTa Dataset: Do vision-language models understand sequential tasks? Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.228810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.228810Z digest=sha256:d1da99b5d1bdd1340420db225483342654fdeedb5552e13e671024342593830f

Observation 6718bc60-6efa-4b08-99ee-8934da57ee5c · outbound

This paper cites The unsur- prising effectiveness of pre-trained vision models for control.

ViSTa Dataset: Do vision-language models understand sequential tasks? The unsur- prising effectiveness of pre-trained vision models for control

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.070961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.233963Z digest=sha256:ea5d8019998f8a5632baa12eba14e08097f3d0e8a3f09c5b42518dd4773a779d

Observation e8819983-5ecb-496f-bbce-4df5d85468f5 · outbound

This paper cites Reward Design with Language Models.

ViSTa Dataset: Do vision-language models understand sequential tasks? Reward Design with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.239421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.239421Z digest=sha256:676427cc0312fd66b7cd679f76a9d389a9b7d30fc313951a2d2d356c2d1f85e2

Observation ed1eaabf-97a8-4d83-ba2c-94c173cec3a2 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

ViSTa Dataset: Do vision-language models understand sequential tasks? RewardBench: Evaluating Reward Models for Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.244847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.244847Z digest=sha256:ea9573f402d02a35e30e5489b9cd65d5f3c9b0e85e68e826a7015d16cbdb7c08

Observation 2d89699d-a082-48fc-9303-2e05cfd4779b · outbound

This paper cites Reward learning from narrated demonstrations.

ViSTa Dataset: Do vision-language models understand sequential tasks? Reward learning from narrated demonstrations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.053066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.251372Z digest=sha256:b1ad3397133d8f6196a2b2d9db70640734df6c2424c052f2fe541b004446f742

Observation 710e80fd-f955-423f-8ebe-7018aaa971bc · outbound

This paper cites Zero-shot reward specification via grounded natural language.

ViSTa Dataset: Do vision-language models understand sequential tasks? Zero-shot reward specification via grounded natural language

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.035370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.256190Z digest=sha256:f7d073db4a393c34061273e7e9894a98c197a65d80a310fe9d163d68d5c76de5

Observation cc57d9c5-0e98-4c8f-95e1-cf4e4564c3f0 · outbound

This paper cites RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback.

ViSTa Dataset: Do vision-language models understand sequential tasks? RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.260990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.260990Z digest=sha256:3cb17c2606720786c14e22830efe54d515fe19f7f1215cbd877f69817befc975

Observation e93e662b-5681-4cfe-b64a-ed702c29e75b · outbound

This paper cites PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain.

ViSTa Dataset: Do vision-language models understand sequential tasks? PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.265925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.265925Z digest=sha256:96888a05d59f9c3cee1bc0f645e77ebbc7d49cc28f3aa1f1f44939cfdf6838c8

Observation bd236a1c-d166-4044-b0fc-81fa4e9e125e · outbound

This paper cites Paxion: Patching Action Knowledge in Video-Language Foundation Models,.

ViSTa Dataset: Do vision-language models understand sequential tasks? Paxion: Patching Action Knowledge in Video-Language Foundation Models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.018311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.271210Z digest=sha256:3491f84b3b4f95073fd02c20f4841b4129dc453b3f876df25cf8ba12cb81173a

Observation 12d3fae1-9b50-49d4-b082-a0c025bed755 · outbound

This paper cites BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.

ViSTa Dataset: Do vision-language models understand sequential tasks? BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.281865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.281865Z digest=sha256:561b140e39257a7c5de8100de737b1edcfadb1a1adee97b98a6c6939134ceeed

Observation 5874afd9-ede9-4663-bd7a-ac3b4119de37 · outbound

This paper cites VirtualHome: Simulating Household Activities via Programs.

ViSTa Dataset: Do vision-language models understand sequential tasks? VirtualHome: Simulating Household Activities via Programs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.286943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.286943Z digest=sha256:c25a8470192054629b18b9058bb63f3e3530a06ea3dcabf648fc2d98be2016cc

Observation 71db451c-40d2-45f9-8b25-7e66452edb7b · outbound

This paper cites Habitat: A Platform for Embodied AI Research, 2019.

ViSTa Dataset: Do vision-language models understand sequential tasks? Habitat: A Platform for Embodied AI Research, 2019

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:48.002322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.292786Z digest=sha256:c3e376eaf90da4969633dcf95e7212eab435b40d52f5d0029e20489366ee37fa

Observation bb1bd04c-aecc-4cc9-89e0-8037a97f1019 · outbound

This paper cites Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots.

ViSTa Dataset: Do vision-language models understand sequential tasks? Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.297556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.297556Z digest=sha256:e751599cb51104e3df9dcff521c30a3fc73f2b2180bdc8925889a47ce22ddc06

Observation 2279356c-1ac0-415b-9d63-798035e8fa5d · outbound

This paper cites ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks.

ViSTa Dataset: Do vision-language models understand sequential tasks? ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.303205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.303205Z digest=sha256:71b4fcb2cdd3872a3131baa82895d1da2cdbdebd816a91dfdd5cf7c7da3dc842

Observation f0770c95-1ffe-4f66-bf58-fe6bfe25b17b · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

ViSTa Dataset: Do vision-language models understand sequential tasks? A Short Note on the Kinetics-700 Human Action Dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.308527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.308527Z digest=sha256:26ecbf45879de227f33d036c559bc70adef9b0368cca47a3187e08ea5677a09d

Observation c5d0ccb6-c9d5-4759-9588-a02d4e5e6e0b · outbound

This paper cites BEDD: the minerl BASALT evaluation and demonstrations dataset for training and benchmarking agents that solve fuzzy tasks.

ViSTa Dataset: Do vision-language models understand sequential tasks? BEDD: the minerl BASALT evaluation and demonstrations dataset for training and benchmarking agents that solve fuzzy tasks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.986056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.314036Z digest=sha256:f82e431b3f0d553d8bab2ad24334a02e52beb7a15f773849c0eca471870227fb

Observation 7826661d-af97-436e-821c-bab9a859cffa · outbound

This paper cites Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks.

ViSTa Dataset: Do vision-language models understand sequential tasks? Skill Reinforcement Learning and Planning for Open-World Long-Horizon Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.319151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.319151Z digest=sha256:1321586a02ad47955266a870a41d805e96a50966713879c2b4503dadcc244790

Observation d2683055-6a7f-4111-9dd9-1f7db5471581 · outbound

This paper cites Learning transferable visual models from natural language supervision.

ViSTa Dataset: Do vision-language models understand sequential tasks? Learning transferable visual models from natural language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.324717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.324717Z digest=sha256:ccd52f1158c409cfe6f97cf2310875b1aee6bf704694c02bee0e8f890d41b1ee

Observation e028cb0c-5dc6-43e0-a6a0-b2a019949f23 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

ViSTa Dataset: Do vision-language models understand sequential tasks? Reproducible scaling laws for contrastive language-image learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.329565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.329565Z digest=sha256:7ae98c829fe0eada3de857a4174da14d8c6d5c3f880ddb438c967de90be55b5d

Observation 7cb16373-0c7a-453d-b7b7-d9eb3609c871 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

ViSTa Dataset: Do vision-language models understand sequential tasks? LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.334727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.334727Z digest=sha256:879614efa9ec5428b1b9fed62534acd5cf08c372941efeac950cabf05ee2b201

Observation b1c1dc7f-dd84-4c26-a1f5-56965d91caf5 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

ViSTa Dataset: Do vision-language models understand sequential tasks? InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.339786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.339786Z digest=sha256:910e538c456fef05f6168413d17a5d446d433a98e2d19db939bbc25f33b1144e

Observation a5599288-1fb0-4b3a-a851-82adc3c08e5a · outbound

This paper cites GPT-4o System Card.

ViSTa Dataset: Do vision-language models understand sequential tasks? GPT-4o System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.344793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.344793Z digest=sha256:b772ad417d3cf792571f1c19f124e6a97b975b04ab2fda2caeadf798e5a48c33

Observation b56b9f6a-7f3e-411b-b0e0-2f5d1da7985b · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.959247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.350936Z digest=sha256:1f8d68357c95818ae5bd0a233807db0bb2d84a87fac1216e6e05a5d2084dd944

Observation 0d463477-81d5-4b03-80d1-0e4d632067d4 · outbound

This paper cites likely does not describe the video.

ViSTa Dataset: Do vision-language models understand sequential tasks? likely does not describe the video

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.941465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.356076Z digest=sha256:8d256f35a262f1ecfdb5fb95d4f7f92dace8ce71824ffae544c62e336b08bdc6

Observation 30a37c82-3666-4da5-88af-79c089e76955 · outbound

This paper cites Player holds an oak fence block in their hands.

ViSTa Dataset: Do vision-language models understand sequential tasks? Player holds an oak fence block in their hands

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.924572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.361398Z digest=sha256:36b34bcfbab53dfc35e9084f03a88d8eb1ab4eeec324e189303f052368c1ddaa

Observation 70e1214b-42bd-4e42-9f27-ab6895ef9a60 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.907865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.366264Z digest=sha256:48b2c79424301297abc5b978870f676e3c485f03679a031201bba5e19ab1fcdd

Observation df5360ce-2569-4aa7-b5ad-0b5ad313d235 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.891764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.371881Z digest=sha256:b55515047c7e43c0469c71c649c323251aeedbda8d56850c3f8873ce6b26a147

Observation a87f0140-0eb4-4f33-af8e-bb5b875d2568 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.876042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.377065Z digest=sha256:3939e753e9a7d6135907af5da8ae09f9b242668c9b5b8a2c90f07bd6654ff435

Observation 57397734-49ee-4851-95fa-7ba559af7cbb · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.858713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.382130Z digest=sha256:226fa966e46f75a5e9bc1193e7229c5055b4194ecae407209866fca4ff884607

Observation 1514cef4-3eb7-405d-aa8d-230ac9932723 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.840209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.387037Z digest=sha256:5d9e17c37a79311bf5a49802a061c93bc37758e241ff12a998afbde4c81b6595

Observation 17260624-6b44-4002-beb6-dc5def08f35c · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.823881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.392457Z digest=sha256:0cb8c3eb8dfcd03c9fc4b77ac11b63ba87d4188192fb754d5bc6dbe194ec6496

Observation 6b87587f-48bf-46ac-b72a-5ea90ab7105e · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.807962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.397470Z digest=sha256:987ddc21c6bce21358a5911c97f9e84aa9b6b991b3b7b929c2900fe38541e62c

Observation 23868eb5-550a-4cdc-aaaf-bc4eb6984541 · outbound

This paper cites an unresolved cited work.

ViSTa Dataset: Do vision-language models understand sequential tasks? Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:47.792703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.402130Z digest=sha256:c06dbf185b0dc208ad94e1a44fad498380f32552159d1c32f12c9cb8a09bbc1b

Observation 89dd84ec-b538-41e0-aba9-af172bd47aad · outbound

This paper cites Figure 10: The prompt used to obtain frame-by-frame descriptions for Minecraft videos.

ViSTa Dataset: Do vision-language models understand sequential tasks? Figure 10: The prompt used to obtain frame-by-frame descriptions for Minecraft videos

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:45:47.776841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:45:47.406893Z digest=sha256:5783934c0f987b29c64f98e4479904b311a6a09e01be4082e5b9747a766aaf96

Observation 70ef0f7b-7277-46a0-89bb-879759f4bd46 · outbound

This paper cites Paxion: Patching Action Knowledge in Video-Language Foundation Models.

ViSTa Dataset: Do vision-language models understand sequential tasks? Paxion: Patching Action Knowledge in Video-Language Foundation Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:47.276058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:47.276058Z digest=sha256:0abbd0d99d207d1d2f1b627d1cb842d209471612fd5d53b53ad15c5c0e5f57ee

Pith citing papers

No inbound Pith citation observations are available.