Pith. sign in

Paper Citation Record · LEDGER

Action- and Language-Conditioned Video Assessment for Embodied Control

As of 19 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2608.08273.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08273 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:14:53.207582Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8661a9ac-a07e-432c-80f7-50b797bc7d98 · outbound

This paper cites Understanding natural language commands for robotic navigation and mobile manipulation.

Action- and Language-Conditioned Video Assessment for Embodied Control Understanding natural language commands for robotic navigation and mobile manipulation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:54.033647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:52.964269Z digest=sha256:bc3d12329e4579f79591907996808869aca39d23c09f742fed7b861db216c6f9

Observation d39cc3b8-9d17-40a9-9453-2abd65512dba · outbound

This paper cites Vision-and- language navigation: Interpreting visually-grounded navigation instructions in real environments.

Action- and Language-Conditioned Video Assessment for Embodied Control Vision-and- language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:54.018019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:52.969980Z digest=sha256:5cc2f9b3f5e70a5bbdafa71c20da3a17d8ac32cd00e3b8ac26c8751f09360320

Observation 78090cf3-5262-43c5-a398-12fe77347a4c · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

Action- and Language-Conditioned Video Assessment for Embodied Control Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:54.004373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:52.974957Z digest=sha256:85b116235382c088869cdbfac80b4b937cd7a550877e18daa2355a54f9c303e8

Observation 647a749b-fd4a-4d8c-9a73-c9f746998b3a · outbound

This paper cites Learning transferable visual models from natural language supervision.

Action- and Language-Conditioned Video Assessment for Embodied Control Learning transferable visual models from natural language supervision

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.990044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:52.980682Z digest=sha256:47ed4498559423b2fef1abdebe4baf0962b4c9e252e53e1d203e6ebacfc95827

Observation b55f337b-9e5b-4307-aceb-65f4ba04ded5 · outbound

This paper cites Zero-shot reward specification via grounded natural language.

Action- and Language-Conditioned Video Assessment for Embodied Control Zero-shot reward specification via grounded natural language

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.976167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:52.985936Z digest=sha256:9a4e5b1cbeccda55adf294a8886f607fb49e3c4b22223a219e7df2650f671a51

Observation f9155e15-4f48-49bb-96ce-76367a7a3dfe · outbound

This paper cites Vision-language models are zero-shot reward models for reinforcement learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Vision-language models are zero-shot reward models for reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.961646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:52.991180Z digest=sha256:486ae3c25d0e6cf20b79a1cc3b0eafcfa49d20ef6099cd42ae5a64d6e758ac14

Observation a8053856-7d00-4b58-9108-ab430e8a554c · outbound

This paper cites R3m: A universal visual representation for robot manipulation.

Action- and Language-Conditioned Video Assessment for Embodied Control R3m: A universal visual representation for robot manipulation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.947687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:52.996765Z digest=sha256:734380ea2b6956070dcb19765ad2014ad19d63813da88fbe13a6d71492b72cf9

Observation 07acc79e-15e9-4b1c-95e9-a1fc83444920 · outbound

This paper cites Liv: Language-image representations and rewards for robotic control.

Action- and Language-Conditioned Video Assessment for Embodied Control Liv: Language-image representations and rewards for robotic control

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.933060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.001334Z digest=sha256:5645e9e36b968d3296935a11251fd76cddeb8c9b4ff26e3f7f0154281b549d5d

Observation 2033ec3d-347e-4ae9-87ba-9d65181455cd · outbound

This paper cites Roboclip: One demonstration is enough to learn robot policies.

Action- and Language-Conditioned Video Assessment for Embodied Control Roboclip: One demonstration is enough to learn robot policies

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.918272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.005797Z digest=sha256:ab8d4bf963aac7b4d92d44b0e75caed14edce89ac33111cac052a99772ad0f77

Observation 2990cf95-b285-4e96-a8a9-fbf19b8bb106 · outbound

This paper cites GPT-4V(ision) System Card.

Action- and Language-Conditioned Video Assessment for Embodied Control GPT-4V(ision) System Card

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.903878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.010476Z digest=sha256:84717f9ffb4a387c214d9f7b14865412f778fc13ebfbc348503a27d9dd6d47da

Observation 9c314c26-8c9a-46f4-a9f0-09e522bf5538 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Action- and Language-Conditioned Video Assessment for Embodied Control Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:14:53.014866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:14:53.014866Z digest=sha256:499f14f51f0ee26f6e5e7ffdb7a0b530ecefbffe8a599c89d5a5a9fee887666e

Observation 3047651b-19a2-4a6d-a31a-50c0f4d60cc7 · outbound

This paper cites Hello GPT-4o.

Action- and Language-Conditioned Video Assessment for Embodied Control Hello GPT-4o

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.889371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.020003Z digest=sha256:60f63bbd8b543ceed537ad38c50026756847e98a4ca7a25a396e03d7585be108

Observation 670f002b-94a5-4577-b3db-bdee28bad407 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Action- and Language-Conditioned Video Assessment for Embodied Control AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:14:53.024968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:14:53.024968Z digest=sha256:37edd8f23f37efdde8b8387f244edd8d732fbfa62bf3cc8704c4065cf63efd7f

Observation 2fb2de0c-0c4f-4006-8ce3-52c4b86c4a62 · outbound

This paper cites Vision-language models for vision tasks: A survey.IEEE Trans.

Action- and Language-Conditioned Video Assessment for Embodied Control Vision-language models for vision tasks: A survey.IEEE Trans

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.874826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.030102Z digest=sha256:d36778fbaa5e4c20ac1fab8b48f61c96a53aae3ed1979a84639f66d34b842747

Observation 2608f3ab-abab-4fcd-85fd-0ad77c527455 · outbound

This paper cites Flamingo: A visual language model for few-shot learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Flamingo: A visual language model for few-shot learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.860356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.034808Z digest=sha256:b5a8e272b2ac48a35b823aff51bcf8ebad30892bb55a9c6a6cfdb4c10f3a9296

Observation 057600a0-43cb-4414-af5b-ce7aaa40080e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Action- and Language-Conditioned Video Assessment for Embodied Control Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:14:53.039341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:14:53.039341Z digest=sha256:ddfa5d7a7fec5b5a7b1615a80cbb95023c5af919b0634160d2122b4a8865c1d4

Observation b3e57602-3110-4556-87d9-25f35a935ef5 · outbound

This paper cites Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification.

Action- and Language-Conditioned Video Assessment for Embodied Control Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.846250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.044011Z digest=sha256:c214a3733bf7af2891a18cc674015e2c93d6a46b6a7bac4a3f3dcbb175463e7f

Observation c04b2b11-fe66-4ca4-b477-435613260d6c · outbound

This paper cites A survey of reinforcement learning informed by natural language.

Action- and Language-Conditioned Video Assessment for Embodied Control A survey of reinforcement learning informed by natural language

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.832557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.048688Z digest=sha256:4be004f8860237df6c7e31fbbae4947b5bc435f29cd2d34b1d2ac60e04136366

Observation f16293f2-7118-4254-9828-d4e78ad76809 · outbound

This paper cites Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation.

Action- and Language-Conditioned Video Assessment for Embodied Control Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.818648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.052911Z digest=sha256:41d1e5a0b2693208dbef04ba6242dd5ff4de5679130c7e7d21b6aa64301dc8e6

Observation 70cd1617-4951-4e84-bc49-a1b9f644e150 · outbound

This paper cites Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robot.

Action- and Language-Conditioned Video Assessment for Embodied Control Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robot

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.803771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.057004Z digest=sha256:5912fe77d65cb77f50be01ada1d42c6418e0de4277a49ab6ce22ccb6ac013a9f

Observation 82f919d6-c83c-4f55-8c1f-490264fa23ef · outbound

This paper cites Walk the talk: Connecting language, knowledge, and action in route instructions.

Action- and Language-Conditioned Video Assessment for Embodied Control Walk the talk: Connecting language, knowledge, and action in route instructions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.789435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.061250Z digest=sha256:50181d391257d155f2045436136814b34fb087d54565d7ce7b60bfbb9e8e41ef

Observation 1bf8c1bc-07e8-40ca-9415-89a8c629f9c1 · outbound

This paper cites Toward understanding natural language directions.

Action- and Language-Conditioned Video Assessment for Embodied Control Toward understanding natural language directions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.774718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.065238Z digest=sha256:0f30af2276fc0edb03913ddb0c9b53669f700ca33a5d82fa4c4657e0d4a34b22

Observation 2f8ffcc9-6ca5-4545-b5ab-b526ed1cb11b · outbound

This paper cites Tell Me Dave: Context-Sensitive Grounding of Natural Language to Manipulation Instructions.

Action- and Language-Conditioned Video Assessment for Embodied Control Tell Me Dave: Context-Sensitive Grounding of Natural Language to Manipulation Instructions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.758600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.069565Z digest=sha256:645cdc2ba204eebd5e55f9612ea7fcf5e9aba6849c3e35d4894bdaa351e3b214

Observation 1c3f86e7-2fc7-4c8a-8989-8f3e92e32445 · outbound

This paper cites Grounding English Commands to Reward Functions.

Action- and Language-Conditioned Video Assessment for Embodied Control Grounding English Commands to Reward Functions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.742233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.073985Z digest=sha256:e9c560e109f4582723a5e99306866b762495b9fbcb3252e7e4cb1266a8ce907e

Observation a6b9ce72-d0fd-4e7b-83fa-7fc28cf49b81 · outbound

This paper cites Learning language-conditioned robot behavior from offline data and crowd-sourced annotation.

Action- and Language-Conditioned Video Assessment for Embodied Control Learning language-conditioned robot behavior from offline data and crowd-sourced annotation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.725996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.078422Z digest=sha256:9435e7852a4e75d464af76c860c32281b2b4c27f115f49e82e9e57de3c13541a

Observation 9a76092d-810a-479b-ba14-b89331bae3fe · outbound

This paper cites Algorithms for inverse reinforcement learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Algorithms for inverse reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.710498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.082907Z digest=sha256:d8907eac323cc90e39cfc7b2fcd54cfe0efab4fff2abfaf8126c3a260ec9c525

Observation d5f7fd72-74aa-4471-9d50-71d25c3d7044 · outbound

This paper cites Apprenticeship learning via inverse reinforcement learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Apprenticeship learning via inverse reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.695498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.088367Z digest=sha256:96aeb4d97ab34a1c950169552d40b1f2c919fe050c0321dd70431ba6dfeb32b0

Observation 92992ae9-aa61-4fef-b507-2a3dfa458998 · outbound

This paper cites Generative adversarial imitation learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Generative adversarial imitation learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.680865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.092958Z digest=sha256:20a26b6c804539165dcf169c1b40c51a3488ed0ae09d3011de5a828bb9b40840

Observation a8447baa-7fb1-4f17-a60d-827673d5ac3e · outbound

This paper cites From language to goals: Inverse reinforcement learning for vision-based instruction following.

Action- and Language-Conditioned Video Assessment for Embodied Control From language to goals: Inverse reinforcement learning for vision-based instruction following

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.667083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.098089Z digest=sha256:fbfd51d3ce81c5e71b505c2ab59a27e3a349f84537a28fee4fb93bca296e75d9

Observation 1626dc02-bb2f-4774-a44e-c3750f02122c · outbound

This paper cites Learning to Understand Goal Specifications by Modelling Reward.

Action- and Language-Conditioned Video Assessment for Embodied Control Learning to Understand Goal Specifications by Modelling Reward

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.651779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.102920Z digest=sha256:7c793f33071b6eb1d9f6e4b5dbff1cc8f568182e04663e696287916804457802

Observation 2c859d64-d627-4053-86b0-65671139af75 · outbound

This paper cites Reward Design with Language Models.

Action- and Language-Conditioned Video Assessment for Embodied Control Reward Design with Language Models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.637245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.107505Z digest=sha256:5e82053efa5ed4841c33a1b574d6533b8eb90d54bd89f024a973b09ce0fbf322

Observation be7ad28a-7900-4234-a019-b335c08756fe · outbound

This paper cites Language to Rewards for Robotic Skill Synthesis.

Action- and Language-Conditioned Video Assessment for Embodied Control Language to Rewards for Robotic Skill Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:14:53.112288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:14:53.112288Z digest=sha256:a72721ad4bf74e0ef038bacefe304f25d4fc9f8b45f0682a1afc0f3e00ef0d94

Observation ef575489-4022-42da-9e92-42acbf469b03 · outbound

This paper cites Text2reward: Automated dense reward function generation for reinforcement learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Text2reward: Automated dense reward function generation for reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.622640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.118463Z digest=sha256:0a84ff44cefb149fe265d6ddfc295ca945f11f531537cd6036eddeb626305a73

Observation 2812868e-82b7-4310-8c34-87a4bf4ed523 · outbound

This paper cites Eureka: Human-level reward design via coding large language models.

Action- and Language-Conditioned Video Assessment for Embodied Control Eureka: Human-level reward design via coding large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.607822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.123096Z digest=sha256:2af5952f79b0dc4e1f64a7e5c2486f587068e9d1bc644c8e2bbf100439ccf0c0

Observation 9152c17e-adbe-4bde-9c6f-75188065e6cb · outbound

This paper cites Robogen: Towards unleashing infinite data for automated robot learning via generative simulation.

Action- and Language-Conditioned Video Assessment for Embodied Control Robogen: Towards unleashing infinite data for automated robot learning via generative simulation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.593269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.127834Z digest=sha256:4659f905ded4589fab284d527604fd6ddc7a4b3eebfc297995cb7319ef9aeec7

Observation f4e19eba-c9bb-4a31-8415-3442f2618c5b · outbound

This paper cites an unresolved cited work.

Action- and Language-Conditioned Video Assessment for Embodied Control Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:14:53.578351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.132841Z digest=sha256:8b144a4bced4c9bd30bf53b83b5c109c121b780bd1b053d6b1c08fce50b205ef

Observation c311f943-f30d-48f8-a359-d46742391e30 · outbound

This paper cites Vision-language models as success detectors.

Action- and Language-Conditioned Video Assessment for Embodied Control Vision-language models as success detectors

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.562755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.137917Z digest=sha256:2508b7b54d119e523eb662dd2b5196bdc3bd349f41179eb3b86a86ea56da6e62

Observation 79ed6e64-f6f7-4650-a39d-e52cb36eda0d · outbound

This paper cites Rl-vlm-f: Reinforcement learning from vision language foundation model feedback.

Action- and Language-Conditioned Video Assessment for Embodied Control Rl-vlm-f: Reinforcement learning from vision language foundation model feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.547176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.142723Z digest=sha256:3074b2006037f83354468d99db632b29c5876448bf1c8ce00bad8d283efee2bd

Observation ea720739-e555-41e2-835a-30e40e9dc585 · outbound

This paper cites Deep reinforcement learning from human preferences.

Action- and Language-Conditioned Video Assessment for Embodied Control Deep reinforcement learning from human preferences

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.531829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.147190Z digest=sha256:5d1f6754bea36d92b83e2f0a26d2f9a2e7728574a4abec96cd7fd0e45c947ee1

Observation e6b72b7a-31fd-4d77-a1aa-174466f1f682 · outbound

This paper cites Interactive learning from policy-dependent human feedback.

Action- and Language-Conditioned Video Assessment for Embodied Control Interactive learning from policy-dependent human feedback

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.516014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.151536Z digest=sha256:fff74e49249db784ce7c7eea8b390e002ba05bb7abbf7456ea77b137ebdc4d96

Observation 5d9ded7e-8af5-4068-a227-11530886fa47 · outbound

This paper cites Learning reward functions from scale feedback.

Action- and Language-Conditioned Video Assessment for Embodied Control Learning reward functions from scale feedback

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.501676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.156022Z digest=sha256:256aca947b059f7084da3c775dcd49b067ef3bffac5d7d983b18b97ad5948b6e

Observation d6e4fbe1-6ec7-465a-aa33-5aad1c0f33c3 · outbound

This paper cites Rating-based reinforcement learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Rating-based reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.486183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.160639Z digest=sha256:b9097f4cf118379ceb3f3ae9ad5af0e0c2aa6c60dc4347ee383aa24bd4fc1d28

Observation f08808fd-23b0-441a-8b15-7802c64250ec · outbound

This paper cites NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions.

Action- and Language-Conditioned Video Assessment for Embodied Control NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.470836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.165188Z digest=sha256:efa3336fe489ed40680de499a49102f5dd7b054fac5a6a51991ed832570cc9b7

Observation 7f87dfa0-47ed-4eaa-a44d-269017156998 · outbound

This paper cites From Representation to Reasoning: Towards Both Evidence and Commonsense Reasoning for Video Question-Answering.

Action- and Language-Conditioned Video Assessment for Embodied Control From Representation to Reasoning: Towards Both Evidence and Commonsense Reasoning for Video Question-Answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.455584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.169788Z digest=sha256:0fb6ad3f345798752a0986d0d2f0fb2a77b2babfb751970ff5c8364a45bd689b

Observation 5a586739-b824-4f3b-951c-a71c053bd80a · outbound

This paper cites Discovering the Real Association: Multimodal Causal Reasoning in Video Question Answering.

Action- and Language-Conditioned Video Assessment for Embodied Control Discovering the Real Association: Multimodal Causal Reasoning in Video Question Answering

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.439686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.174719Z digest=sha256:3da6da40c0e98d9404098f42e3d7cccd9dad291a8b80dffa5f96e7cc927bc056

Observation 3ea084ff-a6a9-42da-9e30-7c2d31fafa24 · outbound

This paper cites MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning.

Action- and Language-Conditioned Video Assessment for Embodied Control MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.422990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.179150Z digest=sha256:abe2057abe729ef97d83b56a3505732a4e74ff6e892beba44d68c1b58420af3d

Observation 62820c07-afc0-4168-ba52-1cadd8b5d15b · outbound

This paper cites Open problems and fundamental limitations of reinforcement learning from human feedback.Trans.

Action- and Language-Conditioned Video Assessment for Embodied Control Open problems and fundamental limitations of reinforcement learning from human feedback.Trans

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.408760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.183259Z digest=sha256:41bcad7e92aa49243321ce6932c75cfe3f0df60504c14fd6e236c88e4e07af8f

Observation 7910f685-215b-4657-9ba9-34acda89ef37 · outbound

This paper cites Visual prompt tuning.

Action- and Language-Conditioned Video Assessment for Embodied Control Visual prompt tuning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.393842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.187182Z digest=sha256:07ded0ef266bc7b93c58a5a29e40caa60a8a00aabd9ead8bdbe8e02d7e8f1f5d

Observation 24a82449-7086-4698-8b78-46dc96409596 · outbound

This paper cites Visual prompting via image inpainting.

Action- and Language-Conditioned Video Assessment for Embodied Control Visual prompting via image inpainting

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.377631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.191045Z digest=sha256:3d6ba2343ee71408b1ae467cf39c8fbf3b4825c92652b9653de762360e59536a

Observation 1bf75583-801e-4e89-aaeb-34e63fb22ce2 · outbound

This paper cites What does clip know about a red circle? visual prompt engineering for vlms.

Action- and Language-Conditioned Video Assessment for Embodied Control What does clip know about a red circle? visual prompt engineering for vlms

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.361646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.195003Z digest=sha256:076898fdbfb281bd4325a9e3e4e54dbea1d04829cf773df3c94c436793d69faa

Observation 286f98ee-f24b-4863-ad9f-f26b17dfc6aa · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Offline reinforcement learning with implicit q-learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.344912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.199150Z digest=sha256:578781417daf523c2349adba46df7a1c61bd604f98033b14ee05d99f56bd0516

Observation 3e54e165-b3e7-47c0-8e8c-c497aadef2c5 · outbound

This paper cites Bootstrap your own skills: Learning to solve new tasks with large language model guidance.

Action- and Language-Conditioned Video Assessment for Embodied Control Bootstrap your own skills: Learning to solve new tasks with large language model guidance

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.328819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.203175Z digest=sha256:a06649bdcd6c5da26730f4e333b937d99027cf10912811d59f942b453c977d45

Observation cd29952b-f632-4979-af5d-f23506e1c664 · outbound

This paper cites Deep residual learning for image recognition.

Action- and Language-Conditioned Video Assessment for Embodied Control Deep residual learning for image recognition

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.311803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T00:14:53.207582Z digest=sha256:444daec14f244c7b27c4ce814fabe606a34b24cdf4c691827ae7885bcf4c9d2f

Pith citing papers

No inbound Pith citation observations are available.