Pith. sign in

Paper Citation Record · LEDGER

Action- and Language-Conditioned Video Assessment for Embodied Control

As of 18 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2608.08273.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08273 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:14:53.207582Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8661a9ac-a07e-432c-80f7-50b797bc7d98 · outbound

This paper cites Understanding natural language commands for robotic navigation and mobile manipulation.

Action- and Language-Conditioned Video Assessment for Embodied Control Understanding natural language commands for robotic navigation and mobile manipulation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:54.033647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:52.964269Z digest=sha256:0bab1ff0b81b0112fdaf3f5daafd7c447426f7c294af7c49cc3ff0962f1accbf

Observation d39cc3b8-9d17-40a9-9453-2abd65512dba · outbound

This paper cites Vision-and- language navigation: Interpreting visually-grounded navigation instructions in real environments.

Action- and Language-Conditioned Video Assessment for Embodied Control Vision-and- language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:54.018019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:52.969980Z digest=sha256:a20e71fdd00d4349833d275213376f1a8abdb9146912fca5e723bff1c50d5963

Observation 78090cf3-5262-43c5-a398-12fe77347a4c · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

Action- and Language-Conditioned Video Assessment for Embodied Control Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:54.004373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:52.974957Z digest=sha256:424f4693d8a9c4d383068a46dac26993a103e1398a479b5ad2088200a5a0370a

Observation 647a749b-fd4a-4d8c-9a73-c9f746998b3a · outbound

This paper cites Learning transferable visual models from natural language supervision.

Action- and Language-Conditioned Video Assessment for Embodied Control Learning transferable visual models from natural language supervision

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.990044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:52.980682Z digest=sha256:6349d5d669e94c1d7df10a16aeded8e0fe0dd71dc6781b96f035c357d85992f6

Observation b55f337b-9e5b-4307-aceb-65f4ba04ded5 · outbound

This paper cites Zero-shot reward specification via grounded natural language.

Action- and Language-Conditioned Video Assessment for Embodied Control Zero-shot reward specification via grounded natural language

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.976167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:52.985936Z digest=sha256:1a7043d6b0ee30f8c69078321be2c79919cd36842845f27df9ae4d5646ce3ffc

Observation f9155e15-4f48-49bb-96ce-76367a7a3dfe · outbound

This paper cites Vision-language models are zero-shot reward models for reinforcement learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Vision-language models are zero-shot reward models for reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.961646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:52.991180Z digest=sha256:2d004ee605ce8807ea010b6067466b0c37a5d274598f06271c756e7a9d4e2949

Observation a8053856-7d00-4b58-9108-ab430e8a554c · outbound

This paper cites R3m: A universal visual representation for robot manipulation.

Action- and Language-Conditioned Video Assessment for Embodied Control R3m: A universal visual representation for robot manipulation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.947687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:52.996765Z digest=sha256:19d2eac5f5a9bda4ac95c03fc4af9f95dcf339892ea2ed0e751439f9b5ef2aba

Observation 07acc79e-15e9-4b1c-95e9-a1fc83444920 · outbound

This paper cites Liv: Language-image representations and rewards for robotic control.

Action- and Language-Conditioned Video Assessment for Embodied Control Liv: Language-image representations and rewards for robotic control

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.933060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.001334Z digest=sha256:dcecc88af8b4c94c93ae696c816f4b1375a0a1217642230bce1523a281ba554d

Observation 2033ec3d-347e-4ae9-87ba-9d65181455cd · outbound

This paper cites Roboclip: One demonstration is enough to learn robot policies.

Action- and Language-Conditioned Video Assessment for Embodied Control Roboclip: One demonstration is enough to learn robot policies

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.918272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.005797Z digest=sha256:9db3cb9b7a2fe4c05f04e78d9dc123ca3e7eb4ac72a6d6ebd8d97a07fa645a3a

Observation 2990cf95-b285-4e96-a8a9-fbf19b8bb106 · outbound

This paper cites GPT-4V(ision) System Card.

Action- and Language-Conditioned Video Assessment for Embodied Control GPT-4V(ision) System Card

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.903878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.010476Z digest=sha256:3748672701cf349d3a5f369a1ad18c30d4b30e9d1d5938d8c078d17a629b12c3

Observation 9c314c26-8c9a-46f4-a9f0-09e522bf5538 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Action- and Language-Conditioned Video Assessment for Embodied Control Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:14:53.014866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:14:53.014866Z digest=sha256:499f14f51f0ee26f6e5e7ffdb7a0b530ecefbffe8a599c89d5a5a9fee887666e

Observation 3047651b-19a2-4a6d-a31a-50c0f4d60cc7 · outbound

This paper cites Hello GPT-4o.

Action- and Language-Conditioned Video Assessment for Embodied Control Hello GPT-4o

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.889371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.020003Z digest=sha256:849df4fa35878bc58dbcc302e5c595cd68a5da02713d16457f08d310ffc2053f

Observation 670f002b-94a5-4577-b3db-bdee28bad407 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Action- and Language-Conditioned Video Assessment for Embodied Control AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:14:53.024968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:14:53.024968Z digest=sha256:37edd8f23f37efdde8b8387f244edd8d732fbfa62bf3cc8704c4065cf63efd7f

Observation 2fb2de0c-0c4f-4006-8ce3-52c4b86c4a62 · outbound

This paper cites Vision-language models for vision tasks: A survey.IEEE Trans.

Action- and Language-Conditioned Video Assessment for Embodied Control Vision-language models for vision tasks: A survey.IEEE Trans

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.874826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.030102Z digest=sha256:3aa8ecd2beb5a49fc3ece8e8466f28c3d150d4764537aab574463675e4a9b268

Observation 2608f3ab-abab-4fcd-85fd-0ad77c527455 · outbound

This paper cites Flamingo: A visual language model for few-shot learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Flamingo: A visual language model for few-shot learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.860356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.034808Z digest=sha256:7623fbbc578ce5cb2a7a7b402fc72dc5274eade5f339bd1765c9c480532b435e

Observation 057600a0-43cb-4414-af5b-ce7aaa40080e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Action- and Language-Conditioned Video Assessment for Embodied Control Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:14:53.039341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:14:53.039341Z digest=sha256:ddfa5d7a7fec5b5a7b1615a80cbb95023c5af919b0634160d2122b4a8865c1d4

Observation b3e57602-3110-4556-87d9-25f35a935ef5 · outbound

This paper cites Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification.

Action- and Language-Conditioned Video Assessment for Embodied Control Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.846250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.044011Z digest=sha256:30a34cfc3585ab4e8b307df0c616159ed3be320e293e9a0d998e8b5ff5d0b4be

Observation c04b2b11-fe66-4ca4-b477-435613260d6c · outbound

This paper cites A survey of reinforcement learning informed by natural language.

Action- and Language-Conditioned Video Assessment for Embodied Control A survey of reinforcement learning informed by natural language

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.832557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.048688Z digest=sha256:ceb0789d99bbf5f11c670e37ecbc27c90d977e8d058d723e0a0bbefc4224d744

Observation f16293f2-7118-4254-9828-d4e78ad76809 · outbound

This paper cites Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation.

Action- and Language-Conditioned Video Assessment for Embodied Control Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.818648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.052911Z digest=sha256:001c963684273f7f82dd95d4ac45799a267917bb7158330aa645f1058f5ecc24

Observation 70cd1617-4951-4e84-bc49-a1b9f644e150 · outbound

This paper cites Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robot.

Action- and Language-Conditioned Video Assessment for Embodied Control Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robot

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.803771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.057004Z digest=sha256:1bcb4482f03ded3d1e69705d3afcc1345e01d12de88c3f108e49234134becbec

Observation 82f919d6-c83c-4f55-8c1f-490264fa23ef · outbound

This paper cites Walk the talk: Connecting language, knowledge, and action in route instructions.

Action- and Language-Conditioned Video Assessment for Embodied Control Walk the talk: Connecting language, knowledge, and action in route instructions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.789435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.061250Z digest=sha256:25ce34cb3eb8bb444dea0206c787d3d5d2b3049eafcede8bccd689461cc0fef2

Observation 1bf8c1bc-07e8-40ca-9415-89a8c629f9c1 · outbound

This paper cites Toward understanding natural language directions.

Action- and Language-Conditioned Video Assessment for Embodied Control Toward understanding natural language directions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.774718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.065238Z digest=sha256:dc9569d03d56a681d2901bbef085ababe695fb0406d48a394c4bbfa8e0998972

Observation 2f8ffcc9-6ca5-4545-b5ab-b526ed1cb11b · outbound

This paper cites Tell Me Dave: Context-Sensitive Grounding of Natural Language to Manipulation Instructions.

Action- and Language-Conditioned Video Assessment for Embodied Control Tell Me Dave: Context-Sensitive Grounding of Natural Language to Manipulation Instructions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.758600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.069565Z digest=sha256:1f68321586d0fb0aaaa5235c40bab96d602e0d69b1e16985feade4940ea212d4

Observation 1c3f86e7-2fc7-4c8a-8989-8f3e92e32445 · outbound

This paper cites Grounding English Commands to Reward Functions.

Action- and Language-Conditioned Video Assessment for Embodied Control Grounding English Commands to Reward Functions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.742233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.073985Z digest=sha256:d190e6e26b0f882f014b8e572d9afac9216312c4d97e384bbcaca91d3877a0ee

Observation a6b9ce72-d0fd-4e7b-83fa-7fc28cf49b81 · outbound

This paper cites Learning language-conditioned robot behavior from offline data and crowd-sourced annotation.

Action- and Language-Conditioned Video Assessment for Embodied Control Learning language-conditioned robot behavior from offline data and crowd-sourced annotation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.725996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.078422Z digest=sha256:6d0597f0312a37228ac5c00cb9b62ad699a5ec61818869c8fff86331ad35afcb

Observation 9a76092d-810a-479b-ba14-b89331bae3fe · outbound

This paper cites Algorithms for inverse reinforcement learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Algorithms for inverse reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.710498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.082907Z digest=sha256:4ad4e94301ae07cd11ba01cdb8a767274e96f83c1740ff2a45af90e8dee66096

Observation d5f7fd72-74aa-4471-9d50-71d25c3d7044 · outbound

This paper cites Apprenticeship learning via inverse reinforcement learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Apprenticeship learning via inverse reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.695498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.088367Z digest=sha256:1c26c7de35b8318a6a96af8bf44052943c0e0ea271df63925878135221466159

Observation 92992ae9-aa61-4fef-b507-2a3dfa458998 · outbound

This paper cites Generative adversarial imitation learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Generative adversarial imitation learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.680865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.092958Z digest=sha256:4a4d3a5a21088db4e73c31aaced9f5f76ae402e8684db1289fed832c8951fa60

Observation a8447baa-7fb1-4f17-a60d-827673d5ac3e · outbound

This paper cites From language to goals: Inverse reinforcement learning for vision-based instruction following.

Action- and Language-Conditioned Video Assessment for Embodied Control From language to goals: Inverse reinforcement learning for vision-based instruction following

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.667083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.098089Z digest=sha256:7541975244143de1692e6815502e5c27982afb35b673e9d37ddfafcefa0c9988

Observation 1626dc02-bb2f-4774-a44e-c3750f02122c · outbound

This paper cites Learning to Understand Goal Specifications by Modelling Reward.

Action- and Language-Conditioned Video Assessment for Embodied Control Learning to Understand Goal Specifications by Modelling Reward

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.651779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.102920Z digest=sha256:3c48e0d13b45ccbcf990c08ea4b7d7c8520750b69f2983838d9d935cc9ccb18d

Observation 2c859d64-d627-4053-86b0-65671139af75 · outbound

This paper cites Reward Design with Language Models.

Action- and Language-Conditioned Video Assessment for Embodied Control Reward Design with Language Models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.637245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.107505Z digest=sha256:1833af05b4addcb2b7b46168256c84d0247c26418dce0753be0d75d25f99c7cd

Observation be7ad28a-7900-4234-a019-b335c08756fe · outbound

This paper cites Language to Rewards for Robotic Skill Synthesis.

Action- and Language-Conditioned Video Assessment for Embodied Control Language to Rewards for Robotic Skill Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:14:53.112288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:14:53.112288Z digest=sha256:a72721ad4bf74e0ef038bacefe304f25d4fc9f8b45f0682a1afc0f3e00ef0d94

Observation ef575489-4022-42da-9e92-42acbf469b03 · outbound

This paper cites Text2reward: Automated dense reward function generation for reinforcement learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Text2reward: Automated dense reward function generation for reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.622640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.118463Z digest=sha256:08013d73da9efc138fbce11a0105278c289384f5c18a7cceb8c2918ddfc83979

Observation 2812868e-82b7-4310-8c34-87a4bf4ed523 · outbound

This paper cites Eureka: Human-level reward design via coding large language models.

Action- and Language-Conditioned Video Assessment for Embodied Control Eureka: Human-level reward design via coding large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.607822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.123096Z digest=sha256:f4a3cac5b0a88a3c8e4aaa508b07b488cd6a8f4bded5ed6d7c39f426d87bba15

Observation 9152c17e-adbe-4bde-9c6f-75188065e6cb · outbound

This paper cites Robogen: Towards unleashing infinite data for automated robot learning via generative simulation.

Action- and Language-Conditioned Video Assessment for Embodied Control Robogen: Towards unleashing infinite data for automated robot learning via generative simulation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.593269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.127834Z digest=sha256:9da5f89809e60312683ef430687e1bf8b99addbcb3446ac61bacd2fa6023caee

Observation f4e19eba-c9bb-4a31-8415-3442f2618c5b · outbound

This paper cites an unresolved cited work.

Action- and Language-Conditioned Video Assessment for Embodied Control Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:14:53.578351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.132841Z digest=sha256:30f64010cd91084e1fe195d78d7441d9336962a0e3b83bd9f8d1bd370999721a

Observation c311f943-f30d-48f8-a359-d46742391e30 · outbound

This paper cites Vision-language models as success detectors.

Action- and Language-Conditioned Video Assessment for Embodied Control Vision-language models as success detectors

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.562755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.137917Z digest=sha256:535096b944ffe105751518148b950ac8b5586c3b5025547c1c96af758861f9a1

Observation 79ed6e64-f6f7-4650-a39d-e52cb36eda0d · outbound

This paper cites Rl-vlm-f: Reinforcement learning from vision language foundation model feedback.

Action- and Language-Conditioned Video Assessment for Embodied Control Rl-vlm-f: Reinforcement learning from vision language foundation model feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.547176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.142723Z digest=sha256:5f747305cbadc80a1f6ddabcfc1a729a3ac93f924f42989f5b434e24ad76401f

Observation ea720739-e555-41e2-835a-30e40e9dc585 · outbound

This paper cites Deep reinforcement learning from human preferences.

Action- and Language-Conditioned Video Assessment for Embodied Control Deep reinforcement learning from human preferences

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.531829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.147190Z digest=sha256:195142434c5a8febef107902bfa47c58fc8386ae01ecb6692faa70985fbd3965

Observation e6b72b7a-31fd-4d77-a1aa-174466f1f682 · outbound

This paper cites Interactive learning from policy-dependent human feedback.

Action- and Language-Conditioned Video Assessment for Embodied Control Interactive learning from policy-dependent human feedback

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.516014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.151536Z digest=sha256:3096772794fccfec253990cdc8d175d799be90c411187c37321132a5f46e4bf2

Observation 5d9ded7e-8af5-4068-a227-11530886fa47 · outbound

This paper cites Learning reward functions from scale feedback.

Action- and Language-Conditioned Video Assessment for Embodied Control Learning reward functions from scale feedback

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.501676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.156022Z digest=sha256:09e7a770c7f4754ae65190ca5db15fbcbb6737bc25cc6af76d5f46eae7b68752

Observation d6e4fbe1-6ec7-465a-aa33-5aad1c0f33c3 · outbound

This paper cites Rating-based reinforcement learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Rating-based reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.486183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.160639Z digest=sha256:f94ebe4eedb6d177fc5e6d4a080450a47dc69ed4178d10eb3f0be30e14302ec0

Observation f08808fd-23b0-441a-8b15-7802c64250ec · outbound

This paper cites NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions.

Action- and Language-Conditioned Video Assessment for Embodied Control NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.470836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.165188Z digest=sha256:8579880c6a63ec73ee17e3c6571ab3842d7954150a28444c325e263cd8d9a01e

Observation 7f87dfa0-47ed-4eaa-a44d-269017156998 · outbound

This paper cites From Representation to Reasoning: Towards Both Evidence and Commonsense Reasoning for Video Question-Answering.

Action- and Language-Conditioned Video Assessment for Embodied Control From Representation to Reasoning: Towards Both Evidence and Commonsense Reasoning for Video Question-Answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.455584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.169788Z digest=sha256:d7e70d9fd8433a485a35ef64a06a04a1445c05a9a0d232edebfe95b1b09d7004

Observation 5a586739-b824-4f3b-951c-a71c053bd80a · outbound

This paper cites Discovering the Real Association: Multimodal Causal Reasoning in Video Question Answering.

Action- and Language-Conditioned Video Assessment for Embodied Control Discovering the Real Association: Multimodal Causal Reasoning in Video Question Answering

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.439686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.174719Z digest=sha256:370db6ce703358c757aaaf9b5fde22774caa19b4741f477a6a0651e0cf1f435d

Observation 3ea084ff-a6a9-42da-9e30-7c2d31fafa24 · outbound

This paper cites MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning.

Action- and Language-Conditioned Video Assessment for Embodied Control MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.422990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.179150Z digest=sha256:b961a1a8ed5b49150b2bddf85f454a6e6a9858d90d14f3b636dd8d8efb869728

Observation 62820c07-afc0-4168-ba52-1cadd8b5d15b · outbound

This paper cites Open problems and fundamental limitations of reinforcement learning from human feedback.Trans.

Action- and Language-Conditioned Video Assessment for Embodied Control Open problems and fundamental limitations of reinforcement learning from human feedback.Trans

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.408760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.183259Z digest=sha256:2c95d7188ca8846f62bd3500a6cc04b00876e0a23ae65913d9ee5607bbc7375b

Observation 7910f685-215b-4657-9ba9-34acda89ef37 · outbound

This paper cites Visual prompt tuning.

Action- and Language-Conditioned Video Assessment for Embodied Control Visual prompt tuning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.393842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.187182Z digest=sha256:0394d8c751f226bf703f7b39f05ef90d31b5abcb06f266eee7cab7d3f81e7df8

Observation 24a82449-7086-4698-8b78-46dc96409596 · outbound

This paper cites Visual prompting via image inpainting.

Action- and Language-Conditioned Video Assessment for Embodied Control Visual prompting via image inpainting

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.377631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.191045Z digest=sha256:0c1f75e14733b8464b29155c272b17c07683bd31645241cb20b2b89ebc53e575

Observation 1bf75583-801e-4e89-aaeb-34e63fb22ce2 · outbound

This paper cites What does clip know about a red circle? visual prompt engineering for vlms.

Action- and Language-Conditioned Video Assessment for Embodied Control What does clip know about a red circle? visual prompt engineering for vlms

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.361646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.195003Z digest=sha256:93ef7bd9ef8dcebf15029bc9724dd3c9e24b5cc42bcea13fadb8f64d6213ebb1

Observation 286f98ee-f24b-4863-ad9f-f26b17dfc6aa · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Action- and Language-Conditioned Video Assessment for Embodied Control Offline reinforcement learning with implicit q-learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.344912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.199150Z digest=sha256:c651c52b75ea17afae988267b03cd5e6b9b7e0cd977f5011e4ff4a09c984f018

Observation 3e54e165-b3e7-47c0-8e8c-c497aadef2c5 · outbound

This paper cites Bootstrap your own skills: Learning to solve new tasks with large language model guidance.

Action- and Language-Conditioned Video Assessment for Embodied Control Bootstrap your own skills: Learning to solve new tasks with large language model guidance

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.328819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.203175Z digest=sha256:b172650a9cbd20bdd613547e4885c0d9b349a768bce1cc2df66cd3c219cb98bc

Observation cd29952b-f632-4979-af5d-f23506e1c664 · outbound

This paper cites Deep residual learning for image recognition.

Action- and Language-Conditioned Video Assessment for Embodied Control Deep residual learning for image recognition

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:14:53.311803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T00:14:53.207582Z digest=sha256:07c95d5728a75b11ca4530d0c898a827f8a587a4e5e843057c19a9b3e82f1b4c

Pith citing papers

No inbound Pith citation observations are available.