Pith. sign in

Paper Citation Record · LEDGER

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models

As of 18 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2504.13351.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.13351 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:15:06.119010Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T22:31:13.391859Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T22:32:10.973323Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b5e9176-f274-4bf5-8f78-19c8935e8aaf · outbound

This paper cites GPT-4 Technical Report.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.823996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.823996Z digest=sha256:a68438bd3317d59a29dca0234fa0ae94bca92a8cf53f5cb38f55e30f2dfd5e6b

Observation 953d85a8-f64c-4f59-9276-8907087ead28 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.829473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.829473Z digest=sha256:d80f5cdf5b4a971dc965784e81a2961c8b146bc9b9c5d4275e659bbc0fe2d93f

Observation 1b7c2986-8ea6-40be-ac2a-4df0580fb00d · outbound

This paper cites Human-to-Robot Imitation in the Wild.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Human-to-Robot Imitation in the Wild

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.834129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.834129Z digest=sha256:7e13a98750db4806daf3c606ce0547e2fabbfb65e0e03886e67fdb14a6783faa

Observation c24c9671-cfc9-4b7a-a0f7-8eda2e820442 · outbound

This paper cites Affordances from human videos as a versatile representation for robotics,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Affordances from human videos as a versatile representation for robotics,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.838737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.838737Z digest=sha256:e5712d665268e1879931fd3df02ed5a66c68fdf81ce88d19dd39f7fba0271397

Observation 6566c015-5722-4ca1-9442-d503edf5b012 · outbound

This paper cites Towards generalizable zero-shot manipulation via translating human interaction plans,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Towards generalizable zero-shot manipulation via translating human interaction plans,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:07.051450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.843956Z digest=sha256:ecafff052b27569e609d288372c1a300418aab9d1b11142e0e1cfa11e9ea5b0d

Observation f66ebc5d-31b3-4d1c-ba56-abeb9bd5b526 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Activitynet: A large-scale video benchmark for human activity understanding,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:07.037710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.849102Z digest=sha256:40f2dc27ac23b008cd6c59142edd0f8afeac2aa3a01f948ee616130b6980aa9d

Observation a6de1d3a-4ac8-4cc1-a888-30e335918dbc · outbound

This paper cites Procedure planning in instructional videos,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Procedure planning in instructional videos,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:07.023693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.854359Z digest=sha256:0aff677114eb14e2d83347d8dd417ee1fb14e1e9e6c18941be847732a8b8c9be

Observation d1a7a79a-8866-4e3c-bfaf-b82b29ebe58d · outbound

This paper cites Learning generalizable robotic reward functions from.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning generalizable robotic reward functions from

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:07.010164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.858352Z digest=sha256:f6c6c1cd4d61c9413a805a3fb3f31ca3ac03f69b3f002a2d81bec1ec6aa6e8af

Observation 1ec69fac-60e0-488d-806f-08500ddda4d0 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.996361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.862463Z digest=sha256:395efeccb5af45ff4b941499d7b1a7d73a63763c2724f54d4df8db351cbcf141

Observation 8e045913-0882-4d8d-a633-ff8ff2c3ea22 · outbound

This paper cites Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.866636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.866636Z digest=sha256:f0ab679a19d1c27138d6cc131259de7cf6a6fd0aee4992db7f44eb50af251557

Observation 562bd8c0-423d-4c76-af4c-4b8d28ab45b8 · outbound

This paper cites Can foundation models perform zero-shot task specification for robot manipulation?.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Can foundation models perform zero-shot task specification for robot manipulation?

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.981665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.871228Z digest=sha256:65f7605cf5c267e2c7e8ed5b1da577f0f9db1c439a73c12a1203ea7e0462036f

Observation 51339741-d041-4973-b24c-15288bb4f70f · outbound

This paper cites Scaling egocentric vision: The epic- kitchens dataset,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Scaling egocentric vision: The epic- kitchens dataset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.967585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.875734Z digest=sha256:b041ebe2a56225449435b2179bfd427f83c645558fffbd925d648d5603da3eaa

Observation 1b415199-27db-4206-a35b-1a90ea1d0228 · outbound

This paper cites Model-based inverse reinforcement learning from visual demonstrations,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Model-based inverse reinforcement learning from visual demonstrations,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.954253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.880003Z digest=sha256:aeb74f7451bbf6f0a5380a2734f6ee5259535077ab8a4d1bb3fa594d8cfdbf54

Observation 07a1c60b-2f91-4c1c-b40c-78ee1eb6c582 · outbound

This paper cites Perceptual Values from Observation.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Perceptual Values from Observation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.884007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.884007Z digest=sha256:9ce7e80a5d9f80a7d542adb34d18d95d5de3b8bc75fa2cb446b80c2ef8f1581c

Observation 2363f667-350e-4d73-98b8-e7f87053b07d · outbound

This paper cites Foundation Models in Robotics: Applications, Challenges, and the Future.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Foundation Models in Robotics: Applications, Challenges, and the Future

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.888688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.888688Z digest=sha256:7619f415023ac2bd30f3fbd1693c5ab35d7ccdb0eb1a78126009beea2be78b63

Observation 73711770-4d0f-494a-ad7c-18cacc6d290a · outbound

This paper cites The” something something.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models The” something something

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.940935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.893962Z digest=sha256:12c9ed31cadcf854be672279fa9d0db9e02506e42919a28b9a939eea65cd1e8e

Observation aba618a8-bbf5-45a9-a975-2baf93792f60 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.928077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.898804Z digest=sha256:8c836e03ca7104019259552ad5a08de986b836cbb9e69e58c4eb09f9ffc14889

Observation 332cdff3-e8e4-46a1-a16e-80336f017610 · outbound

This paper cites Gu et al., Rt-trajectory: Robotic task generalization via hindsight trajectory sketches , 2023.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Gu et al., Rt-trajectory: Robotic task generalization via hindsight trajectory sketches , 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.903048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.903048Z digest=sha256:1f1e527d332e6dee37cb8a235346cb192938af7bc4fa5c3fc287538d01515504

Observation 9df6f5cb-68e4-4f18-ae1b-b25a61b06d3e · outbound

This paper cites Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.907167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.907167Z digest=sha256:4954616723150749039b64186514e45a3422d31b7297772da914df92872de2ca

Observation 153673ef-0b34-408b-aea2-b7b00219c727 · outbound

This paper cites Neural task graphs: Generalizing to unseen tasks from a single video demonstration,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Neural task graphs: Generalizing to unseen tasks from a single video demonstration,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.906112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.911401Z digest=sha256:c12f2a7c65c285ff33aeafb8cc5101c72e0b414488d9d510f40d4763a709e97a

Observation 4e09e530-64a9-4abf-bffc-429cbe81a950 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.893030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.915400Z digest=sha256:0becbb89a54446f41e34b7f124c751f8c58496d8a3fc2b9fb4e2c6435924fbad

Observation ad73f741-e752-4740-8e4b-86eedd14fd18 · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.920527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.920527Z digest=sha256:acb86cf3f8739ba491ccc747a6c0e9d38601682504df9bbffdec52ff49dac3ab

Observation 626817ea-66c6-4e8f-bc38-27a963ee36c8 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.924813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.924813Z digest=sha256:bf7fdc029178465f769defd187d508873671665f8d56321f5719d3114c5658e5

Observation dea3d730-a29c-4b51-b11c-129b69886541 · outbound

This paper cites Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.929096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.929096Z digest=sha256:a800fb5de13f1fbe1ece69beff6abc58247dca4a6c3bf7759dc1dd88049de8f0

Observation ec8b220b-43bd-46c8-8859-31c04e85b903 · outbound

This paper cites Prompting visual-language models for efficient video understanding,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Prompting visual-language models for efficient video understanding,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.879575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.933595Z digest=sha256:de69b57677c55f619b4d291314904d7cec66e836e8e4af155b7a3b0b5824b611

Observation 0f4db8b3-28f7-406b-9960-8089c67be071 · outbound

This paper cites Human action recognition and prediction: A survey,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Human action recognition and prediction: A survey,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.865923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.937756Z digest=sha256:078e1d6c0c2f4aa149b2e9d1e0481edba913189106dec60e2174a7f1b308b2b5

Observation 7ffb5d53-ac82-452c-b8ac-cc0832839844 · outbound

This paper cites Anticipating human activities using object affordances for reactive robotic response,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Anticipating human activities using object affordances for reactive robotic response,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.853103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.941701Z digest=sha256:7e64790eebf03e98139fd159eea61b9ac9cb5d4daeef8de297f581866e7c4cde

Observation edf41b8e-010d-427b-9192-4445b94ad7da · outbound

This paper cites Graph inverse reinforcement learning from diverse videos,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Graph inverse reinforcement learning from diverse videos,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.840008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.945640Z digest=sha256:f79609377fc9bc776951e9c627b81c3cefeb419718024d122e427bfda9f649eb

Observation 524e4606-1f1d-494e-a216-e69ba2bf4290 · outbound

This paper cites HAKE: Human Activity Knowledge Engine.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models HAKE: Human Activity Knowledge Engine

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.949799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.949799Z digest=sha256:b899287921accdf910cd9d494f7587557f79712308569009817c7af440e14125

Observation 43e694a3-e04c-4a4c-9f6f-834dba6310ba · outbound

This paper cites Code as policies: Language model programs for embodied control,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Code as policies: Language model programs for embodied control,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.826654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.954444Z digest=sha256:eddfb28b986e1f71938cf14223ec2f0b3469120b18e881738ef33f4d0fd18d73

Observation ddfeea2d-bcaf-4cdb-8549-2177c55fee32 · outbound

This paper cites Learning to Learn Faster from Human Feedback with Language Model Predictive Control.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning to Learn Faster from Human Feedback with Language Model Predictive Control

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.958643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.958643Z digest=sha256:c60fb0a98c3be2d7da632db0229b90759179e136b6166c08abac6a11cef2520b

Observation 029c77d3-91c1-44d6-a7fb-7be115bd0f9d · outbound

This paper cites Text2motion: From natural language instructions to feasible plans,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Text2motion: From natural language instructions to feasible plans,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.813390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.963232Z digest=sha256:48d44bcda32f4ceb010b2f7f7bf28c55204c0d6ce742015760a0874ba88aa748

Observation 8334d377-0654-4f39-8059-f769e0eaf7d6 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.967780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.967780Z digest=sha256:d5b602862ee2189a0b2abc00acb1e366243b18daa7dcf29aafef63dd1b2dacdf

Observation af4a4e74-bada-4530-8546-c78ec375f196 · outbound

This paper cites Imitation from observation: Learning to imitate behaviors from raw video via context translation,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Imitation from observation: Learning to imitate behaviors from raw video via context translation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.799915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.972747Z digest=sha256:c774041552cd933e12a9c918313e63cabd5fd44500bb0dce906be431ba169cae

Observation 62ae18b9-ac15-436c-b5d5-dc9e7809140d · outbound

This paper cites VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.977805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.977805Z digest=sha256:3bd58acb19427d5e662834cc648a20566f53aea4eb97c64980dab87a4f86d5a1

Observation ca246feb-1907-4498-abb9-91b4e55c323f · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.982335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.982335Z digest=sha256:e1ab4d3afe705be2305bb4d021eff438f68491f3b1d96a94dda480f669eb7f21

Observation 8012cf00-5557-49e5-a3d7-f5eac04cd3be · outbound

This paper cites R+X: Retrieval and Execution from Everyday Human Videos.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models R+X: Retrieval and Execution from Everyday Human Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.986843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.986843Z digest=sha256:25af2022c70131d576685d5f813ea4adfa2ec8c506c047056664b7c1c6f7351d

Observation ecf863bd-b2cf-4e7f-92a8-d099dd57f61f · outbound

This paper cites Reconstructing hands in 3D with transformers,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Reconstructing hands in 3D with transformers,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.785927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.991085Z digest=sha256:1128c0c1056466845ec9ed408093ffa112a99b8b1b142e533508d9643b7d8dcd

Observation 3d0306c3-cfbf-44a0-ad5e-08931ec73954 · outbound

This paper cites Planning with large language models via corrective re-prompting,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Planning with large language models via corrective re-prompting,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.772228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:05.995264Z digest=sha256:0e59dc60f56b87a89f82b33c08fa6a5e771de5c98078059d1c407f0846dc6ff3

Observation ce308fe6-2918-458c-a391-e35be8176faa · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.999412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.999412Z digest=sha256:b23a2cf2b56cbbb98e6c0651ed2ddad770a9d7e6b756b43eb0dd0efd5d5a30c0

Observation 494da36c-0ae1-49eb-b631-6315d0951129 · outbound

This paper cites First-person activity forecasting with online inverse reinforcement learning,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models First-person activity forecasting with online inverse reinforcement learning,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.756810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.003775Z digest=sha256:82b8d1e0665556b82a640c12cc6c042eaf26783d0055260dde4bcf3a6cac742c

Observation 35af371c-7f12-42ac-9df6-63f1efbfe7c0 · outbound

This paper cites Reinforcement Learning with Videos: Combining Offline Observations with Interaction.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Reinforcement Learning with Videos: Combining Offline Observations with Interaction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.008300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.008300Z digest=sha256:334b28049b540a54ed903e48001439a026601fc0f51f65d570f825769641aac0

Observation c017f492-57b7-4e75-bb03-e249e5bd876e · outbound

This paper cites Learning predictive models from observation and interaction,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning predictive models from observation and interaction,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.743315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.012623Z digest=sha256:a23a345ea394bcbd4026fc561d8d8b542577f561b112c7b847bef64fb3d8b796

Observation fbfba823-2a36-4ea3-a700-7213cf749988 · outbound

This paper cites Time-contrastive networks: Self- supervised learning from video,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Time-contrastive networks: Self- supervised learning from video,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.729483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.016846Z digest=sha256:e31d34713892bedd6f56efa902ec1636ba94aa48e85c27c9aa547dd2fb8d5d9d

Observation 5936cc87-13c4-47de-8a41-75dc2ba438d8 · outbound

This paper cites Unsupervised Perceptual Rewards for Imitation Learning.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Unsupervised Perceptual Rewards for Imitation Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.020797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.020797Z digest=sha256:b566b53835d31a79076fe9f709c1f40cfdc8e89f4d0c123c45a3b3ab271e4b74

Observation 5d5cca8d-92fe-425d-bbe4-94d46b2f6314 · outbound

This paper cites Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.716265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.025092Z digest=sha256:9ec2e392692f0d2ff08bab8d6b25d649a3e7093635879c7234138288c778fbd1

Observation 53ea4167-f4bc-4391-9a69-c5eb279eaa1c · outbound

This paper cites Concept2robot: Learning manipulation concepts from instructions and human demonstrations,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Concept2robot: Learning manipulation concepts from instructions and human demonstrations,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.702427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.029196Z digest=sha256:9a2e469e4fb43579eab7072c38554c54fe88cd5781c9bb4e5b08dda74de350d3

Observation 64ac1d08-59bd-40d4-ae6d-1577697185fd · outbound

This paper cites Third-person visual imitation learning via decoupled hierarchical controller,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Third-person visual imitation learning via decoupled hierarchical controller,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.033284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.033284Z digest=sha256:fcdac4959d841b218969b2ca722bfc977cd0af3af8254cbc0ac4da3740e42b37

Observation 7fdefc1e-e9de-46d6-b321-ccb76c1d509b · outbound

This paper cites Videodex: Learning dexterity from internet videos,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Videodex: Learning dexterity from internet videos,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.679337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.037238Z digest=sha256:32e3b88a9f20731e9d59787471c26ff21b0e716361a8c717ca01961643d4d01d

Observation 9cbb0e9d-2e59-4dce-b269-e94561a12242 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Cliport: What and where pathways for robotic manipulation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.666143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.041470Z digest=sha256:d27ddcc380eb921f1d51efd5e9fac83b670b37c74b4a7e213b33b995dc44d2db

Observation b100e8b7-b69d-47f4-8f2b-9c497e515b8d · outbound

This paper cites Generalized planning in pddl domains with pretrained large language models,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Generalized planning in pddl domains with pretrained large language models,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.652030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.045413Z digest=sha256:9257566a72ee4e38c549f5badb981f127639abdbeeb3cfafb5c4584c5c091f5e

Observation c8ce1d2a-f309-4ba7-9fbe-079e8c5f9771 · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Progprompt: Generating situated robot task plans using large language models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.638380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.049345Z digest=sha256:4cde031decedd9553d15ffab3460618e57f8f5e3d5aa0b7b8f8ba3ba36d8a7e4

Observation 5e08acf3-1ab6-4dbe-b8fb-d79b87d0f243 · outbound

This paper cites AVID: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models AVID: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.053396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.053396Z digest=sha256:c61d9b4db59ea95930bf178c7a27d6c661012cc39bb09ed132c1a7c1388640ee

Observation bae1268d-aecf-4e21-afd8-befc75498813 · outbound

This paper cites Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.624307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.058134Z digest=sha256:14cfed533b1c408e9d6f670aa8ec8a64019664dd9c92d404ae8f463486ea4499

Observation 70c2604c-d6d6-4a9b-820f-01e4eaac77c2 · outbound

This paper cites MimicPlay: Long-Horizon Imitation Learning by Watching Human Play.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models MimicPlay: Long-Horizon Imitation Learning by Watching Human Play

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.062178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.062178Z digest=sha256:3974271faa9b741aeb43af482a66e48b3c5f40a3fdedc7ef0bdc7de9ac3ca557

Observation e029ff48-0892-4c15-ae94-e09bb9b4a6a5 · outbound

This paper cites Temporal segment networks for action recognition in videos,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Temporal segment networks for action recognition in videos,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.609454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.066608Z digest=sha256:6c48dfc4a105ad0932beacb10250fe4f342d1c89fce0de0a91943d523e9fb579

Observation ccfc0c3a-e781-4fc8-b1e8-42124a900da9 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.070735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.070735Z digest=sha256:fe166ed9fa598cff0f773e6dc3e5bd26dbe9ad57ee3a845a16081d5e0da22254

Observation 9bad7664-3736-4466-96e9-284f630f4666 · outbound

This paper cites You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.075484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.075484Z digest=sha256:7749bc0d06b14eebc5c8308a49f2d89d4c1ddd687e24c2b8f8badbfcf3aac16d

Observation 5a8437c9-a347-4a60-bd11-f753709bfdbd · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Any-point Trajectory Modeling for Policy Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.080003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.080003Z digest=sha256:56a4e5d3d40859bd75e2fb7b91d496efe54d9160fc9144abe0ef27f31a1b6d97

Observation 0fc53998-4711-45aa-89a6-21b936be4c15 · outbound

This paper cites Learning by watching: Physical imitation of manipulation skills from human videos,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning by watching: Physical imitation of manipulation skills from human videos,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.594823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.084224Z digest=sha256:68e03c1eb0cccd89113cfc9be6575c4557c1db3271a0b65e57f86282d6f4c5fa

Observation c7e02b1c-fd93-45ed-83c9-62b7ff4b66f5 · outbound

This paper cites R-c3d: Region convolutional 3d network for temporal activity detection,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models R-c3d: Region convolutional 3d network for temporal activity detection,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.580137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.088928Z digest=sha256:767f1d69fc873ac886df03ea0f0d2aff0c17bda24277e591a8af74e8e59e5693

Observation 543b7814-6507-4f88-a2f9-a6cb528eb278 · outbound

This paper cites Xskill: Cross embodiment skill discovery,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Xskill: Cross embodiment skill discovery,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.566205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.092975Z digest=sha256:330ed9f70248b88fadfd5aa367457aceff179df7a01ea2e1ec695119137119f8

Observation b81a5940-fb40-436a-aafb-63f13630f2c4 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.097274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.097274Z digest=sha256:ea98e1ecc40f9245f1b195aece50d2b76e8bf2377c2a69b6e5ab6ec252073a60

Observation f5a88b61-30fd-43b7-85df-92a144c2dde7 · outbound

This paper cites Language to Rewards for Robotic Skill Synthesis.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Language to Rewards for Robotic Skill Synthesis

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.101518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.101518Z digest=sha256:511313672a787648f8897ba301e07a3a995071b3e68ec931c4f292a9f700793b

Observation b4d1615e-b21a-4e42-8b67-7f36d58670b9 · outbound

This paper cites Xirl: Cross-embodiment inverse rein- forcement learning,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Xirl: Cross-embodiment inverse rein- forcement learning,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.551893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.105941Z digest=sha256:eb0caf091860568dd3919eadf502ac7cc89449640ae1860f0b9851715929728b

Observation 8f7de078-182a-49d8-938e-d4883ebcce9f · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.110419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.110419Z digest=sha256:93d2ed11eb422bc752f53bdd3a9a67d99217546d935fbb963e3c7eb3e6c8fa67

Observation dda92637-4a19-45f2-bcf8-3250d59a31c1 · outbound

This paper cites Actionformer: Localizing moments of actions with transformers,.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Actionformer: Localizing moments of actions with transformers,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:15:06.538148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:15:06.114777Z digest=sha256:fe3a3832f29e2a2bce9a1cae2c046e1991cf7e4978f8793903cfb0aa1a440903

Observation 530ccaf9-7f76-4434-9786-94972fb0ade5 · outbound

This paper cites Vision-based Manipulation from Single Human Video with Open-World Object Graphs.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Vision-based Manipulation from Single Human Video with Open-World Object Graphs

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:06.119010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:06.119010Z digest=sha256:713640fe802e11045877d4fadfaa9108dd67131177320ae0f21963d14df16609

Pith citing papers

Observation 4aaef4b5-f425-411a-b59c-31eadb2ff0d2 · inbound

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views cites this paper.

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:32:10.976303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T22:31:13.391859Z digest=sha256:23b36a06a04778b070e3dfbb3091fa0b2e05f84679c6538f1fdfa253b2d74ebb