Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:15:06.119010Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2504.13351.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:15:06.119010Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-17T22:31:13.391859Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T22:32:10.973323Z
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8b5e9176-f274-4bf5-8f78-19c8935e8aaf · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 953d85a8-f64c-4f59-9276-8907087ead28 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b7c2986-8ea6-40be-ac2a-4df0580fb00d · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Human-to-Robot Imitation in the Wild
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24c9671-cfc9-4b7a-a0f7-8eda2e820442 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Affordances from human videos as a versatile representation for robotics,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6566c015-5722-4ca1-9442-d503edf5b012 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Towards generalizable zero-shot manipulation via translating human interaction plans,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f66ebc5d-31b3-4d1c-ba56-abeb9bd5b526 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Activitynet: A large-scale video benchmark for human activity understanding,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a6de1d3a-4ac8-4cc1-a888-30e335918dbc · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Procedure planning in instructional videos,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d1a7a79a-8866-4e3c-bfaf-b82b29ebe58d · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning generalizable robotic reward functions from
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1ec69fac-60e0-488d-806f-08500ddda4d0 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e045913-0882-4d8d-a633-ff8ff2c3ea22 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 562bd8c0-423d-4c76-af4c-4b8d28ab45b8 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Can foundation models perform zero-shot task specification for robot manipulation?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 51339741-d041-4973-b24c-15288bb4f70f · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Scaling egocentric vision: The epic- kitchens dataset,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1b415199-27db-4206-a35b-1a90ea1d0228 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Model-based inverse reinforcement learning from visual demonstrations,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 07a1c60b-2f91-4c1c-b40c-78ee1eb6c582 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Perceptual Values from Observation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2363f667-350e-4d73-98b8-e7f87053b07d · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Foundation Models in Robotics: Applications, Challenges, and the Future
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73711770-4d0f-494a-ad7c-18cacc6d290a · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models The” something something
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aba618a8-bbf5-45a9-a975-2baf93792f60 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Ego4d: Around the world in 3,000 hours of egocentric video,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 332cdff3-e8e4-46a1-a16e-80336f017610 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Gu et al., Rt-trajectory: Robotic task generalization via hindsight trajectory sketches , 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9df6f5cb-68e4-4f18-ae1b-b25a61b06d3e · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 153673ef-0b34-408b-aea2-b7b00219c727 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Neural task graphs: Generalizing to unseen tasks from a single video demonstration,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4e09e530-64a9-4abf-bffc-429cbe81a950 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ad73f741-e752-4740-8e4b-86eedd14fd18 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 626817ea-66c6-4e8f-bc38-27a963ee36c8 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Inner Monologue: Embodied Reasoning through Planning with Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dea3d730-a29c-4b51-b11c-129b69886541 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec8b220b-43bd-46c8-8859-31c04e85b903 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Prompting visual-language models for efficient video understanding,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f4db8b3-28f7-406b-9960-8089c67be071 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Human action recognition and prediction: A survey,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7ffb5d53-ac82-452c-b8ac-cc0832839844 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Anticipating human activities using object affordances for reactive robotic response,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation edf41b8e-010d-427b-9192-4445b94ad7da · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Graph inverse reinforcement learning from diverse videos,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 524e4606-1f1d-494e-a216-e69ba2bf4290 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models HAKE: Human Activity Knowledge Engine
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43e694a3-e04c-4a4c-9f6f-834dba6310ba · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Code as policies: Language model programs for embodied control,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ddfeea2d-bcaf-4cdb-8549-2177c55fee32 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning to Learn Faster from Human Feedback with Language Model Predictive Control
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 029c77d3-91c1-44d6-a7fb-7be115bd0f9d · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Text2motion: From natural language instructions to feasible plans,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8334d377-0654-4f39-8059-f769e0eaf7d6 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af4a4e74-bada-4530-8546-c78ec375f196 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Imitation from observation: Learning to imitate behaviors from raw video via context translation,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 62ae18b9-ac15-436c-b5d5-dc9e7809140d · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca246feb-1907-4498-abb9-91b4e55c323f · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8012cf00-5557-49e5-a3d7-f5eac04cd3be · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models R+X: Retrieval and Execution from Everyday Human Videos
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecf863bd-b2cf-4e7f-92a8-d099dd57f61f · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Reconstructing hands in 3D with transformers,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3d0306c3-cfbf-44a0-ad5e-08931ec73954 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Planning with large language models via corrective re-prompting,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ce308fe6-2918-458c-a391-e35be8176faa · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 494da36c-0ae1-49eb-b631-6315d0951129 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models First-person activity forecasting with online inverse reinforcement learning,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 35af371c-7f12-42ac-9df6-63f1efbfe7c0 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Reinforcement Learning with Videos: Combining Offline Observations with Interaction
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c017f492-57b7-4e75-bb03-e249e5bd876e · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning predictive models from observation and interaction,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fbfba823-2a36-4ea3-a700-7213cf749988 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Time-contrastive networks: Self- supervised learning from video,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5936cc87-13c4-47de-8a41-75dc2ba438d8 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Unsupervised Perceptual Rewards for Imitation Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d5cca8d-92fe-425d-bbe4-94d46b2f6314 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 53ea4167-f4bc-4391-9a69-c5eb279eaa1c · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Concept2robot: Learning manipulation concepts from instructions and human demonstrations,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 64ac1d08-59bd-40d4-ae6d-1577697185fd · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Third-person visual imitation learning via decoupled hierarchical controller,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fdefc1e-e9de-46d6-b321-ccb76c1d509b · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Videodex: Learning dexterity from internet videos,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9cbb0e9d-2e59-4dce-b269-e94561a12242 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Cliport: What and where pathways for robotic manipulation,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b100e8b7-b69d-47f4-8f2b-9c497e515b8d · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Generalized planning in pddl domains with pretrained large language models,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c8ce1d2a-f309-4ba7-9fbe-079e8c5f9771 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Progprompt: Generating situated robot task plans using large language models,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5e08acf3-1ab6-4dbe-b8fb-d79b87d0f243 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models AVID: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae1268d-aecf-4e21-afd8-befc75498813 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 70c2604c-d6d6-4a9b-820f-01e4eaac77c2 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models MimicPlay: Long-Horizon Imitation Learning by Watching Human Play
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e029ff48-0892-4c15-ae94-e09bb9b4a6a5 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Temporal segment networks for action recognition in videos,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ccfc0c3a-e781-4fc8-b1e8-42124a900da9 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bad7664-3736-4466-96e9-284f630f4666 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models You Only Demonstrate Once: Category-Level Manipulation from Single Visual Demonstration
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a8437c9-a347-4a60-bd11-f753709bfdbd · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Any-point Trajectory Modeling for Policy Learning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc53998-4711-45aa-89a6-21b936be4c15 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Learning by watching: Physical imitation of manipulation skills from human videos,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c7e02b1c-fd93-45ed-83c9-62b7ff4b66f5 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models R-c3d: Region convolutional 3d network for temporal activity detection,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 543b7814-6507-4f88-a2f9-a6cb528eb278 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Xskill: Cross embodiment skill discovery,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b81a5940-fb40-436a-aafb-63f13630f2c4 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5a88b61-30fd-43b7-85df-92a144c2dde7 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Language to Rewards for Robotic Skill Synthesis
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4d1615e-b21a-4e42-8b67-7f36d58670b9 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Xirl: Cross-embodiment inverse rein- forcement learning,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8f7de078-182a-49d8-938e-d4883ebcce9f · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dda92637-4a19-45f2-bcf8-3250d59a31c1 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Actionformer: Localizing moments of actions with transformers,
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 530ccaf9-7f76-4434-9786-94972fb0ade5 · outbound
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models Vision-based Manipulation from Single Human Video with Open-World Object Graphs
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aaef4b5-f425-411a-b59c-31eadb2ff0d2 · inbound
Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.