Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T04:27:53.257043Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2607.27180.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T04:27:53.257043Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6ab25c09-464f-4e97-bf1f-5ee681b944f6 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6861f628-8709-404a-a0c8-b3902d8df2d1 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5eeda64-7f19-4f92-80c4-e96181dae722 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Turner, Eric Undersander, and Tsung-Yen Yang
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aa073d81-8d68-4964-98de-ad1b587e7176 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? SpatialVLM : Endowing vision-language models with spatial reasoning capabilities
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e374e4c5-7954-439f-8bc9-b776274cfdf5 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? LoTa-Bench : Benchmarking language-oriented task planners for embodied agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8292f7df-5bfe-42c3-a35a-f19f9cc0e854 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Umo: Unified in-context learning unlocks motion foundation model priors
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0bf41762-0765-4b22-9521-7b9495f08f82 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Bullet physics simulation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6a06d3cd-4200-422d-8e44-0f91d5c528a3 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Moving by looking: Towards vision-driven avatar motion generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation da4f47b3-2073-4f65-bba5-ccecfc645fc8 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5bc2a7c2-77e5-4910-8028-bd549859a556 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Manipulate-anything: Automating real-world robots using vision-language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 095dfea5-81cc-40ad-8356-a69877df88ad · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Sanketi, Dorsa Sadigh, Chelsea Finn, and Sergey Levine
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 456824bd-d734-45ca-8459-3aac11782433 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Generating diverse and natural 3d human motions from text
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eb38bdd-4631-4ab4-9fb1-64027ef29cb7 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? MoMask : Generative masked modeling of 3d human motions
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7794d1f2-a8cc-4948-aca3-45d37c17222f · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdfb89a2-9142-4ec9-932b-d9dfa3b7ebb2 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af4541cf-c2f4-488b-8491-bc5e1a61fcdb · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? VoxPoser : Composable 3d value maps for robotic manipulation with language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cb7a4c46-791b-4cae-955a-3f4ccb7768dd · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Inner monologue: Embodied reasoning through planning with language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44b99768-d4d5-49c3-80c5-a0109b96d0fa · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Como: Controllable motion generation through language guided pose code editing
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6b290da6-ee40-4ae9-be7e-fe714aa8d0c9 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Do as i can, not as i say: Grounding language in robotic affordances
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8388b5cf-ba48-47ea-89fb-4a03d3d66643 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? IAM: Identity-Aware Human Motion and Shape Joint Generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c858dc4b-21cd-4c2d-bd0e-5d028bf41d97 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Chang, and Manolis Savva
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53816b94-5b39-493e-8ab6-a7c37085c324 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Foster, Pannag R
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3b3463b-aa57-4496-b26f-f08d8a18dc20 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? MolmoAct: Action Reasoning Models that can Reason in Space
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 654f3d73-4a64-4859-8225-8005446494cb · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? BEHAVIOR-1K : A benchmark for embodied AI with 1,000 everyday activities and realistic simulation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b14b3cb5-72ac-49e0-b444-a6d3609feed1 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d179ca96-d1d7-49ea-bcb0-a307c43329b0 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Embodied agent interface: Benchmarking LLMs for embodied decision making
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b6ddff6-51e1-449c-b1cc-f2c41190a46e · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Llamo: Scaling pretrained language models for unified motion understanding and generation with continuous autoregressive tokens
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eb01743d-6420-4798-a91b-78f6e8169446 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Genhsi: Controllable generation of human-scene interaction videos
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ee7b26eb-85ec-4bc1-b01a-50cc68fbbba5 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Code as policies: Language model programs for embodied control
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dacf82d1-d40d-43ee-9137-9db50bbaaf7b · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187746f4-ff8d-47eb-bbc6-8f18e60f9021 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d128dc5-a29f-4764-a2f1-91c139c6ae33 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? A survey on vision-language-action models for embodied AI
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe402b2-5200-43fd-b1d5-c2e165f74644 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation def29d3b-0927-4b45-a1ce-5179a2add6b8 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? VirtualHome : Simulating household activities via programs
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c1c0467c-3d70-4ec6-b36b-dc28c675986e · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakrishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara Rai, and Roozbeh Mottaghi
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 185de51e-d99e-463b-ac14-77f4d674669e · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Habitat: A platform for embodied AI research
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb40735d-4807-495d-8c7f-8efac0cf7ac8 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? ALFRED : A benchmark for interpreting grounded instructions for everyday tasks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f7444fe4-4ebb-4543-859f-c584b270ba71 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Bailando: 3 D dance generation by actor-critic GPT with choreographic memory
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4a14de67-bbf9-4f72-af83-dd893038eb70 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Bailando++: 3 D dance GPT with choreographic memory
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0d3f88fb-894d-45f7-a2aa-3a0a19f60417 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Duolando: Follower GPT with off-policy reinforcement learning for dance accompaniment
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6cd900b1-e2e8-4fcd-a52f-e75692678fb7 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Half-Physics: Enabling Kinematic 3D Human Model with Physical Interactions
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a154a73-2625-4d63-96df-3574c6949044 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Sadler, Wei-Lun Chao, and Yu Su
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 21f1e0cb-c8f9-4e01-beee-5ebcf618b83c · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9979aaa9-cf98-430c-9d65-8678ed2ff504 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Cradle: Empowering Foundation Agents Towards General Computer Control
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b70710c0-b998-42d1-b684-19c919fbbe9c · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1b7e799c-b6ee-494b-ac5a-b9ed8c0f6087 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Voyager: An open-ended embodied agent with large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3c074f47-6e14-450c-bf47-633fb999e084 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Describe, explain, plan and select: Interactive planning with LLMs enables open-world multi-task agents
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ad0f9f70-d0d0-453c-ba0c-88a6e417d331 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Text2interact: High-fidelity and diverse text-to-two-person interaction generation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffcc9334-c785-4714-ba36-c1197e77bbba · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? OmniControl : Control any joint at any time for human motion generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e3d3914-2827-4669-838d-89d3bc8b3f73 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e90fffc-49ea-4a6e-9185-94db7348aa6b · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? EmbodiedBench : Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e3695817-2f10-44fd-987d-948aeb894aef · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? PhysDiff : Physics-guided human motion diffusion model
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abee7c03-b433-43e5-be08-9769e454be1c · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? VideoGameBench: Can Vision-Language Models complete popular video games?
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec644601-6dbe-46d6-83ec-76a2517f540c · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Egoreact: Egocentric video-driven 3d human reaction generation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8832dcaf-2362-47ed-9ad6-5ffde4835f41 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? The wanderings of odysseus in 3d scenes
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 59f315d1-c39d-4469-8de1-0cafdf9c5f23 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e5da0000-ac44-4a54-998e-dc1069ad0619 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? Synthesizing diverse human motions in 3d indoor scenes
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 81391ef6-c7d9-482e-8da4-8b14d2e44362 · outbound
HumanCLAW: Can Vision-Language Models Act Through a Body? RT-2 : Vision-language-action models transfer web knowledge to robotic control
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.