Pith. sign in

Paper Citation Record · LEDGER

HumanCLAW: Can Vision-Language Models Act Through a Body?

As of 10 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2607.27180.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.27180 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:27:53.257043Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact3
  • verified fuzzy30
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ab25c09-464f-4e97-bf1f-5ee681b944f6 · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

HumanCLAW: Can Vision-Language Models Act Through a Body? Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.527603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.070094Z digest=sha256:ba11427590b2845d27cdee24cfa8df7d446284062a55e66e016f9a644e30c00e

Observation 6861f628-8709-404a-a0c8-b3902d8df2d1 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

HumanCLAW: Can Vision-Language Models Act Through a Body? $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.075264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.075264Z digest=sha256:7407372dcdac0b49f2938aa8b00792108274ae22c888201c7d75f0b1333a491f

Observation a5eeda64-7f19-4f92-80c4-e96181dae722 · outbound

This paper cites Turner, Eric Undersander, and Tsung-Yen Yang.

HumanCLAW: Can Vision-Language Models Act Through a Body? Turner, Eric Undersander, and Tsung-Yen Yang

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.518174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.078982Z digest=sha256:7e7d0552ecf617ff048a2e118772565884a7d59d9940fef396ded561f8116e1b

Observation aa073d81-8d68-4964-98de-ad1b587e7176 · outbound

This paper cites SpatialVLM : Endowing vision-language models with spatial reasoning capabilities.

HumanCLAW: Can Vision-Language Models Act Through a Body? SpatialVLM : Endowing vision-language models with spatial reasoning capabilities

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.508738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.082197Z digest=sha256:2a56dce6108ba87c0bb77ac0bceece94ffb0dbe79ff95703753e36b49879ef72

Observation e374e4c5-7954-439f-8bc9-b776274cfdf5 · outbound

This paper cites LoTa-Bench : Benchmarking language-oriented task planners for embodied agents.

HumanCLAW: Can Vision-Language Models Act Through a Body? LoTa-Bench : Benchmarking language-oriented task planners for embodied agents

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.499028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.086299Z digest=sha256:d48126cb1207efbc81317eff036810eadc822a1cc0604b2cfcbfd540e5687bc2

Observation 8292f7df-5bfe-42c3-a35a-f19f9cc0e854 · outbound

This paper cites Umo: Unified in-context learning unlocks motion foundation model priors.

HumanCLAW: Can Vision-Language Models Act Through a Body? Umo: Unified in-context learning unlocks motion foundation model priors

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-05T04:27:54.192337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.089644Z digest=sha256:353557ddcbe1ffe7cc0bbef71f35fd04b942bffec8d2e0028556ae9d7f4e8e8f

Observation 0bf41762-0765-4b22-9521-7b9495f08f82 · outbound

This paper cites Bullet physics simulation.

HumanCLAW: Can Vision-Language Models Act Through a Body? Bullet physics simulation

Reference 7

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T04:27:54.127642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.093124Z digest=sha256:250d46100d01e52e7bce733ad9d632f6369eb5a08bf3fda432e9e4d8f7e14351

Observation 6a06d3cd-4200-422d-8e44-0f91d5c528a3 · outbound

This paper cites Moving by looking: Towards vision-driven avatar motion generation.

HumanCLAW: Can Vision-Language Models Act Through a Body? Moving by looking: Towards vision-driven avatar motion generation

Reference 8

Resolution
verified exact
raw_fallback, observed 2026-08-05T04:27:54.043473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.096083Z digest=sha256:2295a2607fd5bb5e4e1449e042adcf3d7d9fe44a51b00c83448ef217253a0860

Observation da4f47b3-2073-4f65-bba5-ccecfc645fc8 · outbound

This paper cites an unresolved cited work.

HumanCLAW: Can Vision-Language Models Act Through a Body? Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:27:54.488985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.099218Z digest=sha256:cb03ff822f8835dc613c4b440fe3a236333c918869121da6b24895104b7788a8

Observation 5bc2a7c2-77e5-4910-8028-bd549859a556 · outbound

This paper cites Manipulate-anything: Automating real-world robots using vision-language models.

HumanCLAW: Can Vision-Language Models Act Through a Body? Manipulate-anything: Automating real-world robots using vision-language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.478561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.102107Z digest=sha256:a2822d9a8d52ec704bb377846486d6843e268d4422e916a90e2125c2c6d552f6

Observation 095dfea5-81cc-40ad-8356-a69877df88ad · outbound

This paper cites Sanketi, Dorsa Sadigh, Chelsea Finn, and Sergey Levine.

HumanCLAW: Can Vision-Language Models Act Through a Body? Sanketi, Dorsa Sadigh, Chelsea Finn, and Sergey Levine

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.106011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.106011Z digest=sha256:83d067cc7468a5c03e72d40739a4c8aa60abefc3e624d5a125fd715d2e6a542a

Observation 456824bd-d734-45ca-8459-3aac11782433 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

HumanCLAW: Can Vision-Language Models Act Through a Body? Generating diverse and natural 3d human motions from text

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.109145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.109145Z digest=sha256:b7041c2a38a4ded6237c2227edb6dc076f79ff0defd94245c1591932b511f98f

Observation 4eb38bdd-4631-4ab4-9fb1-64027ef29cb7 · outbound

This paper cites MoMask : Generative masked modeling of 3d human motions.

HumanCLAW: Can Vision-Language Models Act Through a Body? MoMask : Generative masked modeling of 3d human motions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.468555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.111900Z digest=sha256:3f13d48fd9af3097590410507c762c0087944ce50c326bea75783197c29ca0fa

Observation 7794d1f2-a8cc-4948-aca3-45d37c17222f · outbound

This paper cites an unresolved cited work.

HumanCLAW: Can Vision-Language Models Act Through a Body? Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.115153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.115153Z digest=sha256:d788725b5e6d1f00961988862406755784a8c886feb6afe7a8af12cebeb9e068

Observation cdfb89a2-9142-4ec9-932b-d9dfa3b7ebb2 · outbound

This paper cites ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop.

HumanCLAW: Can Vision-Language Models Act Through a Body? ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.118055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.118055Z digest=sha256:d8cd8607913a625aab5062481d3ccdfbfad0dfe86049620c83e6897ef0793c0e

Observation af4541cf-c2f4-488b-8491-bc5e1a61fcdb · outbound

This paper cites VoxPoser : Composable 3d value maps for robotic manipulation with language models.

HumanCLAW: Can Vision-Language Models Act Through a Body? VoxPoser : Composable 3d value maps for robotic manipulation with language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.458537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.121290Z digest=sha256:d6b24e7e14c6ba3df3680e643170c8365a4fbdb5bb1850ce816d7f014f643154

Observation cb7a4c46-791b-4cae-955a-3f4ccb7768dd · outbound

This paper cites Inner monologue: Embodied reasoning through planning with language models.

HumanCLAW: Can Vision-Language Models Act Through a Body? Inner monologue: Embodied reasoning through planning with language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.448682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.124955Z digest=sha256:7d2d3d05b75bf4693df178f2c71ad4b921f197d849e1b936dc9e7ee06485da3e

Observation 44b99768-d4d5-49c3-80c5-a0109b96d0fa · outbound

This paper cites Como: Controllable motion generation through language guided pose code editing.

HumanCLAW: Can Vision-Language Models Act Through a Body? Como: Controllable motion generation through language guided pose code editing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.439160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.127826Z digest=sha256:7d2367ccc7b98d48c7486a9e2f6007b734782246879d51aad7b1562953d361ab

Observation 6b290da6-ee40-4ae9-be7e-fe714aa8d0c9 · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

HumanCLAW: Can Vision-Language Models Act Through a Body? Do as i can, not as i say: Grounding language in robotic affordances

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.429200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.130780Z digest=sha256:dcad0aea8056c209d21436e6c23691dec3932433aef035516bc9003a6d436459

Observation 8388b5cf-ba48-47ea-89fb-4a03d3d66643 · outbound

This paper cites IAM: Identity-Aware Human Motion and Shape Joint Generation.

HumanCLAW: Can Vision-Language Models Act Through a Body? IAM: Identity-Aware Human Motion and Shape Joint Generation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:27:53.814598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.133792Z digest=sha256:78c6776ceb1f88abb4503f979ccad34d1ddaac4b1381917a3bce60acf557aa41

Observation c858dc4b-21cd-4c2d-bd0e-5d028bf41d97 · outbound

This paper cites Chang, and Manolis Savva.

HumanCLAW: Can Vision-Language Models Act Through a Body? Chang, and Manolis Savva

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.137106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.137106Z digest=sha256:399d1e2c39f442950181ad52a196ce8a622063b5bac2137180c0c4afb7d62c3c

Observation 53816b94-5b39-493e-8ab6-a7c37085c324 · outbound

This paper cites Foster, Pannag R.

HumanCLAW: Can Vision-Language Models Act Through a Body? Foster, Pannag R

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.419383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.140217Z digest=sha256:b191747bcbc7a958ae5e670174aadc4981851a0d8da500530bced052b60f1c69

Observation f3b3463b-aa57-4496-b26f-f08d8a18dc20 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

HumanCLAW: Can Vision-Language Models Act Through a Body? MolmoAct: Action Reasoning Models that can Reason in Space

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.143134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.143134Z digest=sha256:19077b2f1efee3c921ddaf7a53c97f61192b93a1b8bd7a61fc2d7c35da230ee3

Observation 654f3d73-4a64-4859-8225-8005446494cb · outbound

This paper cites BEHAVIOR-1K : A benchmark for embodied AI with 1,000 everyday activities and realistic simulation.

HumanCLAW: Can Vision-Language Models Act Through a Body? BEHAVIOR-1K : A benchmark for embodied AI with 1,000 everyday activities and realistic simulation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.410027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.146428Z digest=sha256:2033462eeb39c8a435c1524fae3843acfc49379381d277e7fdee99eb61575130

Observation b14b3cb5-72ac-49e0-b444-a6d3609feed1 · outbound

This paper cites Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations.

HumanCLAW: Can Vision-Language Models Act Through a Body? Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.149718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.149718Z digest=sha256:9e776962ecbafff55c63edac77ca5f1d7b49f8807f19e296ef1a05584a0d0b4e

Observation d179ca96-d1d7-49ea-bcb0-a307c43329b0 · outbound

This paper cites Embodied agent interface: Benchmarking LLMs for embodied decision making.

HumanCLAW: Can Vision-Language Models Act Through a Body? Embodied agent interface: Benchmarking LLMs for embodied decision making

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.153319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.153319Z digest=sha256:ee91054730f01d78f1c813994f28f2aa9eb8229fad85833e09cfaf28dd64e39a

Observation 7b6ddff6-51e1-449c-b1cc-f2c41190a46e · outbound

This paper cites Llamo: Scaling pretrained language models for unified motion understanding and generation with continuous autoregressive tokens.

HumanCLAW: Can Vision-Language Models Act Through a Body? Llamo: Scaling pretrained language models for unified motion understanding and generation with continuous autoregressive tokens

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.400013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.156338Z digest=sha256:95689e546ff670825bf8cfd174ca77aa595678e28a301189c59d349e19349a16

Observation eb01743d-6420-4798-a91b-78f6e8169446 · outbound

This paper cites Genhsi: Controllable generation of human-scene interaction videos.

HumanCLAW: Can Vision-Language Models Act Through a Body? Genhsi: Controllable generation of human-scene interaction videos

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.390146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.159675Z digest=sha256:f673776bb94346579af0b3c001a8612ab7e5212947cbe544f136c10011bf8897

Observation ee7b26eb-85ec-4bc1-b01a-50cc68fbbba5 · outbound

This paper cites Code as policies: Language model programs for embodied control.

HumanCLAW: Can Vision-Language Models Act Through a Body? Code as policies: Language model programs for embodied control

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.380611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.162728Z digest=sha256:e7d217db100c2095782cfac6cb7e50b702ba5fef8dccc056e52333825da52964

Observation dacf82d1-d40d-43ee-9137-9db50bbaaf7b · outbound

This paper cites Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning.

HumanCLAW: Can Vision-Language Models Act Through a Body? Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.165796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.165796Z digest=sha256:1e4336eaa199b038b3fbdf56192385c32d4ec4d4f4eb5060502617347e528493

Observation 187746f4-ff8d-47eb-bbc6-8f18e60f9021 · outbound

This paper cites VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents.

HumanCLAW: Can Vision-Language Models Act Through a Body? VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.169379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.169379Z digest=sha256:045bfbc90f6a3f6fab64826b8a04ae7ce052e98bdb19658f24c4d4fbd83d6c7f

Observation 6d128dc5-a29f-4764-a2f1-91c139c6ae33 · outbound

This paper cites A survey on vision-language-action models for embodied AI.

HumanCLAW: Can Vision-Language Models Act Through a Body? A survey on vision-language-action models for embodied AI

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.172713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.172713Z digest=sha256:cc9e1e580f3d7dc0c2b1842dcbdc8edb5302ddaf0686955241828f1773a8abe9

Observation 3fe402b2-5200-43fd-b1d5-c2e165f74644 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

HumanCLAW: Can Vision-Language Models Act Through a Body? GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.176051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.176051Z digest=sha256:a6cd0402930b64989604623894d887dfa158fb0656cbdf1c2a641f1fa0c56d03

Observation def29d3b-0927-4b45-a1ce-5179a2add6b8 · outbound

This paper cites VirtualHome : Simulating household activities via programs.

HumanCLAW: Can Vision-Language Models Act Through a Body? VirtualHome : Simulating household activities via programs

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.371114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.179226Z digest=sha256:0e4263d823586e993d76d15960813f105f8db7a9510ba12c2e6cb45167bbc710

Observation c1c0467c-3d70-4ec6-b36b-dc28c675986e · outbound

This paper cites Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakrishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara Rai, and Roozbeh Mottaghi.

HumanCLAW: Can Vision-Language Models Act Through a Body? Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakrishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara Rai, and Roozbeh Mottaghi

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.361693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.182274Z digest=sha256:a12ffe7c309ed26881f66d8ddb8b1ec0a2a4f73515db9ac5fff53225d9701824

Observation 185de51e-d99e-463b-ac14-77f4d674669e · outbound

This paper cites Habitat: A platform for embodied AI research.

HumanCLAW: Can Vision-Language Models Act Through a Body? Habitat: A platform for embodied AI research

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.185235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.185235Z digest=sha256:1b033342a23dc75a8d12813867bc1745818b09edea6308595144f233d2e1bb1b

Observation fb40735d-4807-495d-8c7f-8efac0cf7ac8 · outbound

This paper cites ALFRED : A benchmark for interpreting grounded instructions for everyday tasks.

HumanCLAW: Can Vision-Language Models Act Through a Body? ALFRED : A benchmark for interpreting grounded instructions for everyday tasks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.352196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.188247Z digest=sha256:e4cea33d8d8e3d589c79373a0675a3b10a3a66795c1aab201d51f71ede172f87

Observation f7444fe4-4ebb-4543-859f-c584b270ba71 · outbound

This paper cites Bailando: 3 D dance generation by actor-critic GPT with choreographic memory.

HumanCLAW: Can Vision-Language Models Act Through a Body? Bailando: 3 D dance generation by actor-critic GPT with choreographic memory

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.342645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.192092Z digest=sha256:04728efc8b3100eff57939cdec64cdddc87956f82781597b3e2b5e4c05d7c755

Observation 4a14de67-bbf9-4f72-af83-dd893038eb70 · outbound

This paper cites Bailando++: 3 D dance GPT with choreographic memory.

HumanCLAW: Can Vision-Language Models Act Through a Body? Bailando++: 3 D dance GPT with choreographic memory

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.333238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.195246Z digest=sha256:2a21240a801943d774009a8eca8cea0a757e4f98471ea062ee274a788810703f

Observation 0d3f88fb-894d-45f7-a2aa-3a0a19f60417 · outbound

This paper cites Duolando: Follower GPT with off-policy reinforcement learning for dance accompaniment.

HumanCLAW: Can Vision-Language Models Act Through a Body? Duolando: Follower GPT with off-policy reinforcement learning for dance accompaniment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.323138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.198233Z digest=sha256:1102dd76984b17e559a90e370aa9b7a590881e71b707b2eb13bd7dcfa101465b

Observation 6cd900b1-e2e8-4fcd-a52f-e75692678fb7 · outbound

This paper cites Half-Physics: Enabling Kinematic 3D Human Model with Physical Interactions.

HumanCLAW: Can Vision-Language Models Act Through a Body? Half-Physics: Enabling Kinematic 3D Human Model with Physical Interactions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.202268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.202268Z digest=sha256:7018d08bb4d8631e82ba40e2b0672bbdba073303445effd41857b02509e528e9

Observation 0a154a73-2625-4d63-96df-3574c6949044 · outbound

This paper cites Sadler, Wei-Lun Chao, and Yu Su.

HumanCLAW: Can Vision-Language Models Act Through a Body? Sadler, Wei-Lun Chao, and Yu Su

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.312655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.205693Z digest=sha256:8466674f8eafabcfa254fc4463b00a6b673079f8f7d787b99f4dc273eaf6c43a

Observation 21f1e0cb-c8f9-4e01-beee-5ebcf618b83c · outbound

This paper cites Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra.

HumanCLAW: Can Vision-Language Models Act Through a Body? Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.302877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.208697Z digest=sha256:45fbdd1f41351ef8611c2b03ef9c582b4b95f6731d012f274eeab4fde8f55dd1

Observation 9979aaa9-cf98-430c-9d65-8678ed2ff504 · outbound

This paper cites Cradle: Empowering Foundation Agents Towards General Computer Control.

HumanCLAW: Can Vision-Language Models Act Through a Body? Cradle: Empowering Foundation Agents Towards General Computer Control

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.212403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.212403Z digest=sha256:7e35da7823d0845cc4154ebda8acc8a8e865ac9c23671deaae9b5dea2237b9ad

Observation b70710c0-b998-42d1-b684-19c919fbbe9c · outbound

This paper cites an unresolved cited work.

HumanCLAW: Can Vision-Language Models Act Through a Body? Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:27:54.292224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.215673Z digest=sha256:21025308293933ca36e66d9fe38f58d998b651c0ac47a3705a03962c59d06d5d

Observation 1b7e799c-b6ee-494b-ac5a-b9ed8c0f6087 · outbound

This paper cites Voyager: An open-ended embodied agent with large language models.

HumanCLAW: Can Vision-Language Models Act Through a Body? Voyager: An open-ended embodied agent with large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.282231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.218692Z digest=sha256:18a25255c0b330a6c24b4721843f5fc622284918faa3843df6d9dfda866073b9

Observation 3c074f47-6e14-450c-bf47-633fb999e084 · outbound

This paper cites Describe, explain, plan and select: Interactive planning with LLMs enables open-world multi-task agents.

HumanCLAW: Can Vision-Language Models Act Through a Body? Describe, explain, plan and select: Interactive planning with LLMs enables open-world multi-task agents

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.272653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.221653Z digest=sha256:601fd62c2d5e4d4e5682e20429d7cf6039cea9f853eafac72497ca805afa64ad

Observation ad0f9f70-d0d0-453c-ba0c-88a6e417d331 · outbound

This paper cites Text2interact: High-fidelity and diverse text-to-two-person interaction generation.

HumanCLAW: Can Vision-Language Models Act Through a Body? Text2interact: High-fidelity and diverse text-to-two-person interaction generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.224699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.224699Z digest=sha256:a9c500cbb5d2c4bbd11a839e4c9c7b01b76e767bcfa604e5bfa4c6b0b6966d9a

Observation ffcc9334-c785-4714-ba36-c1197e77bbba · outbound

This paper cites OmniControl : Control any joint at any time for human motion generation.

HumanCLAW: Can Vision-Language Models Act Through a Body? OmniControl : Control any joint at any time for human motion generation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.262342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.228626Z digest=sha256:a9477fe648f2b5a8b136e440f2bbc65150b860f5d934db18e12b3c932c95a2b6

Observation 2e3d3914-2827-4669-838d-89d3bc8b3f73 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

HumanCLAW: Can Vision-Language Models Act Through a Body? Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.231578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.231578Z digest=sha256:7a1e0596f7a3c73501f577acb019c36e2f8afb9843d63fd11fe31309893033f4

Observation 8e90fffc-49ea-4a6e-9185-94db7348aa6b · outbound

This paper cites EmbodiedBench : Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents.

HumanCLAW: Can Vision-Language Models Act Through a Body? EmbodiedBench : Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.251796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.234858Z digest=sha256:7115c77ad20e74ee210382a34304f376ce9127f22066f14e87e403dc1496370b

Observation e3695817-2f10-44fd-987d-948aeb894aef · outbound

This paper cites PhysDiff : Physics-guided human motion diffusion model.

HumanCLAW: Can Vision-Language Models Act Through a Body? PhysDiff : Physics-guided human motion diffusion model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.237794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.237794Z digest=sha256:729be18fc6d329f4b4912907c61ca85de79cdd365da825cd56006ef386f1715b

Observation abee7c03-b433-43e5-be08-9769e454be1c · outbound

This paper cites VideoGameBench: Can Vision-Language Models complete popular video games?.

HumanCLAW: Can Vision-Language Models Act Through a Body? VideoGameBench: Can Vision-Language Models complete popular video games?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.240874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.240874Z digest=sha256:78b9ccc22d6a4b10f864f75072e6344c5ff736f2ea22cf6ca00d6684c4ccb2cb

Observation ec644601-6dbe-46d6-83ec-76a2517f540c · outbound

This paper cites Egoreact: Egocentric video-driven 3d human reaction generation.

HumanCLAW: Can Vision-Language Models Act Through a Body? Egoreact: Egocentric video-driven 3d human reaction generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:53.244099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:53.244099Z digest=sha256:83227bb511f5c9fb44ed221f8ec083653173a958c0a0f363c86c45af53287f7c

Observation 8832dcaf-2362-47ed-9ad6-5ffde4835f41 · outbound

This paper cites The wanderings of odysseus in 3d scenes.

HumanCLAW: Can Vision-Language Models Act Through a Body? The wanderings of odysseus in 3d scenes

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.241970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.247056Z digest=sha256:2fe6f4fbcb8515f5be18ced1ed6b4ffd1ccbc9fb11ac8f0e06e7e57f23cb1dcf

Observation 59f315d1-c39d-4469-8de1-0cafdf9c5f23 · outbound

This paper cites an unresolved cited work.

HumanCLAW: Can Vision-Language Models Act Through a Body? Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:27:54.231966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.249978Z digest=sha256:49f2eb72f5c1ccef5ecdffdf3e7a8256a652d095f6936955d53599904ffdc487

Observation e5da0000-ac44-4a54-998e-dc1069ad0619 · outbound

This paper cites Synthesizing diverse human motions in 3d indoor scenes.

HumanCLAW: Can Vision-Language Models Act Through a Body? Synthesizing diverse human motions in 3d indoor scenes

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.222112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.253873Z digest=sha256:20ed31a1008b2e7f7fb140dd022ee81a764a0d3e4186927ec4c81325e3c11539

Observation 81391ef6-c7d9-482e-8da4-8b14d2e44362 · outbound

This paper cites RT-2 : Vision-language-action models transfer web knowledge to robotic control.

HumanCLAW: Can Vision-Language Models Act Through a Body? RT-2 : Vision-language-action models transfer web knowledge to robotic control

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:54.211622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T04:27:53.257043Z digest=sha256:35d03550f34779b41c28d7016495370158d5af936cd91bea017dda56dc38f88a

Pith citing papers

No inbound Pith citation observations are available.