← back to paper
arxiv: 2607.27180 · 2 revisions
HumanCLAW: Can Vision-Language Models Act Through a Body?