Pith. sign in

Llamo: Scaling pretrained language models for unified motion understanding and generation with continuous autoregressive tokens

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CV 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

HumanCLAW: Can Vision-Language Models Act Through a Body?

cs.CV · 2026-07-29 · conditional · novelty 7.0

A new full-body benchmark shows that current VLMs can recognize targets but cannot reliably tell where their own body is, whether it arrived, or whether it collided; the best solves only 16.8% of episodes.

citing papers explorer

Showing 1 of 1 citing paper.

  • HumanCLAW: Can Vision-Language Models Act Through a Body? cs.CV · 2026-07-29 · conditional · none · ref 27

    A new full-body benchmark shows that current VLMs can recognize targets but cannot reliably tell where their own body is, whether it arrived, or whether it collided; the best solves only 16.8% of episodes.