A single pretrained autoregressive VLM that generates chain-of-thought reasoning and robot action tokens in one stream matches or beats VLM-plus-flow-matching baselines across seven benchmarks.
Cot-vla: Visualchain-of-thoughtreasoningforvision-language-action models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
G0.5: One Autoregressive Stream for Robot Reasoning and Action
A single pretrained autoregressive VLM that generates chain-of-thought reasoning and robot action tokens in one stream matches or beats VLM-plus-flow-matching baselines across seven benchmarks.