Pith. sign in

Cot-vla: Visual chain-of-thought reasoning for vision-language-action models

9 Pith papers cite this work. Polarity classification is still indexing.

9 Pith papers citing it

citation-role summary

background 1 baseline 1

citation-polarity summary

fields

cs.RO 5 cs.CV 4

years

2026 8 2025 1

representative citing papers

Native Video-Action Pretraining for Generalizable Robot Control

cs.RO · 2026-07-09 · conditional · novelty 5.0

A video-action foundation model pretrained natively with a causal diffusion transformer and semantic visual-action tokenizer reports improved few-shot robot manipulation and 225 Hz asynchronous closed-loop control.

QuoVLA: Quotient Space for Vision-Language-Action Models

cs.CV · 2026-05-24 · unverdicted · novelty 5.0

QuoVLA introduces a quotient-space framework that compresses VLM latents into action-sufficient representations via quantization and dual-branch design for better VLA generalization.

Causal World Modeling for Robot Control

cs.CV · 2026-01-29 · unverdicted · novelty 5.0

LingBot-VA combines video world modeling with policy learning via Mixture-of-Transformers, closed-loop rollouts, and asynchronous inference to improve robot manipulation in simulation and real settings.

citing papers explorer

Showing 9 of 9 citing papers.