Pith. sign in

REVIEW 3 cited by

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.09268 v2 pith:DBBMOISH submitted 2025-02-13 cs.RO cs.LG

classification cs.ROcs.LG
keywords modelgevrmperturbationsexternalinternalrobotsignificantvisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the rapid development of embodied artificial intelligence, significant progress has been made in vision-language-action (VLA) models for general robot decision-making. However, the majority of existing VLAs fail to account for the inevitable external perturbations encountered during deployment. These perturbations introduce unforeseen state information to the VLA, resulting in inaccurate actions and consequently, a significant decline in generalization performance. The classic internal model control (IMC) principle demonstrates that a closed-loop system with an internal model that includes external input signals can accurately track the reference input and effectively offset the disturbance. We propose a novel closed-loop VLA method GEVRM that integrates the IMC principle to enhance the robustness of robot visual manipulation. The text-guided video generation model in GEVRM can generate highly expressive future visual planning goals. Simultaneously, we evaluate perturbations by simulating responses, which are called internal embeddings and optimized through prototype contrastive learning. This allows the model to implicitly infer and distinguish perturbations from the external environment. The proposed GEVRM achieves state-of-the-art performance on both standard and perturbed CALVIN benchmarks and shows significant improvements in realistic robot tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A 0.5B parameter latent world-action model trains end-to-end on a single GPU and reaches 90.48% average success on 50 RoboTwin 2.0 tasks with a language-free Visual Transition Token for task specification.

  2. Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    Projection-aware activation steering using logistic regression recovers honesty and compassion under malicious prompts while preserving coherence and benchmark performance better than uniform steering.

  3. Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation

    cs.RO 2025-08 reject novelty 5.0 of 10

    Long-VLA uses phase-aware camera masking and a phase token to let a single end-to-end diffusion VLA handle 10-step manipulation, reporting large gains on its own L-CALVIN benchmark and on real-world tasks.

Pith tools