Pith. sign in

REVIEW 4 cited by

VIOLA: Imitation Learning for Vision-Based Manipulation with Object Proposal Priors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.11339 v2 pith:U3XEHKSP submitted 2022-10-20 cs.RO

classification cs.RO
keywords violalearningimitationobjectapproachmanipulationmodelobject-centric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We introduce VIOLA, an object-centric imitation learning approach to learning closed-loop visuomotor policies for robot manipulation. Our approach constructs object-centric representations based on general object proposals from a pre-trained vision model. VIOLA uses a transformer-based policy to reason over these representations and attend to the task-relevant visual factors for action prediction. Such object-based structural priors improve deep imitation learning algorithm's robustness against object variations and environmental perturbations. We quantitatively evaluate VIOLA in simulation and on real robots. VIOLA outperforms the state-of-the-art imitation learning methods by $45.8\%$ in success rate. It has also been deployed successfully on a physical robot to solve challenging long-horizon tasks, such as dining table arrangement and coffee making. More videos and model details can be found in supplementary material and the project website: https://ut-austin-rpl.github.io/VIOLA .

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 12 citations worldwide. Full citation record

  1. SAIL: Faster-than-Demonstration Execution of Imitation Learning Policies

    cs.RO 2025-06 conditional novelty 7.0 of 10

    A full-stack speed-adaptation system lets imitation-learned robot policies execute up to 3-4x faster than human demonstrations while preserving task success rates.

  2. Compositional Diffusion with Guided Search for Long-Horizon Planning

    cs.RO 2025-12 conditional novelty 6.0 of 10

    CDGS adds population-based search and likelihood-based pruning to compositional diffusion, enabling long-horizon planning from short-horizon models across robot manipulation, panoramas, and video.

  3. RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A generative model and wrist camera turn human hand videos into robot gripper demonstrations that train manipulation policies at success rates close to those trained on real gripper data.

  4. PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    PSALM-V autonomously learns PDDL action preconditions and effects by executing LLM-proposed plans, predicting error messages, and refining a tree-structured belief, raising ALFRED plan success from 37% to 74% and indu...

Pith tools