Pith. sign in

REVIEW 13 cited by

HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.05273 v3 pith:5BWHFRB7 submitted 2024-09-12 cs.CV cs.AIcs.RO

classification cs.CVcs.AIcs.RO
keywords hirttaskscontrolmodelssuccessbackendsdynamicfeatures
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large Vision-Language-Action (VLA) models, leveraging powerful pre trained Vision-Language Models (VLMs) backends, have shown promise in robotic control due to their impressive generalization ability. However, the success comes at a cost. Their reliance on VLM backends with billions of parameters leads to high computational costs and inference latency, limiting the testing scenarios to mainly quasi-static tasks and hindering performance in dynamic tasks requiring rapid interactions. To address these limitations, this paper proposes HiRT, a Hierarchical Robot Transformer framework that enables flexible frequency and performance trade-off. HiRT keeps VLMs running at low frequencies to capture temporarily invariant features while enabling real-time interaction through a high-frequency vision-based policy guided by the slowly updated features. Experiment results in both simulation and real-world settings demonstrate significant improvements over baseline methods. Empirically, in static tasks, we double the control frequency and achieve comparable success rates. Additionally, on novel real-world dynamic ma nipulation tasks which are challenging for previous VLA models, HiRT improves the success rate from 48% to 75%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models

    cs.CV 2026-01 conditional novelty 7.0 of 10

    Using a simple action-token adapter, nine VLMs are compared as robot policy backbones, showing general VLM ability transfers poorly to control and the vision encoder is the key bottleneck.

  2. Prediction with Action: Visual Policy Learning via Joint Denoising Process

    cs.RO 2024-11 conditional novelty 7.0 of 10

    PAD jointly denoises future images and robot actions in a single diffusion transformer, using video co-training to improve multi-task imitation learning.

  3. Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection

    cs.RO 2026-08 conditional novelty 6.0 of 10

    EMS, a dual-system VLA framework with RL-trained switching, achieves near-large-model success on LIBERO at high effective command rate.

  4. Robix: A Unified Model for Robot Interaction, Reasoning and Planning

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A three-stage-trained VLM unifies robot planning and dialogue, and beats commercial VLMs on the authors' interactive-task benchmarks.

  5. SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    SwitchVLA trains a vision-language-action policy to handle mid-execution instruction changes by conditioning on contact state and a three-way behavior mode, using only existing single-task demonstrations.

  6. Reinforced Reasoning for Embodied Planning

    cs.AI 2025-05 conditional novelty 6.0 of 10

    An SFT-plus-GRPO recipe lifts a 7B VLM to 35.6 percent success on EB-ALFRED versus 22.0 for GPT-4o-mini and 33.7 for Qwen2.5-VL-72B, with smaller but consistent gains on unseen EB-Habitat.

  7. Hume: Introducing System-2 Thinking in Visual-Language-Action Model

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A dual-system vision-language-action model that improves robot control by ranking multiple sampled action chunks with a learned value function before fast execution.

  8. UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Combining multimodal understanding with future image prediction in one autoregressive model improves vision-language-action policy success rates in simulation and real-world manipulation.

  9. FSD-VLN: Fast-Slow Dual-System Modeling for Aerial Long-Horizon Vision-Language Navigation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    An asynchronous fast-slow dual-system with DiT action modeling and time-weighted loss doubles unseen aerial VLN success rates and halves decision latency in simulation.

  10. Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A dual-system framework with a structured subtask interface, event-balanced training, and inference harness enables VLM-guided long-horizon robotic manipulation, achieving 95.5% on LIBERO-Long and 65% on real-world ch...

  11. Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning

    cs.RO 2025-06 conditional novelty 5.0 of 10

    FiS-VLA embeds a diffusion-based action module into the final transformer blocks of a vision-language model, achieving 69% mean success on RLBench and a claimed 117.7 Hz control frequency.

  12. OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation

    cs.RO 2025-05 conditional novelty 5.0 of 10

    OpenHelix shows that a frozen vision-language model with a prompt-tuned token and an auxiliary action-prediction head beats full fine-tuning on CALVIN language generalization while training far fewer parameters.

  13. Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey

    cs.RO 2025-10 conditional novelty 4.0 of 10

    A survey that groups VLA efficiency techniques into four categories: model architecture, perception features, action generation, and training/inference strategies.

Pith tools