Pith. sign in

Look, focus, act: Efficient and robust robot learning via human gaze and foveated vision transform- ers

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

citation-role summary

baseline 1

citation-polarity summary

years

2026 4

roles

baseline 1

polarities

baseline 1

representative citing papers

Policy-based Foveated Imaging and Perception

cs.CV · 2026-06-01 · unverdicted · novelty 6.0

A task-aware policy learned via reinforcement learning allocates high-resolution pixels on dual-stream sensors in real time, outperforming fixed or non-predictive baselines under tight pixel budgets in both simulation and 200 MP hardware tests.

GazeVLA: Learning Human Intention for Robotic Manipulation

cs.RO · 2026-04-24 · unverdicted · novelty 6.0

GazeVLA pretrains on large human egocentric datasets to capture gaze-based intention, then finetunes on limited robot data with chain-of-thought reasoning to achieve better robotic manipulation performance than baselines.

HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations

cs.RO · 2026-03-03 · unverdicted · novelty 6.0

HoMMI learns whole-body mobile manipulation policies from robot-free human demonstrations by augmenting UMI with egocentric sensing and bridging the embodiment gap through an agnostic visual representation, relaxed head actions, and a whole-body controller.

citing papers explorer

Showing 4 of 4 citing papers.

  • Policy-based Foveated Imaging and Perception cs.CV · 2026-06-01 · unverdicted · none · ref 179 · internal anchor

    A task-aware policy learned via reinforcement learning allocates high-resolution pixels on dual-stream sensors in real time, outperforming fixed or non-predictive baselines under tight pixel budgets in both simulation and 200 MP hardware tests.

  • GazeVLA: Learning Human Intention for Robotic Manipulation cs.RO · 2026-04-24 · unverdicted · none · ref 19 · internal anchor

    GazeVLA pretrains on large human egocentric datasets to capture gaze-based intention, then finetunes on limited robot data with chain-of-thought reasoning to achieve better robotic manipulation performance than baselines.

  • Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games cs.LG · 2026-04-06 · conditional · none · ref 10 · internal anchor

    Across 20 Atari games, removing peripheral visual input drops human-action prediction accuracy by median 35–44%, far more than removing gaze maps (~2%) or past states (1.5–15%).

  • HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations cs.RO · 2026-03-03 · unverdicted · none · ref 8 · internal anchor

    HoMMI learns whole-body mobile manipulation policies from robot-free human demonstrations by augmenting UMI with egocentric sensing and bridging the embodiment gap through an agnostic visual representation, relaxed head actions, and a whole-body controller.