Pith. sign in

REVIEW 1 cited by

Privileged Sensing Scaffolds Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14853 v1 pith:UN6ATTY2 submitted 2024-05-23 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords privilegedagentsrobotscaffolderscaffoldingsensingsensorssensory
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We need to look at our shoelaces as we first learn to tie them but having mastered this skill, can do it from touch alone. We call this phenomenon "sensory scaffolding": observation streams that are not needed by a master might yet aid a novice learner. We consider such sensory scaffolding setups for training artificial agents. For example, a robot arm may need to be deployed with just a low-cost, robust, general-purpose camera; yet its performance may improve by having privileged training-time-only access to informative albeit expensive and unwieldy motion capture rigs or fragile tactile sensors. For these settings, we propose "Scaffolder", a reinforcement learning approach which effectively exploits privileged sensing in critics, world models, reward estimators, and other such auxiliary components that are only used at training time, to improve the target policy. For evaluating sensory scaffolding agents, we design a new "S3" suite of ten diverse simulated robotic tasks that explore a wide range of practical sensor setups. Agents must use privileged camera sensing to train blind hurdlers, privileged active visual perception to help robot arms overcome visual occlusions, privileged touch sensors to train robot hands, and more. Scaffolder easily outperforms relevant prior baselines and frequently performs comparably even to policies that have test-time access to the privileged sensors. Website: https://penn-pal-lab.github.io/scaffolder/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GoStop: Reinforcement Learning for Adaptive Temporal Aggregation in Event-Based Feature Tracking

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An RL agent that adaptively decides when to accumulate events and when to run tracking inference improves event-based feature tracking on a new dynamic benchmark, but the gains are less consistent on an existing benchmark.

Pith tools