REVIEW 8 cited by
RRL: Resnet as representation for Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The ability to autonomously learn behaviors via direct interactions in uninstrumented environments can lead to generalist robots capable of enhancing productivity or providing care in unstructured settings like homes. Such uninstrumented settings warrant operations only using the robot's proprioceptive sensor such as onboard cameras, joint encoders, etc which can be challenging for policy learning owing to the high dimensionality and partial observability issues. We propose RRL: Resnet as representation for Reinforcement Learning -- a straightforward yet effective approach that can learn complex behaviors directly from proprioceptive inputs. RRL fuses features extracted from pre-trained Resnet into the standard reinforcement learning pipeline and delivers results comparable to learning directly from the state. In a simulated dexterous manipulation benchmark, where the state of the art methods fail to make significant progress, RRL delivers contact rich behaviors. The appeal of RRL lies in its simplicity in bringing together progress from the fields of Representation Learning, Imitation Learning, and Reinforcement Learning. Its effectiveness in learning behaviors directly from visual inputs with performance and sample efficiency matching learning directly from the state, even in complex high dimensional domains, is far from obvious.
Forward citations
Cited by 8 Pith papers
-
MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
Trained only on unlabeled human play videos, MimicDroid lets a GR1 humanoid perform new manipulation tasks from one to three demonstration videos, with roughly twice the real-world success of prior video-conditioned methods.
-
V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control
V-Simba, a visual RL architecture combining layer normalization, weight decay, and a distributional critic, matches or outperforms complex baselines on 29 continuous control tasks while using less compute.
-
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
UAD distills affordance knowledge from vision-language models and DINOv2 features into a lightweight task-conditioned model that predicts pixel-level manipulation regions and improves few-shot imitation learning gener...
-
MuST: Multi-Head Skill Transformer for Long-Horizon Dexterous Manipulation with Skill Progress
MuST adds per-skill action heads and a progress-guided skill selector to the Octo robot policy, improving long-horizon pick-and-pack success from about 32% to 90% in one simulated setting.
-
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
SAM2Act reports 86.8% average success across 18 RLBench tasks, and the memory variant SAM2Act+ reaches 94.3% on the new MemoryBench tasks.
-
When Should We Prefer State-to-Visual DAgger Over Visual Reinforcement Learning?
State-to-Visual DAgger outperforms visual RL on hard manipulation tasks and is more stable and faster in wall-clock time, but offers little sample-efficiency benefit on easy tasks.
-
4D Visual Pre-training for Robot Learning
A next-frame point-cloud diffusion pre-training method (FVP) improves DP3 and RDT-1B manipulation success rates on the paper's own tasks.
-
Efficient Reinforcement Learning Through Adaptively Pretrained Visual Encoder
APE pretrains a ResNet18 encoder with adaptively selected augmentations and freezes its early layers during policy learning, improving sample efficiency of DreamerV3 and DrQ-v2 on several visual RL benchmarks.
Discussion (0). Continue with ORCID to comment.