REVIEW 3 cited by
Align-RUDDER: Learning From Few Demonstrations by Reward Redistribution
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Reinforcement learning algorithms require many samples when solving complex hierarchical tasks with sparse and delayed rewards. For such complex tasks, the recently proposed RUDDER uses reward redistribution to leverage steps in the Q-function that are associated with accomplishing sub-tasks. However, often only few episodes with high rewards are available as demonstrations since current exploration strategies cannot discover them in reasonable time. In this work, we introduce Align-RUDDER, which utilizes a profile model for reward redistribution that is obtained from multiple sequence alignment of demonstrations. Consequently, Align-RUDDER employs reward redistribution effectively and, thereby, drastically improves learning on few demonstrations. Align-RUDDER outperforms competitors on complex artificial tasks with delayed rewards and few demonstrations. On the Minecraft ObtainDiamond task, Align-RUDDER is able to mine a diamond, though not frequently. Code is available at https://github.com/ml-jku/align-rudder. YouTube: https://youtu.be/HO-_8ZUl-UY
Forward citations
Cited by 3 Pith papers
-
Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning
Temporal GRPO splits a robot rollout into detectable task stages and applies separate group-relative policy advantages to each stage's action interval, improving success rates by 7 points on average over matched baselines.
-
Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning
LLM-generated code produces compact multi-dimensional latent rewards that improve temporal and multi-agent credit assignment in episodic reinforcement learning, outperforming standard return-decomposition baselines an...
-
Agent-Temporal Credit Assignment for Optimal Policy Preservation in Sparse Multi-Agent Reinforcement Learning
TAR2 redistributes sparse multi-agent rewards both across time and across agents, but its optimal-policy-preservation proof depends on a trajectory-dependent 'potential' and is not valid.
Discussion (0). Continue with ORCID to comment.