REVIEW 12 cited by
Universal Successor Features Approximators
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The ability of a reinforcement learning (RL) agent to learn about many reward functions at the same time has many potential benefits, such as the decomposition of complex tasks into simpler ones, the exchange of information between tasks, and the reuse of skills. We focus on one aspect in particular, namely the ability to generalise to unseen tasks. Parametric generalisation relies on the interpolation power of a function approximator that is given the task description as input; one of its most common form are universal value function approximators (UVFAs). Another way to generalise to new tasks is to exploit structure in the RL problem itself. Generalised policy improvement (GPI) combines solutions of previous tasks into a policy for the unseen task; this relies on instantaneous policy evaluation of old policies under the new reward function, which is made possible through successor features (SFs). Our proposed universal successor features approximators (USFAs) combine the advantages of all of these, namely the scalability of UVFAs, the instant inference of SFs, and the strong generalisation of GPI. We discuss the challenges involved in training a USFA, its generalisation properties and demonstrate its practical benefits and transfer abilities on a large-scale domain in which the agent has to navigate in a first-person perspective three-dimensional environment.
Forward citations
Cited by 12 Pith papers
-
Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning
Independent per-agent successor-feature composition can be unsafe in cooperative teams; synchronized composition is safe but inflexible; the MA-USFA hierarchy claims both safety and flexibility.
-
Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines
Multitask Preplay replays experience from pursued tasks as starting points for counterfactual simulation of unpursued tasks to learn predictive representations that support fast generalization in humans and machines.
-
Goal-Conditioned Agents that Learn Everything All at Once
LEO enables efficient all-goals learning in goal-conditioned RL by jointly predicting for all goals in one network pass, yielding >250x speedup over relabelling and better performance on Craftax.
-
Robust Remote Reinforcement Learning over Unreliable Communication Channels using Homomorphic State Encoding
HR3L enables robust remote RL training over unreliable channels via homomorphic state encoding without gradient exchange, outperforming prior methods in sample efficiency and adapting to packet loss, delays, and bandw...
-
VUSFA:Variational Universal Successor Features Approximator to Improve Transfer DRL for Target Driven Visual Navigation
VUSFA combines universal successor features, a successor-feature-dependent policy, and a variational information bottleneck to improve target-driven visual navigation transfer in the AI2THOR simulator.
-
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.
-
When Dynamics Shift, Robust Task Inference Wins: Offline Imitation Learning with Behavior Foundation Models Revisited
Robust minimax task inference in BFMs achieves dynamics-shift robustness from nominal offline data alone and outperforms standard baselines.
-
Adaptive Policy Backbone via Shared Network
Adapting only linear layers before and after a frozen shared backbone is enough to transfer to out-of-distribution continuous-control tasks, with a theoretical argument and MuJoCo experiments.
-
Zero-Shot Reinforcement Learning Under Partial Observability
Behavior foundation models with GRU memory outperform memory-free zero-shot RL baselines in most partially observable ExORL settings, but the advantage is inconsistent on Cheetah.
-
Intention-Conditioned Flow Occupancy Models
InFOM applies flow matching to model intention-conditioned occupancy measures for RL pre-training, reporting 1.8x median return gains and 36% higher success rates on benchmarks.
-
Balancing Plasticity and Stability with Fast and Slow Successor Features
Synaptic consolidation applied to multi-timescale successor features yields better performance than plasticity-focused methods in RL under gradual environmental drift.
-
MAGIK: Mapping to Analogous Goals via Imagination-enabled Knowledge Transfer
MAGIK reuses a source RL policy for new analogous tasks by using a semi-supervised VAE to imagine target observations in source form, achieving zero-shot transfer in MiniGrid and Reacher.
Discussion (0). Continue with ORCID to comment.