Pith. sign in

REVIEW 3 cited by

Extreme Q-Learning: MaxEnt RL without Entropy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.02328 v2 pith:2QYOEIHD submitted 2023-01-05 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords entropyextremeonlineq-learningactionsalgorithmsdirectlyemph
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern Deep Reinforcement Learning (RL) algorithms require estimates of the maximal Q-value, which are difficult to compute in continuous domains with an infinite number of possible actions. In this work, we introduce a new update rule for online and offline RL which directly models the maximal value using Extreme Value Theory (EVT), drawing inspiration from economics. By doing so, we avoid computing Q-values using out-of-distribution actions which is often a substantial source of error. Our key insight is to introduce an objective that directly estimates the optimal soft-value functions (LogSumExp) in the maximum entropy RL setting without needing to sample from a policy. Using EVT, we derive our \emph{Extreme Q-Learning} framework and consequently online and, for the first time, offline MaxEnt Q-learning algorithms, that do not explicitly require access to a policy or its entropy. Our method obtains consistently strong performance in the D4RL benchmark, outperforming prior works by \emph{10+ points} on the challenging Franka Kitchen tasks while offering moderate improvements over SAC and TD3 on online DM Control tasks. Visualizations and code can be found on our website at https://div99.github.io/XQL/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion

    cs.RO 2025-06 conditional novelty 6.0 of 10

    DoublyAware combines conformal trajectory filtering with a group-relative policy constraint to improve sample efficiency of TD-MPC for simulated humanoid locomotion.

  2. Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood

    cs.LG 2025-06 conditional novelty 5.0 of 10

    SQOG adds a noise-based smoothing loss that pulls out-of-distribution action values toward neighboring in-sample values, improving Q-estimation and offline RL performance.

  3. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools