REVIEW 16 cited by
Soft Actor-Critic for Discrete Action Settings
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Soft Actor-Critic for Discrete Action Settings
read the original abstract
Soft Actor-Critic is a state-of-the-art reinforcement learning algorithm for continuous action settings that is not applicable to discrete action settings. Many important settings involve discrete actions, however, and so here we derive an alternative version of the Soft Actor-Critic algorithm that is applicable to discrete action settings. We then show that, even without any hyperparameter tuning, it is competitive with the tuned model-free state-of-the-art on a selection of games from the Atari suite.
Forward citations
Cited by 16 Pith papers
-
ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
ACPO decomposes the joint policy gradient into per-agent terms allowing independent actor training that collectively forms a joint gradient step in CTDE-based MARL.
-
Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection
B2FF pre-generates a milestone bank of familiar future states from the clean initial observation and uses a recoverability-aware selector to guide VLA policies back from deviations, raising average success rate from 5...
-
Your GFlowNet Secretly Learns an Optimal Transport Plan
Minimum-flow GFlowNets on graphs encode optimal transport plans, with the learned policy recovering the optimal coupling between source and target distributions.
-
R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial Observability
R2PS combines a proof that dynamic programming remains optimal under asynchronous evader moves, a belief preservation mechanism for partial observability, and integration into equilibrium policy generalization to prod...
-
Practical Graph Optimisation and AI-Driven Models for Active Directory Security Hardening
New game-theoretic and optimization models for honeypot placement, temporal decoy placement, and human-in-the-loop edge removal on Active Directory attack graphs, with hardness proofs and scalable heuristics.
-
ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
ACPO claims an exact decentralized decomposition of the cooperative MARL joint policy gradient via serialized agent-chained beliefs, but a key proof step is invalid, so the central claim is not established.
-
Retry Policy Gradients in Continuous Action Spaces
ReMAC applies pathwise estimators to retry objectives in continuous RL, reshaping gradients to increase policy entropy and matching SAC performance without explicit regularization.
-
Emotion Entanglement and Bayesian Inference for Multi-Dimensional Emotion Understanding
Hybrid RL-MPC trained on the full hybrid action space parametrizes continuous MPC via discrete rollouts and a critic terminal cost, yielding near-MINLP F1 strategies with recursive feasibility under a structural assumption.
-
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
Shows entropy coupling limits DSAC on discrete tasks and introduces a generalized actor-critic framework with m-step critics and novel entropy-regularized objectives that perform robustly on Atari.
-
FactorLibrary: From Polynomials to Circuits via Recursive Subgoals
FactorLibrary stores reusable subexpressions to help RL agents (especially PPO+MCTS top-down) find certified optimal arithmetic circuits for polynomials up to complexity 8 at 91.8% success rate.
-
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.
-
Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning
Qreg+NWLU improves forgetting mitigation and knowledge transfer in value-based multi-cyclic CRL by using dynamic Q-value rehearsal and immediate regularization instead of waiting after the first task.
-
Relative Entropy Pathwise Policy Optimization
REPPO is an on-policy RL method that combines pathwise policy gradients with relative entropy constraints to achieve stable training and high sample efficiency without replay buffers.
-
Event-Driven Reinforcement Learning Enables Long-Horizon Control in Semiconductor Fabrication
Event-driven RL framework for semiconductor manufacturing control shows throughput and utilization gains in high-fidelity simulations under offline and online training.
-
MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments
MacroNav learns multi-scale navigation-centric representations through multi-task self-supervised learning and combines them with graph-based reinforcement learning for efficient action selection, reporting gains in s...
-
Deep Reinforcement Learning: From First Principles to Reasoning Models
A textbook survey of deep reinforcement learning, from Bellman foundations to DQN, PPO, MuZero, offline RL, and reasoning models, with UAV/SD-WAN examples throughout.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.