B2FF pre-generates a milestone bank of familiar future states from the clean initial observation and uses a recoverability-aware selector to guide VLA policies back from deviations, raising average success rate from 56.3% to 74.0% on failure-injected LIBERO.
hub
Soft actor-critic for discrete action settings
11 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Minimum-flow GFlowNets on graphs encode optimal transport plans, with the learned policy recovering the optimal coupling between source and target distributions.
R2PS combines a proof that dynamic programming remains optimal under asynchronous evader moves, a belief preservation mechanism for partial observability, and integration into equilibrium policy generalization to produce real-time pursuer policies that zero-shot generalize to unseen graphs.
ReMAC applies pathwise estimators to retry objectives in continuous RL, reshaping gradients to increase policy entropy and matching SAC performance without explicit regularization.
Shows entropy coupling limits DSAC on discrete tasks and introduces a generalized actor-critic framework with m-step critics and novel entropy-regularized objectives that perform robustly on Atari.
FactorLibrary stores reusable subexpressions to help RL agents (especially PPO+MCTS top-down) find certified optimal arithmetic circuits for polynomials up to complexity 8 at 91.8% success rate.
Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.
Qreg+NWLU improves forgetting mitigation and knowledge transfer in value-based multi-cyclic CRL by using dynamic Q-value rehearsal and immediate regularization instead of waiting after the first task.
REPPO is an on-policy RL method that combines pathwise policy gradients with relative entropy constraints to achieve stable training and high sample efficiency without replay buffers.
Event-driven RL framework for semiconductor manufacturing control shows throughput and utilization gains in high-fidelity simulations under offline and online training.
MacroNav learns multi-scale navigation-centric representations through multi-task self-supervised learning and combines them with graph-based reinforcement learning for efficient action selection, reporting gains in success rate and path efficiency over prior methods.
citing papers explorer
-
Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection
B2FF pre-generates a milestone bank of familiar future states from the clean initial observation and uses a recoverability-aware selector to guide VLA policies back from deviations, raising average success rate from 56.3% to 74.0% on failure-injected LIBERO.
-
Your GFlowNet Secretly Learns an Optimal Transport Plan
Minimum-flow GFlowNets on graphs encode optimal transport plans, with the learned policy recovering the optimal coupling between source and target distributions.
-
R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial Observability
R2PS combines a proof that dynamic programming remains optimal under asynchronous evader moves, a belief preservation mechanism for partial observability, and integration into equilibrium policy generalization to produce real-time pursuer policies that zero-shot generalize to unseen graphs.
-
Retry Policy Gradients in Continuous Action Spaces
ReMAC applies pathwise estimators to retry objectives in continuous RL, reshaping gradients to increase policy entropy and matching SAC performance without explicit regularization.
-
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
Shows entropy coupling limits DSAC on discrete tasks and introduces a generalized actor-critic framework with m-step critics and novel entropy-regularized objectives that perform robustly on Atari.
-
FactorLibrary: From Polynomials to Circuits via Recursive Subgoals
FactorLibrary stores reusable subexpressions to help RL agents (especially PPO+MCTS top-down) find certified optimal arithmetic circuits for polynomials up to complexity 8 at 91.8% success rate.
-
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Survey mapping RL techniques onto LLM training and highlighting gaps in value-based, off-policy, and bootstrapping methods.
-
Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning
Qreg+NWLU improves forgetting mitigation and knowledge transfer in value-based multi-cyclic CRL by using dynamic Q-value rehearsal and immediate regularization instead of waiting after the first task.
-
Relative Entropy Pathwise Policy Optimization
REPPO is an on-policy RL method that combines pathwise policy gradients with relative entropy constraints to achieve stable training and high sample efficiency without replay buffers.
-
Event-Driven Reinforcement Learning Enables Long-Horizon Control in Semiconductor Fabrication
Event-driven RL framework for semiconductor manufacturing control shows throughput and utilization gains in high-fidelity simulations under offline and online training.
-
MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments
MacroNav learns multi-scale navigation-centric representations through multi-task self-supervised learning and combines them with graph-based reinforcement learning for efficient action selection, reporting gains in success rate and path efficiency over prior methods.