Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:13:57.717789Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2501.09080.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:13:57.717789Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-15T08:41:52.119405Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T08:45:19.217416Z
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f0146d23-1cdb-4608-9a76-fdb7bbe79c2a · outbound
Average-Reward Soft Actor-Critic Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2a6e11c-6dc1-4c5a-b861-10be3ea7f69c · outbound
Average-Reward Soft Actor-Critic Stochastic first-order methods for average-reward Markov decision processes
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6c73f66b-967c-4f76-8a4d-f29c5bb2f753 · outbound
Average-Reward Soft Actor-Critic Lillicrap, Jonathan J
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1b6fd44b-c44c-4368-894a-37436948a5c3 · outbound
Average-Reward Soft Actor-Critic Discounted Reinforcement Learning Is Not an Optimization Problem
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66e598d2-f89e-43f0-9f97-8c10a4339abe · outbound
Average-Reward Soft Actor-Critic Controllability-Aware Unsupervised Skill Discovery
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4054c8d6-6ce6-4337-a5a7-ceaf8b0314e0 · outbound
Average-Reward Soft Actor-Critic Relative entropy and free energy dualities: Con- nections to path integral and kl control
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2dcb0544-e426-4be8-bd63-58865362d663 · outbound
Average-Reward Soft Actor-Critic Behavior Regularized Offline Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 992949d2-b9ae-474e-a18b-77742679f4f9 · outbound
Average-Reward Soft Actor-Critic Efficient Reinforcement Learning with Large Language Model Priors
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da70aefb-c78e-47c8-8757-d9ca2eed2c29 · outbound
Average-Reward Soft Actor-Critic Finite sample analysis of average-reward TD learning and Q-learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f5849538-2447-48e1-8bb2-cd64b643977e · outbound
Average-Reward Soft Actor-Critic reward scale
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 255b3875-7684-45df-b7fc-a09b7b8aa4b3 · outbound
Average-Reward Soft Actor-Critic Reward Tweaking: Maximizing the Total Reward While Planning for Short Horizons
Reference 1999
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1d97c46c-cecf-417d-b851-c37836d391bf · outbound
Average-Reward Soft Actor-Critic Dense dynamics-aware reward synthesis: Integrating prior experience with demonstrations
Reference 2003
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2ce61fd6-8b07-45bb-aece-fa51b7471a3f · outbound
Average-Reward Soft Actor-Critic Mujoco: A physics engine for model-based control
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da48ec9a-95f8-4a39-876d-784452b9a9a8 · outbound
Average-Reward Soft Actor-Critic ∞X k=1 r(st+k, at+k) − 1 β log π(at+k|st+k) π0(at+k|st+k) − θπ # , Qπ(st+1, at+1) = r(st+1, at+1) − θπ + E p,π
Reference 2010
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 711b2ad6-1711-4dfa-8ed3-b693e82bc4ae · outbound
Average-Reward Soft Actor-Critic Proximal Policy Optimization Algorithms
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caafb751-ac41-488d-84a2-853866f4051f · outbound
Average-Reward Soft Actor-Critic Soft Actor-Critic Algorithms and Applications
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7e82671-4a27-4b1f-b938-ed127239bce1 · outbound
Average-Reward Soft Actor-Critic RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 94bd5da1-694f-4bd3-85f0-935f9e338448 · outbound
Average-Reward Soft Actor-Critic A unified view of entropy-regularized Markov decision processes
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d17c5b1-a7cd-4e07-989b-bc2d03398a94 · outbound
Average-Reward Soft Actor-Critic Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ced4dc-16f7-4a88-a920-be7660fe61c2 · outbound
Average-Reward Soft Actor-Critic Dueling network architectures for deep reinforcement learning
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97ac2ec5-a628-4b86-a794-32b9c58f7522 · outbound
Average-Reward Soft Actor-Critic Making Reinforcement Learning Work on Swimmer
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b4e59e9f-f86a-4533-bde0-f99391d668fa · outbound
Average-Reward Soft Actor-Critic Prioritized Experience Replay
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db8c2468-c875-4db1-b9ea-441555c33241 · outbound
Average-Reward Soft Actor-Critic The dependence of effective planning horizon on model accuracy
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4d84dc87-68c2-43ee-b53f-4d0c58b22abf · outbound
Average-Reward Soft Actor-Critic What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f999352-14a3-4199-b542-99f5bf01bf69 · inbound
Learning Adaptive Parameter Policies for Nonlinear Bayesian Filtering Average-Reward Soft Actor-Critic
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3fb8d06e-83b7-44b8-bd36-ffba1af6ba03 · inbound
When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy Average-Reward Soft Actor-Critic
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.