Pith. sign in

Paper Citation Record · LEDGER

Average-Reward Soft Actor-Critic

As of 16 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2501.09080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09080 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:13:57.717789Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T08:41:52.119405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T08:45:19.217416Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact5
  • verified fuzzy6
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0146d23-1cdb-4608-9a76-fdb7bbe79c2a · outbound

This paper cites Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels.

Average-Reward Soft Actor-Critic Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.551979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.551979Z digest=sha256:83b116c8469569e322318fbf1f43c8972c2b5ca4653f2dcbf557397ba2fc3197

Observation f2a6e11c-6dc1-4c5a-b861-10be3ea7f69c · outbound

This paper cites Stochastic first-order methods for average-reward Markov decision processes.

Average-Reward Soft Actor-Critic Stochastic first-order methods for average-reward Markov decision processes

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:13:58.151359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:13:57.565005Z digest=sha256:56a210019508183e9276a07163b5ca442b769af5a8c3fa9a4b611a534a6e2d30

Observation 6c73f66b-967c-4f76-8a4d-f29c5bb2f753 · outbound

This paper cites Lillicrap, Jonathan J.

Average-Reward Soft Actor-Critic Lillicrap, Jonathan J

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.778193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:13:57.574195Z digest=sha256:b2486b30df239826dcaf68a9ec2c343ee6b2d42a954eaa42628c1ca5e5adfbb0

Observation 1b6fd44b-c44c-4368-894a-37436948a5c3 · outbound

This paper cites Discounted Reinforcement Learning Is Not an Optimization Problem.

Average-Reward Soft Actor-Critic Discounted Reinforcement Learning Is Not an Optimization Problem

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.580973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.580973Z digest=sha256:f400dcc9a09b72e37bb75b274184bd31b4f0217bd5aa575b59e398545dae20e9

Observation 66e598d2-f89e-43f0-9f97-8c10a4339abe · outbound

This paper cites Controllability-Aware Unsupervised Skill Discovery.

Average-Reward Soft Actor-Critic Controllability-Aware Unsupervised Skill Discovery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.602182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.602182Z digest=sha256:ea17f5231ddb5e95eca545ca5496a9ff27a99659f339a640c90d91abee6654a5

Observation 4054c8d6-6ce6-4337-a5a7-ceaf8b0314e0 · outbound

This paper cites Relative entropy and free energy dualities: Con- nections to path integral and kl control.

Average-Reward Soft Actor-Critic Relative entropy and free energy dualities: Con- nections to path integral and kl control

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.746542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:13:57.655247Z digest=sha256:a748c035969ff6a874bbf98d012cafc9006597dd040d681f0067a8bab9437b6b

Observation 2dcb0544-e426-4be8-bd63-58865362d663 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Average-Reward Soft Actor-Critic Behavior Regularized Offline Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.685681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.685681Z digest=sha256:091d6b8ebd022b7c24a2fa6fba0f420b6752e6efc44e6ecf3e5f08ae1e89f7df

Observation 992949d2-b9ae-474e-a18b-77742679f4f9 · outbound

This paper cites Efficient Reinforcement Learning with Large Language Model Priors.

Average-Reward Soft Actor-Critic Efficient Reinforcement Learning with Large Language Model Priors

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.693280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.693280Z digest=sha256:ff58c376009a89b938e85b700725806934c84efae501c97d530b3dceb222fe74

Observation da70aefb-c78e-47c8-8757-d9ca2eed2c29 · outbound

This paper cites Finite sample analysis of average-reward TD learning and Q-learning.

Average-Reward Soft Actor-Critic Finite sample analysis of average-reward TD learning and Q-learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.665375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:13:57.700341Z digest=sha256:b105e48f33e93c9d39c75d871498b17e1b5f41aba3ae7abe691db2fae937050f

Observation f5849538-2447-48e1-8bb2-cd64b643977e · outbound

This paper cites reward scale.

Average-Reward Soft Actor-Critic reward scale

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.613200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:13:57.717789Z digest=sha256:78755ae26326ae2d615acd6017b5faeaa8b3cd23f111b5ddd39f94874efb16d2

Observation 255b3875-7684-45df-b7fc-a09b7b8aa4b3 · outbound

This paper cites Reward Tweaking: Maximizing the Total Reward While Planning for Short Horizons.

Average-Reward Soft Actor-Critic Reward Tweaking: Maximizing the Total Reward While Planning for Short Horizons

Reference 1999

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:13:57.873267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:13:57.638508Z digest=sha256:0d2ca0210f31ba644629c3ea7fffb4cdfc040fd210a42cad47a66444eb2503b1

Observation 1d97c46c-cecf-417d-b851-c37836d391bf · outbound

This paper cites Dense dynamics-aware reward synthesis: Integrating prior experience with demonstrations.

Average-Reward Soft Actor-Critic Dense dynamics-aware reward synthesis: Integrating prior experience with demonstrations

Reference 2003

Resolution
verified exact
raw_fallback, observed 2026-08-10T20:13:58.392412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:13:57.536401Z digest=sha256:ea56a7f20c915e45c29022854a68806b2627d0850649de125b7f86d338b38ce8

Observation 2ce61fd6-8b07-45bb-aece-fa51b7471a3f · outbound

This paper cites Mujoco: A physics engine for model-based control.

Average-Reward Soft Actor-Critic Mujoco: A physics engine for model-based control

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.667433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.667433Z digest=sha256:8dfab008fb9028c767aa2f4359b47e6e2126dee2d9a1e4d0aec3fff97f69dcb3

Observation da48ec9a-95f8-4a39-876d-784452b9a9a8 · outbound

This paper cites ∞X k=1 r(st+k, at+k) − 1 β log π(at+k|st+k) π0(at+k|st+k) − θπ # , Qπ(st+1, at+1) = r(st+1, at+1) − θπ + E p,π.

Average-Reward Soft Actor-Critic ∞X k=1 r(st+k, at+k) − 1 β log π(at+k|st+k) π0(at+k|st+k) − θπ # , Qπ(st+1, at+1) = r(st+1, at+1) − θπ + E p,π

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.641196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:13:57.708275Z digest=sha256:83bb6c244b58361fe2fbf0534c8c9be1548f62550d2ca31ddf6a60688a612db0

Observation 711b2ad6-1711-4dfa-8ed3-b693e82bc4ae · outbound

This paper cites Proximal Policy Optimization Algorithms.

Average-Reward Soft Actor-Critic Proximal Policy Optimization Algorithms

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.618379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.618379Z digest=sha256:b76e671ff2feb9829e538e26ee00a0e1539f205e51bad22159214b62cc200ef1

Observation caafb751-ac41-488d-84a2-853866f4051f · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Average-Reward Soft Actor-Critic Soft Actor-Critic Algorithms and Applications

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.506033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.506033Z digest=sha256:d4a8ae80d68c38376a17e63a0858dc46befbae72d95405a2e5872fdbf6b8ca50

Observation b7e82671-4a27-4b1f-b938-ed127239bce1 · outbound

This paper cites RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning.

Average-Reward Soft Actor-Critic RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning

Reference 2018

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:13:58.447879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:13:57.513390Z digest=sha256:22f34cd321875520023ef61c52f1e985ea79817885087e75aa5b8115379917f7

Observation 94bd5da1-694f-4bd3-85f0-935f9e338448 · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

Average-Reward Soft Actor-Critic A unified view of entropy-regularized Markov decision processes

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.593356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.593356Z digest=sha256:6d75dfeae4a5df2afdcbbb574ce87cfa6b1bc98cff2dff5e06a3e78ec07759af

Observation 1d17c5b1-a7cd-4e07-989b-bc2d03398a94 · outbound

This paper cites Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review.

Average-Reward Soft Actor-Critic Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.558208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.558208Z digest=sha256:4fcce5d7b26a399aa955db72060b1e5e440030bc2a8017e96c1e5901ee5a12b1

Observation 71ced4dc-16f7-4a88-a920-be7660fe61c2 · outbound

This paper cites Dueling network architectures for deep reinforcement learning.

Average-Reward Soft Actor-Critic Dueling network architectures for deep reinforcement learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.676411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.676411Z digest=sha256:99adcf802a0b03ef03fee23572dbcc9418eaeb9d6a9215ef366a0ef884fe4573

Observation 97ac2ec5-a628-4b86-a794-32b9c58f7522 · outbound

This paper cites Making Reinforcement Learning Work on Swimmer.

Average-Reward Soft Actor-Critic Making Reinforcement Learning Work on Swimmer

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:13:58.551293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:13:57.499060Z digest=sha256:78486a0e15702bf4023e9cbe8377a7197e19e9fa241206ba6f942c9c8f42b10b

Observation b4e59e9f-f86a-4533-bde0-f99391d668fa · outbound

This paper cites Prioritized Experience Replay.

Average-Reward Soft Actor-Critic Prioritized Experience Replay

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.609321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.609321Z digest=sha256:33a18a6939f0cf2f8ec2b9505305253591c5b1c801643d0a285e43a9660df593

Observation db8c2468-c875-4db1-b9ea-441555c33241 · outbound

This paper cites The dependence of effective planning horizon on model accuracy.

Average-Reward Soft Actor-Critic The dependence of effective planning horizon on model accuracy

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.807736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:13:57.528380Z digest=sha256:025747770a3d1375bf5058be3e2f00d58727ddafa390e5718c92a9c59a0ec4e1

Observation 4d84dc87-68c2-43ee-b53f-4d0c58b22abf · outbound

This paper cites What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study.

Average-Reward Soft Actor-Critic What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.486495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.486495Z digest=sha256:c4b91b7babf99c7bd3eb4c65d3261f0c1844f0482f163d3f4444af417ba12732

Pith citing papers

Observation 2f999352-14a3-4199-b542-99f5bf01bf69 · inbound

Learning Adaptive Parameter Policies for Nonlinear Bayesian Filtering cites this paper.

Learning Adaptive Parameter Policies for Nonlinear Bayesian Filtering Average-Reward Soft Actor-Critic

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:45:19.219042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T08:41:52.119405Z digest=sha256:c3ebcf2b4bd0f7158158914681138d360825fe5500d9581356789d0ba63256c5

Observation 3fb8d06e-83b7-44b8-bd36-ffba1af6ba03 · inbound

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy cites this paper.

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy Average-Reward Soft Actor-Critic

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:52:22.146444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:50:13.653022Z digest=sha256:f9d3b2616eb3cc18ac4ce4d8b62f27f405947e75a8c76577f5258d4c2341671b