Pith. sign in

Paper Citation Record · LEDGER

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs

As of 18 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2504.11997.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.11997 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:42:09.234480Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf20c210-1e8f-4150-a5b6-21453d6537d1 · outbound

This paper cites Improved algorithms for linear stochastic bandits.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Improved algorithms for linear stochastic bandits

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.620630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.136777Z digest=sha256:bac5674a4635102faafd3ec6cfb631986aeddb5a20467575a51897e42196e6ee

Observation 5f0c9a4d-d416-4a2a-9d54-0666df91980e · outbound

This paper cites Near-opt imal regret bounds for reinforcement learn- ing.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Near-opt imal regret bounds for reinforcement learn- ing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.607023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.141280Z digest=sha256:38f4df5d03ba80b67198660b1353dfc218b8c16fdb41fc08b5943fe84bd6e4ab

Observation 1ea6fd89-cdc6-48f4-8e41-2abc1102bf37 · outbound

This paper cites Model-based reinforcement learning with value-targeted regression.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Model-based reinforcement learning with value-targeted regression

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.593152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.145088Z digest=sha256:875754840f2031df19931ee862a55329f88ff803c0c4dc86efb66ee966e163d9

Observation 002223ee-3be3-4914-972c-8bb3ae082ce9 · outbound

This paper cites REGAL: a regularizati on based algorithm for reinforcement learn- ing in weakly communicating MDPs.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs REGAL: a regularizati on based algorithm for reinforcement learn- ing in weakly communicating MDPs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.579572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.148984Z digest=sha256:e14c9511e7dfdeca20e2b185d914dcdb4b1dc4f843f4389b63a60a0f382de5d9

Observation 31c5fe95-e827-45d8-9aee-9d599ebdc74a · outbound

This paper cites Learning Infinite- Horizon Average-Reward Linear Mixture MDPs of Bounded Span.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Learning Infinite- Horizon Average-Reward Linear Mixture MDPs of Bounded Span

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.567360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.153118Z digest=sha256:ff4164d6d6093d43a2a0613c93e8b10c66e3d7081e650173bdcc2ac9a25fc185

Observation 3c9707ba-a4c4-41bd-9b5b-39656b819037 · outbound

This paper cites Efficient bias-span- constrained exploration-exploitation in reinforcement l earning.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Efficient bias-span- constrained exploration-exploitation in reinforcement l earning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.554902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.157133Z digest=sha256:15b4999b0e1e1a09e367383270228dba448ae312021670561468e27486386f69

Observation d93434ac-c213-49e1-85fb-85fb567db066 · outbound

This paper cites Inven tory management in supply chains: a re- inforcement learning approach.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Inven tory management in supply chains: a re- inforcement learning approach

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.542022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.161662Z digest=sha256:647aead3ed60370cc35f74b0d2330c85a292c5bcd57089cb3b9bc731d1afa696

Observation 4fda9f0f-7d61-40fc-a096-0838a5d95c09 · outbound

This paper cites Can deep reinforce- ment learning improve inventory management? performance o n lost sales, dual-sourcing, and multi- echelon problems.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Can deep reinforce- ment learning improve inventory management? performance o n lost sales, dual-sourcing, and multi- echelon problems

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.529036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.165411Z digest=sha256:fd50e3dd4638d90cf01e8c26f0934836a36b28697cccdf8e5261e5a4659c6d77

Observation 98eaf7ce-6c45-4b91-965b-db82a068ab8d · outbound

This paper cites Reinforcement learning for long-run a verage cost.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Reinforcement learning for long-run a verage cost

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.515365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.168956Z digest=sha256:be1000ca28f8214dafccc1c44e1f13107e35660ef9e638a432b9c7be5b0bbebf

Observation a373189d-04ed-4aea-bb3f-17d2dc51226c · outbound

This paper cites Sample-effi cient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Sample-effi cient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.502455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.172453Z digest=sha256:67fd89d5bcd3fe781f613d194fcea1a4ff2068ab0f20a6f71cdbc430df2cba2d

Observation 24fdec2c-618e-4828-b6ce-4d8d17d54ab2 · outbound

This paper cites Reinforcement Learn- ing for Infinite-Horizon Average-Reward Linear MDPs via App roximation by Discounted-Reward MDPs.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Reinforcement Learn- ing for Infinite-Horizon Average-Reward Linear MDPs via App roximation by Discounted-Reward MDPs

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.489896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.175983Z digest=sha256:736c4668b0a5bc16f46fa94c911feb8c03014021c0f02cf935b653df25a511a7

Observation e391adbc-7cf0-4878-8159-94883c14b200 · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Provably efficient reinforcement learning with linear function approximation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.475838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.179550Z digest=sha256:b105679890f4a7253929e157fc294423e8ebc39a521f385182006734b64363de

Observation 657a964b-9e31-40c9-b403-68f9bd3a29ce · outbound

This paper cites Towards tight bounds on th e sample complexity of average-reward MDPs.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Towards tight bounds on th e sample complexity of average-reward MDPs

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.461767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.183813Z digest=sha256:38939fc8c1398d6a514b1967f4df56deef6119f222e21026fd116ca601e1618b

Observation cedffd33-f0b3-448f-8bf8-2e0fd7300a92 · outbound

This paper cites Reinforcement learning based routin g in networks: Review and classification of approaches.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Reinforcement learning based routin g in networks: Review and classification of approaches

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.446829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.187438Z digest=sha256:e4026a021ec18c9167e6a7c2b8a7258e3f263fed0b2500da1bd2efd2ebfa921d

Observation 0c11398c-0252-4f53-9d54-307eef909972 · outbound

This paper cites Sample complexity of reinforcement learning using linearly combined model ensembles.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Sample complexity of reinforcement learning using linearly combined model ensembles

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.431697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.191051Z digest=sha256:b67796ad272c463c89a1c49976b9c8586f50350304ef91a962e579c1531f9381

Observation 1dbf06e3-abd3-41d1-b432-3413b6dc6477 · outbound

This paper cites Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:42:09.194938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:42:09.194938Z digest=sha256:7cf1697e6b74d96aaa68fa2f0a07ca29cc65edbce2cf56aff491ed6fdb2f05f4

Observation 217f2d16-2c1e-46c1-8c2e-141cdeb8a9cc · outbound

This paper cites Optimal Sample Complexity for Average Reward Markov Decision Processes.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Optimal Sample Complexity for Average Reward Markov Decision Processes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:42:09.198958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:42:09.198958Z digest=sha256:9c49f16c6a9debafa4aac5d9d1e1c639023d05bd54f6253d89e3285919424b3f

Observation 67fd2d37-9401-4a37-83f0-86ede1737eaa · outbound

This paper cites Learning infinite-horizon average-reward mdps with linear function approximation.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Learning infinite-horizon average-reward mdps with linear function approximation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.417503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.203038Z digest=sha256:222e1d1c820f08792c5a031313925693b8c90a2ef3acfd94a1905fbfb402deeb

Observation 53a1d58c-2ca8-4d3b-b05c-fcaf71db6ffc · outbound

This paper cites Model-free reinforcement learning in infinite-horizon average-rewar d markov decision processes.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Model-free reinforcement learning in infinite-horizon average-rewar d markov decision processes

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.402404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.206780Z digest=sha256:9f749cf6ab99ec0ce57623ac45f748d2878860c7868c1181303cac2bf56909f3

Observation d2731aac-4d0a-458c-b9a5-3e15acf4d278 · outbound

This paper cites Nearly minimax o ptimal regret for learning infinite- horizon average-reward mdps with linear function approxim ation.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Nearly minimax o ptimal regret for learning infinite- horizon average-reward mdps with linear function approxim ation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.388691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.210522Z digest=sha256:85c9a15bd4a0a6a4bad1036d210338b0e5e24b0e9b69a4391a535e90508bda1a

Observation 0345c7d6-3486-466a-958a-fabedd69edf8 · outbound

This paper cites Joint optimiz ation of preventive maintenance and production scheduling for multi-state production systems based on reinforcement learning.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Joint optimiz ation of preventive maintenance and production scheduling for multi-state production systems based on reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.373977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.214369Z digest=sha256:efc68de9c2f123eb56ada5906b6fa6d71418532eba16cc61f004ed4db4629901

Observation 990b21e2-c469-466b-a0bc-df1453fd3a51 · outbound

This paper cites Regret minimization for reinforcement learning by evaluating the optimal bias function.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Regret minimization for reinforcement learning by evaluating the optimal bias function

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.360091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.217863Z digest=sha256:f06eb35db2bd6e59b4e7b9019e1d503369cd0e6ff7c506cb1c6315093ddd9eb0

Observation d57704ee-93b7-40f9-8cb2-ac9df73c03b4 · outbound

This paper cites Sharper Model-free Reinf orcement Learning for Average-reward Markov Decision Processes.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Sharper Model-free Reinf orcement Learning for Average-reward Markov Decision Processes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.346138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.221490Z digest=sha256:8db3e9e1834ed4c8219e6694362b0fb3251ceab4923afac148aea4c9549f0e13

Observation 486aa4c3-e495-49a0-8ec6-1b840927354f · outbound

This paper cites Span-Based Optimal Sample Complexity for Average Reward MDPs.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Span-Based Optimal Sample Complexity for Average Reward MDPs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:42:09.225245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:42:09.225245Z digest=sha256:afb5134452bc38215a75b8f5435d4488ccc9461c3cd895f9af3befc57adb9c92

Observation 5c22267d-d48d-4aa0-9489-d9849a5e0a23 · outbound

This paper cites If n is odd, we can take φ n = 0 and similar argument holds.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs If n is odd, we can take φ n = 0 and similar argument holds

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.331024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.229998Z digest=sha256:cd03238fcbe8459cdc347c094a42b841e382dc84b5cfe5f0fb97f45f3dbb7676

Observation f18c35c9-61d7-4e3d-be48-0c3248b617e2 · outbound

This paper cites By Lemma 6, we have for t≥ 4, Qt u(s, a)≤ r(s, a) + γ[P V t u+1](s, a) + 2β‖ϕ (s, a)‖Λ −1 t + 2(mt−3− mt).

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs By Lemma 6, we have for t≥ 4, Qt u(s, a)≤ r(s, a) + γ[P V t u+1](s, a) + 2β‖ϕ (s, a)‖Λ −1 t + 2(mt−3− mt)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:42:09.314883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:42:09.234480Z digest=sha256:f533d297f2ef005fc0cfc5dc03eb467517fb9ca090a10cfdc818c591a2c7edf9

Pith citing papers

No inbound Pith citation observations are available.