Pith. sign in

Paper Citation Record · LEDGER

Natural Policy Gradient for Average Reward Non-Stationary RL

As of 23 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2504.16415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16415 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:04.170613Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy44
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 653de7a0-e14a-4d30-a844-79114d016b31 · outbound

This paper cites write newline.

Natural Policy Gradient for Average Reward Non-Stationary RL write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:03.948118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:03.948118Z digest=sha256:bf0a58494895ce6b2b0b8824564037b50fa174ff09aa797942a9a7d65749d5a2

Observation ca2878ab-e549-467a-8b06-6bd76f9bb7c8 · outbound

This paper cites M., Lee, J.

Natural Policy Gradient for Average Reward Non-Stationary RL M., Lee, J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.907969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.953166Z digest=sha256:0fe5698f471d5913cb83333165bcbfde835ea60db2c92f6ae4eee48c1fd234fb

Observation 93917d66-04eb-49ce-8e98-a9e3f8e805b4 · outbound

This paper cites U., and Aggarwal, V.

Natural Policy Gradient for Average Reward Non-Stationary RL U., and Aggarwal, V

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.896660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.957028Z digest=sha256:3a24eafe1a34633b351f7c079f9727dc51054d70d113ced9f5007a8403d81f41

Observation c4622e78-580e-4ef6-bc66-5a57b67e4a53 · outbound

This paper cites First-order methods in optimization.

Natural Policy Gradient for Average Reward Non-Stationary RL First-order methods in optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:03.961027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:03.961027Z digest=sha256:0bcff7ae7cfe54b3a320ac64755e8ca845cc70c61e9da44321dfc25845f58dc7

Observation b8b4b070-4b30-47ad-8c5e-97c0be17014c · outbound

This paper cites Stochastic multi-armed-bandit problem with non-stationary rewards.

Natural Policy Gradient for Average Reward Non-Stationary RL Stochastic multi-armed-bandit problem with non-stationary rewards

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.879255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.964686Z digest=sha256:7d6c8b1ebf3f22be12de4c4dbd2b74b5dede8534ad86c77c1cdde178f197127d

Observation dcaceccf-a89b-488c-aa53-53760d493ceb · outbound

This paper cites S., Ghavamzadeh, M., and Lee, M.

Natural Policy Gradient for Average Reward Non-Stationary RL S., Ghavamzadeh, M., and Lee, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.868979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.968267Z digest=sha256:f184093200aa4f5c781a1eb149d68f1c163aed719f5b414c3a12ab0ebe753e13

Observation e11a1424-fac8-4ed4-af34-327cebf3369a · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Natural Policy Gradient for Average Reward Non-Stationary RL Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.856558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.972827Z digest=sha256:3b66055044b037605921b5e2bb7182a8c577a374a0a2185eb36a9e7e44e7cfc4

Observation 94b9d0e9-78b0-4e12-881a-09704368f686 · outbound

This paper cites Fast global convergence of natural policy gradient methods with entropy regularization.

Natural Policy Gradient for Average Reward Non-Stationary RL Fast global convergence of natural policy gradient methods with entropy regularization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.843310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.977907Z digest=sha256:54d1618ac1c163069a31b3fdd6a81f9a27164bb8d4ac1be56736ee3ff27786f6

Observation b8489771-fe9f-4d25-94fe-812dc81e94fa · outbound

This paper cites Optimizing for the future in non-stationary mdps.

Natural Policy Gradient for Average Reward Non-Stationary RL Optimizing for the future in non-stationary mdps

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.830462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.982030Z digest=sha256:54f30ef92afb2e2247dfd3903c6dbdc4d995baa8ae9857d41b18367464626de7

Observation 1676c12f-fb0f-4024-b1f2-2becc8f34351 · outbound

This paper cites Stabilizing reinforcement learning in dynamic environment with application to online recommendation.

Natural Policy Gradient for Average Reward Non-Stationary RL Stabilizing reinforcement learning in dynamic environment with application to online recommendation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.817636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.985559Z digest=sha256:037ed4c4f469a3ad66e96df86e746c8788c617688a3cbdbeb8081d2edaa45108

Observation 4ac1d03a-6728-4324-918d-4e140bdeb12d · outbound

This paper cites and Zhao, L.

Natural Policy Gradient for Average Reward Non-Stationary RL and Zhao, L

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.805055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.989285Z digest=sha256:4e125a4f33e5ec56fa4558d0b6264747cae76f156a99ced75fec2d772812ac99

Observation 3f1ef721-8f11-483f-853b-a48936e43dce · outbound

This paper cites C., Simchi-Levi, D., and Zhu, R.

Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.793923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.993078Z digest=sha256:5794aa1eb4e38987248431e0043b0bb180ef6a6bd9e65bc8541d7c2a729688ab

Observation 85046599-b60b-49c4-ba6e-2e3e49c817e6 · outbound

This paper cites C., Simchi-Levi, D., and Zhu, R.

Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.782754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:03.996510Z digest=sha256:77429901c9ffe1e0ffa8c2ab8ae8a0365ad875975b7ce2383ceea85c7c4d7951

Observation 24689d8c-b7d8-492a-821f-96c1c02e42a6 · outbound

This paper cites C., Simchi-Levi, D., and Zhu, R.

Natural Policy Gradient for Average Reward Non-Stationary RL C., Simchi-Levi, D., and Zhu, R

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.770734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.000031Z digest=sha256:071d0310fb72bbdc33f96ebe40825421078a2170c9cec82944b2edf8d47c8f40

Observation ef86e6ac-8325-4271-913e-b48613ffe31c · outbound

This paper cites A kernel-based approach to non-stationary reinforcement learning in metric spaces.

Natural Policy Gradient for Average Reward Non-Stationary RL A kernel-based approach to non-stationary reinforcement learning in metric spaces

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.757771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.003317Z digest=sha256:4cba6678988436b9ea1bae995d67e90a6e7517abe90b61d1fa952d07c2d8fa77

Observation 207f88a5-a57e-4d80-8411-061ad28e0c8c · outbound

This paper cites M., and Mansour, Y.

Natural Policy Gradient for Average Reward Non-Stationary RL M., and Mansour, Y

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.744455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.006825Z digest=sha256:a7efa125b99c560a445d77836e3f698fd1133e269f945e6313ff38e92011a98d

Observation c201b2fe-8135-476a-891f-74025763357b · outbound

This paper cites Dynamic regret of policy optimization in non-stationary environments.

Natural Policy Gradient for Average Reward Non-Stationary RL Dynamic regret of policy optimization in non-stationary environments

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.732565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.010119Z digest=sha256:709127bc2e753947a16b022c9e500a2f5a1bab98a34955f95e4b10b2140ac407

Observation 18ed8316-7435-4fe5-968f-26ddfac355b2 · outbound

This paper cites Non-stationary reinforcement learning under general function approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Non-stationary reinforcement learning under general function approximation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.715426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.013406Z digest=sha256:f981ab8baf2ddc00c99a95c79e998c9932a034d8491caa2545fde1c4a727845b

Observation 2cd4f8d5-c7b2-4f18-9148-9aa8005410d1 · outbound

This paper cites A sliding-window algorithm for markov decision processes with arbitrarily changing rewards and transitions, 2018.

Natural Policy Gradient for Average Reward Non-Stationary RL A sliding-window algorithm for markov decision processes with arbitrarily changing rewards and transitions, 2018

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.702291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.016639Z digest=sha256:7b021469b16b5784feeab74889cbd091385b4de992469548f757328f1d9ca7c9

Observation a1f95f88-4cb7-4d29-bcea-79402c0638f2 · outbound

This paper cites On Upper-Confidence Bound Policies for Non-Stationary Bandit Problems.

Natural Policy Gradient for Average Reward Non-Stationary RL On Upper-Confidence Bound Policies for Non-Stationary Bandit Problems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.020144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.020144Z digest=sha256:d6308dd82317202a60bdff79a68bba6458671609ca5bec3c19e5de453d9d29a4

Observation 7444774b-c490-49f7-aaaf-9a767560eb1f · outbound

This paper cites Transient Non-Stationarity and Generalisation in Deep Reinforcement Learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Transient Non-Stationarity and Generalisation in Deep Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.023886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.023886Z digest=sha256:4394a541dced4ea38e8fa8a8a510b19aea6d4b639314c2b23fcdb8a0e85a19cf

Observation 80a783cc-88cd-4381-84ea-1d15b75b8b8c · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Near-optimal regret bounds for reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.689556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.027462Z digest=sha256:dce1e121d7630555f966a472559b9eef9fea3d0584b07d1c2bdeba845296eae9

Observation a950b67b-098c-4223-82f4-d8ccde3256a1 · outbound

This paper cites Efficient reinforcement learning for routing jobs in heterogeneous queueing systems.

Natural Policy Gradient for Average Reward Non-Stationary RL Efficient reinforcement learning for routing jobs in heterogeneous queueing systems

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.676827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.031248Z digest=sha256:a63625bbc6f56164a156f1d0552836b42c4d13712010c0cbf6bb862552a4914e

Observation 94ac8006-84f1-4ffe-89ca-0fdd2be74acf · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.664726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.034636Z digest=sha256:ef887a2cddd636e5d12ae440399067ca6ec62c3a977b919d25238f807493d8ff

Observation 6615cda7-1c04-45a2-b2d9-8ba0e1b443be · outbound

This paper cites and Qian, P.

Natural Policy Gradient for Average Reward Non-Stationary RL and Qian, P

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.651057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.038113Z digest=sha256:b34261888f5655440652ea40484b3e6b89fbae9d9ec6c09b66646db57a812150

Observation eeeb929f-f14a-4ae9-ae11-9f6b52eda3b4 · outbound

This paper cites Towards continual reinforcement learning: A review and perspectives.

Natural Policy Gradient for Average Reward Non-Stationary RL Towards continual reinforcement learning: A review and perspectives

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.041523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.041523Z digest=sha256:74928e08b9bafd7ea1818ce5c3fe06f4705d7a991080f368bedb2a54def9850f

Observation ec2f33a2-8e55-4d43-8873-7adaaf4d9deb · outbound

This paper cites R., Varma, S.

Natural Policy Gradient for Average Reward Non-Stationary RL R., Varma, S

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.628726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.045034Z digest=sha256:b45513aa1e867d0156e5e411e3c59d6f4c57ff8f2e6c5c00794edea93d038c3d

Observation f71f109a-8aa0-4ed3-a96a-09d665fa23b5 · outbound

This paper cites T., Romberg, J., and Maguluri, S.

Natural Policy Gradient for Average Reward Non-Stationary RL T., Romberg, J., and Maguluri, S

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.614699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.048824Z digest=sha256:c57c22962f58d9d6098e8edee6a3437133ae473dd031a32d8f807d3491d1af6a

Observation 2923654e-62a3-433f-a4e0-a7a9239c85be · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.601039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.052793Z digest=sha256:a8abc194c8dd1562b599cf778792bf9cdef9006ea3b2bb206f982149d105da67

Observation c939c320-c385-424b-a86d-8883a306e65f · outbound

This paper cites Improved regret bound and experience replay in regularized policy iteration.

Natural Policy Gradient for Average Reward Non-Stationary RL Improved regret bound and experience replay in regularized policy iteration

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.588694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.056203Z digest=sha256:7dd21f4fd34b8bb26ee4c8cc33bca88688ec832004d91d93dfd3aa519c38f619

Observation 3e2a53e9-02f4-47ac-a315-6d51d8646982 · outbound

This paper cites and Rachelson, E.

Natural Policy Gradient for Average Reward Non-Stationary RL and Rachelson, E

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.575504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.059581Z digest=sha256:bd52806b3829d588da6816cf653516d25a3abb165b88509c0e10c4d84573ae0e

Observation b373027a-0703-405c-88fe-b887980b8a9c · outbound

This paper cites Pausing policy learning in non-stationary reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Pausing policy learning in non-stationary reinforcement learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.562565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.062965Z digest=sha256:2a812e66e0278c64814e2a069fc6ccea7fa9737ad3a5c1fbecbf7ceff8a957da

Observation 1c993825-9e01-4eda-aa3b-696f9a48ed80 · outbound

This paper cites Rl-qn: A reinforcement learning framework for optimal control of queueing systems.

Natural Policy Gradient for Average Reward Non-Stationary RL Rl-qn: A reinforcement learning framework for optimal control of queueing systems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.551164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.066666Z digest=sha256:fa24174aae8b137f59bee1da346b7fbdd235026aeb3a7771817330abca281bfc

Observation 7bc2e48e-317b-48d7-82ac-eef77431d013 · outbound

This paper cites A Definition of Non-Stationary Bandits.

Natural Policy Gradient for Average Reward Non-Stationary RL A Definition of Non-Stationary Bandits

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.069961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.069961Z digest=sha256:78fd59ee96a9b39aa5e7db99216f35145220044bc7d6d577b9ed856be148f442

Observation 96129abb-a2ab-488b-9a3a-a86e9ec7b14e · outbound

This paper cites Nonstationary bandit learning via predictive sampling.

Natural Policy Gradient for Average Reward Non-Stationary RL Nonstationary bandit learning via predictive sampling

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.538438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.073909Z digest=sha256:aae3f97f172577a99097d2df31b226c5a7cc9161cf26b2332d0de427fecd05b3

Observation fa0791ca-66de-4294-b7d7-95b4148e8c84 · outbound

This paper cites Average reward reinforcement learning: Foundations, algorithms, and empirical results.

Natural Policy Gradient for Average Reward Non-Stationary RL Average reward reinforcement learning: Foundations, algorithms, and empirical results

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.527413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.078377Z digest=sha256:ae5f2ac23525b55ded658f991b219649d02cfb3573e384ce78cd6c8644856dca

Observation 3b38bb50-96f6-4444-9da8-efc50c204f34 · outbound

This paper cites Model-free nonstationary reinforcement learning: Near-optimal regret and applications in multiagent reinforcement learning and inventory control.

Natural Policy Gradient for Average Reward Non-Stationary RL Model-free nonstationary reinforcement learning: Near-optimal regret and applications in multiagent reinforcement learning and inventory control

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.514105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.082744Z digest=sha256:a4b7d058c55b2246a1e44661839796a4c7a0127a88602cd33410b4fe279dbbed

Observation 7c4e1aa2-4ae1-4cfd-8133-12f97bc87932 · outbound

This paper cites New insights and perspectives on the natural gradient method.

Natural Policy Gradient for Average Reward Non-Stationary RL New insights and perspectives on the natural gradient method

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.501063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.087368Z digest=sha256:5d7332e966af990c8ffe697d575ad19c8503e3d005c339084344df2688039db8

Observation f9fd8b3d-866e-48f3-aea2-a2e05a6eed6d · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.488647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.091564Z digest=sha256:7073a9b270162375fb43c5d266429a397aaf2358bb53355e0aee158a74bac000

Observation 64f089b1-aab5-4b53-a9fe-86759e30c5f5 · outbound

This paper cites and Srikant, R.

Natural Policy Gradient for Average Reward Non-Stationary RL and Srikant, R

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.477406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.095137Z digest=sha256:fdcc13787a7daf6f8bfc08c8d72c02c977800575fc8ff7c8a845cbebca35952d

Observation d8631203-b9e7-4f10-ae79-c455b2175960 · outbound

This paper cites Performance bounds for policy-based average reward reinforcement learning algorithms.

Natural Policy Gradient for Average Reward Non-Stationary RL Performance bounds for policy-based average reward reinforcement learning algorithms

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.465214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.098637Z digest=sha256:53fe75457494bab20a1d325b4455c0066e784e0273103dad3c7af6cd9e1bc359

Observation e13e4eda-4d78-4a9e-90b6-7f579e8d7dd8 · outbound

This paper cites Bridging the gap between value and policy based reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Bridging the gap between value and policy based reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.452784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.102009Z digest=sha256:f66c03f037616f9975c9cb30afd2373d84d203928488c5cd0a0d6ff01e8882a6

Observation 1c888d79-729b-481e-af21-e2e8af216279 · outbound

This paper cites Variational regret bounds for reinforcement learning.

Natural Policy Gradient for Average Reward Non-Stationary RL Variational regret bounds for reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.440305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.105527Z digest=sha256:f338de55446b77d0f3f1598edbb485613868ecf5036ac149cd53baf36bac4978

Observation 0c74fd38-3b86-46e4-b764-21e61e672db4 · outbound

This paper cites A survey of reinforcement learning algorithms for dynamically varying environments.

Natural Policy Gradient for Average Reward Non-Stationary RL A survey of reinforcement learning algorithms for dynamically varying environments

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.428052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.109453Z digest=sha256:84f9398a20ab803b27cd061907063b15bca1216a41b2b8a2c38fdaeafe8d642e

Observation 7cc4bef4-0572-4bc7-a6af-9ed68e9df8c2 · outbound

This paper cites and Papadimitriou, C.

Natural Policy Gradient for Average Reward Non-Stationary RL and Papadimitriou, C

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.413135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.113199Z digest=sha256:0de247058cf873e95a15e5f071faae77141733d73576718f3331504686999949

Observation 92027896-4a5b-4ed7-91e8-dc88be160dbe · outbound

This paper cites Reinforcement learning for humanoid robotics.

Natural Policy Gradient for Average Reward Non-Stationary RL Reinforcement learning for humanoid robotics

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.400798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.117009Z digest=sha256:b27c4f57892cc35d8ab8777112cf643bf2f15a63fdb3f14f435d04ad63546fb2

Observation de98ca8f-217c-48db-8cd9-5e3ca019b59e · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.120637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.120637Z digest=sha256:53e9c298d35132cbad528503057301720269f4db5a43678d7d28ba5dcaaee721

Observation 12bbe108-bada-4cc1-be69-f1a3ae9ba9e0 · outbound

This paper cites Taming Non-stationary Bandits: A Bayesian Approach.

Natural Policy Gradient for Average Reward Non-Stationary RL Taming Non-stationary Bandits: A Bayesian Approach

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.124835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.124835Z digest=sha256:46733140ef7f4dfaf01a7e4c14b91dd055202caaf1eb614ac7ae26c8d83eb024

Observation 0b1eae28-da1a-457a-badc-0404906eb8a7 · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.381015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.129502Z digest=sha256:f6a40d7e188a76d9243c97c743bec9baf0b00c37fb9db6324f23aa1c251a2a63

Observation 9f6fd08f-133a-43e5-8029-9d32cbd1cfcf · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

Natural Policy Gradient for Average Reward Non-Stationary RL S., McAllester, D., Singh, S., and Mansour, Y

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.369819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.133253Z digest=sha256:4a440f97af1db1eb841dcb4538230041705a945d3a9de0070b3a3a337c035c12

Observation 07f8b68d-865c-44b6-a5a9-c4a505f98284 · outbound

This paper cites Efficient Learning in Non-Stationary Linear Markov Decision Processes.

Natural Policy Gradient for Average Reward Non-Stationary RL Efficient Learning in Non-Stationary Linear Markov Decision Processes

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:16:04.223437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.136748Z digest=sha256:79f8867ffd2c3a1d116fc759d8785c1b517e396da352af44a32f4ef309ce3e57

Observation 609a6bef-86e7-4ab2-b128-633b6fc78cab · outbound

This paper cites Non-asymptotic analysis for single-loop ( N atural) actor-critic with compatible function approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Non-asymptotic analysis for single-loop ( N atural) actor-critic with compatible function approximation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.358298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.143651Z digest=sha256:dd58f95431db4c3a207ee7e80def8d032de561487254a851edc5d7f96257c7e6

Observation 6bbed65e-980d-4ddb-80f2-b7a6683b5676 · outbound

This paper cites and Luo, H.

Natural Policy Gradient for Average Reward Non-Stationary RL and Luo, H

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.345302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.147668Z digest=sha256:8b67c9b0f4d6ec1c8b13f630e3fb07f2b442212cca88a92450503e0121afbc15

Observation 38c9fcba-22e8-40c3-8b8f-b6b677a123cf · outbound

This paper cites F., Zhang, W., Xu, P., and Gu, Q.

Natural Policy Gradient for Average Reward Non-Stationary RL F., Zhang, W., Xu, P., and Gu, Q

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.332643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.151138Z digest=sha256:ab340edf14c6b381af745da43203f44e34d116c2385edc25c3419369a83a45a7

Observation 12087280-777f-4d97-8e43-0fc39d8cad9a · outbound

This paper cites M., Golmohammadi, A., Shi, Y., et al.

Natural Policy Gradient for Average Reward Non-Stationary RL M., Golmohammadi, A., Shi, Y., et al

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.320024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.155857Z digest=sha256:d844f45d536fb2b8fc147be1736f3504d1a380c5ac1988de81dbc3dc4c67c4d2

Observation 855a2975-1c4f-4ce8-97d8-4a99ce175245 · outbound

This paper cites Multi-agent reinforcement learning: A selective overview of theories and algorithms.

Natural Policy Gradient for Average Reward Non-Stationary RL Multi-agent reinforcement learning: A selective overview of theories and algorithms

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.306025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.159451Z digest=sha256:cfc84b268b7996ffc87bb4543fac5341307bf7fe88db0a40dbdf5dd24456d485

Observation fc320535-15a3-44f1-b4ee-59dfacc79af9 · outbound

This paper cites an unresolved cited work.

Natural Policy Gradient for Average Reward Non-Stationary RL Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:16:04.292073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.162985Z digest=sha256:54b988e6d48769c881199aef21c282f00122713c8f8d395abd3ac443922e5988

Observation 0f857b11-1235-487a-af37-095c7afe56d3 · outbound

This paper cites Nonstationary Reinforcement Learning with Linear Function Approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Nonstationary Reinforcement Learning with Linear Function Approximation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:04.166805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:16:04.166805Z digest=sha256:826f38c2f42bbfdc4e6764bb27c9d3a760920112a125f04d10a9d620d5883b32

Observation 01959a49-5c02-454e-9b45-0d99969c0422 · outbound

This paper cites Finite-sample analysis for sarsa with linear function approximation.

Natural Policy Gradient for Average Reward Non-Stationary RL Finite-sample analysis for sarsa with linear function approximation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:16:04.279920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:16:04.170613Z digest=sha256:4fe511b68b80081c8bd259f35df5c26edbeb1ae556c27498bf9ffcad4f19053f

Pith citing papers

No inbound Pith citation observations are available.