Pith. sign in

Paper Citation Record · LEDGER

Average Reward Reinforcement Learning for Wireless Radio Resource Management

As of 23 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2501.06700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06700 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:58:03.027593Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe418e7b-8936-457b-be7e-cf0d22d8cbb8 · outbound

This paper cites Reinforcement learning: An introduction by richards’ sutton,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Reinforcement learning: An introduction by richards’ sutton,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.622682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.867946Z digest=sha256:3bf7f9fedd3be104b949527d5fc939a943eb0431cbb2d582cab73a3a2cabc013

Observation 63e51f31-3ce3-41dd-a0f5-29f0c4a04cee · outbound

This paper cites Average-reward off- policy policy evaluation with function approximation,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Average-reward off- policy policy evaluation with function approximation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.606153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.873578Z digest=sha256:093dd2952936f9e29ea9a4130142ac94dd237aea55b3acc1f73958bf26f70606

Observation 9be1150f-6663-4a75-9b9c-539ee364bde6 · outbound

This paper cites Off-policy average reward actor-critic with deterministic policy search,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Off-policy average reward actor-critic with deterministic policy search,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.589049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.878783Z digest=sha256:d83e1409fe088b8d32d627e525fe7c84bfc9aa53f90ad9fcd27d4005e26d679e

Observation 153af67c-6a51-4bce-a795-01b689aa1b11 · outbound

This paper cites Power control for a network of access points,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Power control for a network of access points,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.573092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.883865Z digest=sha256:96c304a41305ba675b0e5eb9e1f2c6737eb478966c7e252177550937278a9fc3

Observation 444c08a9-f2dd-430c-992a-1da342f5233d · outbound

This paper cites Base station employing shared resources among antenna units,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Base station employing shared resources among antenna units,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.556489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.888817Z digest=sha256:3c2641cf81ee813474829ad5b324383eea5207fa049ee1572e6df5ac8482f06f

Observation 832225f1-a78d-4760-b37c-37bcf3a5e033 · outbound

This paper cites Methods and apparatus for power management in a wireless communication system,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Methods and apparatus for power management in a wireless communication system,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.538941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.893915Z digest=sha256:5405039a99803a5046add7ed8e1d46f664fadd6160f9ef9e8f7bd9cb5c7204ed

Observation 9b85fed7-2f67-4e5f-a05b-86b4a65ceb9e · outbound

This paper cites Generalized global bandit and its application in cellular coverage optimization,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Generalized global bandit and its application in cellular coverage optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.522036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.899554Z digest=sha256:b62e3819eba9ac35a56436fceace1cacaf22af9c73b087207643fd3138f2f4ff

Observation 3048fc99-bf64-4f0d-87f0-7383e2230e45 · outbound

This paper cites A non-stationary online learning approach to mobility management,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management A non-stationary online learning approach to mobility management,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.505499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.904069Z digest=sha256:7c706ed618b28447f7a2dc9b86b706f9f6988c1a799cbd920c483842d306c508

Observation 8b685e6b-9a91-481d-822e-971cf0f84b39 · outbound

This paper cites A Deep Q-Learning Method for Downlink Power Allocation in Multi-Cell Networks.

Average Reward Reinforcement Learning for Wireless Radio Resource Management A Deep Q-Learning Method for Downlink Power Allocation in Multi-Cell Networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:02.908832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:02.908832Z digest=sha256:a58678e0c5f6923a03f0c8e31f5b40ffbf4aa85cc5bfdd3c7bac6c1ef03ae73a

Observation 8f0986a6-d45d-476f-b06d-17ed7a5f424d · outbound

This paper cites Power allocation in multi-user cellular networks with deep Q learning approach,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Power allocation in multi-user cellular networks with deep Q learning approach,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.487857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.914367Z digest=sha256:fc5a866cc1237dc0cceee66eb7db60127c6766c2d293db81dfdbf15b62c6aec5

Observation d86cf87a-3445-470b-a252-e99e5069f05f · outbound

This paper cites Joint power control and channel allocation for interference mitigation based on reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Joint power control and channel allocation for interference mitigation based on reinforcement learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.469648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.919206Z digest=sha256:41efdce2b55c03fbc7a001e13a4dce2803a1e651ea59976d7ef917303d538fe3

Observation e896803b-aba4-4573-b966-ed5796e76b39 · outbound

This paper cites Deep actor-critic learning for distributed power control in wireless mobile networks,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Deep actor-critic learning for distributed power control in wireless mobile networks,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.447250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.923840Z digest=sha256:b160ed919804157de0e5c96a6c2fb64ecd0ac31c7c180f4bd196fa6392a0741d

Observation a9e7afcb-a20d-4767-806d-f055c673c843 · outbound

This paper cites Deep reinforcement learning based wire- less network optimization: A comparative study,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Deep reinforcement learning based wire- less network optimization: A comparative study,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.429992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.928335Z digest=sha256:af268c165d5c6ca6237996eccf207ea911d0159c94a0660dff97545a4a2fa11c

Observation 4a8fb1ba-fd12-4558-8c98-551e9435361f · outbound

This paper cites ColO- RAN: Developing Machine Learning-based xApps for Open RAN Closed-loop Control on Programmable Experimental Platforms,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management ColO- RAN: Developing Machine Learning-based xApps for Open RAN Closed-loop Control on Programmable Experimental Platforms,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.412476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.933233Z digest=sha256:21b1bbd8121e5f576ba84b3641a9e641904a57df2fb68db72171dea07e7f7167

Observation 0efcb88b-4d34-4582-8df1-cf93c067722c · outbound

This paper cites FlexRAN: A flexible and programmable platform for software- defined radio access networks,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management FlexRAN: A flexible and programmable platform for software- defined radio access networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.396691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.938283Z digest=sha256:e2c1a7c76fa63ed4114d87364ca0b26fc1cc440a041da55ba31d91e5c95ed533

Observation 04d1d7ad-6a01-41cf-9b35-9849bb3fec9e · outbound

This paper cites Deep reinforcement learning for joint spectrum and power allocation in cellular networks,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Deep reinforcement learning for joint spectrum and power allocation in cellular networks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.376074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.943681Z digest=sha256:3b09fa2e39ffe64326aa445a96e4e2b65d4e97a82a8050b5d67588d5b9679d69

Observation 8c8a43ac-33d9-4d35-aeda-e0329a01963f · outbound

This paper cites Multi-agent deep reinforcement learning for dynamic power allocation in wireless networks,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Multi-agent deep reinforcement learning for dynamic power allocation in wireless networks,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.359596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.948386Z digest=sha256:c5f75cfccfe23a9c19d4bdd8ecef994211e28994bb570ed9772cb4ff59d71bd6

Observation abbb5159-d0a4-4f5a-a14c-376c7d667aa0 · outbound

This paper cites Multi- agent reinforcement learning for wireless user scheduling: Performance, scalablility, and generalization,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Multi- agent reinforcement learning for wireless user scheduling: Performance, scalablility, and generalization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.340416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.953229Z digest=sha256:4d0f288e84deb3487cbaacabb94f38a4b4e4a210c790c1f65597ab4b8f6534de

Observation 4786d838-3339-4a54-a55a-9d0ad12459ac · outbound

This paper cites Resource management in wireless networks via multi-agent deep reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Resource management in wireless networks via multi-agent deep reinforcement learning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:02.958639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:02.958639Z digest=sha256:21d291c8df23cdaf92898f0c349236cf5d78c837809e60e2d48e39884dd2ac63

Observation 9f6fe38c-60b8-4321-a8e6-407c26949b68 · outbound

This paper cites Distributed MARL for scheduling in conflict graphs,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Distributed MARL for scheduling in conflict graphs,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.311863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.963929Z digest=sha256:7901ea51a3b06f9e4f240dc0e502c08048e37bf6815f2e7d9a8e1d4c09207a79

Observation d2de51e2-2bdd-4ffa-9194-8accdc97981a · outbound

This paper cites Offline reinforcement learning for wireless network optimization with mixture datasets,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Offline reinforcement learning for wireless network optimization with mixture datasets,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.295227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.970809Z digest=sha256:01a9264e5a13c2eec5dc47b6e4c8025e2ce46309bd054a6ab555daa462a4998f

Observation 3daa3c11-e5c9-4b28-bc45-5af3768177b2 · outbound

This paper cites Advancing RAN slicing with offline reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Advancing RAN slicing with offline reinforcement learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.279124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.976751Z digest=sha256:25e51f3fb3df430f80ce9a5c04989c34ee4c0745f81108d715662c5086dae855

Observation 0fc23e0a-4c89-4b71-abaf-532fb6834614 · outbound

This paper cites Mean-variance policy iteration for risk-averse reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Mean-variance policy iteration for risk-averse reinforcement learning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.261064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.982057Z digest=sha256:a6517d4538c2c6e6cfe755cc81e9ed798c2d6f2a4269b4165634af433f33d304

Observation e0aed05d-01fa-49d7-8e9f-d6d13de8f23a · outbound

This paper cites Averaged-dqn: Variance reduc- tion and stabilization for deep reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Averaged-dqn: Variance reduc- tion and stabilization for deep reinforcement learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.243701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.987694Z digest=sha256:7ee9c2e726e313bace90a6b35951ce45a73668146f4985f947f2d6e72599d54e

Observation 38faea7d-0567-4d26-9138-308f0ede6ca3 · outbound

This paper cites Average-Reward Reinforcement Learning with Trust Region Methods.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Average-Reward Reinforcement Learning with Trust Region Methods

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:02.992888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:02.992888Z digest=sha256:fd38808ac273586f22a001cd52a5163870ce5dc37993003189addf800af86178

Observation 79263734-8de1-48bf-9c0d-79864f8f6157 · outbound

This paper cites The NS-3 network simulator,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management The NS-3 network simulator,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.224731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:02.997779Z digest=sha256:6b611bb81a0fe1045f484160d69b1ccdffb12740c6f9ea8c245e6ffc9f45c2ed

Observation bf85f0c0-a91e-49a3-b59b-75d2d16428ae · outbound

This paper cites Network slicing architecture,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Network slicing architecture,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.206772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:03.002821Z digest=sha256:018867080d063b637bb16925e28cbd01ee4545e92502d03e69a479d45ef1d06b

Observation a0c4bc3e-559f-45c2-8445-4840f43886d2 · outbound

This paper cites NetworkGym: Democratizing Network AI via Sim-aaS,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management NetworkGym: Democratizing Network AI via Sim-aaS,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.188710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:03.007765Z digest=sha256:55c3e51b619dd2395887f3bedcc422c8bfcf3731ebd3645ee7eb581301d26a2e

Observation 77cd05a6-0482-4fbd-8745-f0df02bb77af · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Soft Actor-Critic Algorithms and Applications

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:03.012480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:03.012480Z digest=sha256:3c61369af77f7e3c73fba881dd4ebfe1ac90859b0b9c333303088e064282d71e

Observation 75f746b8-3394-412e-9bdd-861ba0a638f8 · outbound

This paper cites A Deeper Look at Discounting Mismatch in Actor-Critic Algorithms.

Average Reward Reinforcement Learning for Wireless Radio Resource Management A Deeper Look at Discounting Mismatch in Actor-Critic Algorithms

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:58:03.078406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:03.017617Z digest=sha256:9b6aab5512cfb26e2632b0f09805367b2fd5b49e34131140a6000e8f597b8028

Observation 8a8d116c-0e34-452f-9949-38b760fd7858 · outbound

This paper cites Revisiting the minimalist approach to offline reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Revisiting the minimalist approach to offline reinforcement learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.171997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:03.022995Z digest=sha256:b1ba6a1601472c04c548fd4613f4b737ac5fbb66616f42f9930038afea4b0c35

Observation 9bc0867d-2531-42e9-b963-a953fdb3ed79 · outbound

This paper cites Supported policy optimization for offline reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Supported policy optimization for offline reinforcement learning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.154181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:58:03.027593Z digest=sha256:fcbf9aabe823312d574793fe9cd28a50932292a6c24c0f410cc9c65419d23f7f

Pith citing papers

No inbound Pith citation observations are available.