Pith. sign in

Paper Citation Record · LEDGER

Average Reward Reinforcement Learning for Wireless Radio Resource Management

As of 16 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2501.06700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06700 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:58:03.027593Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe418e7b-8936-457b-be7e-cf0d22d8cbb8 · outbound

This paper cites Reinforcement learning: An introduction by richards’ sutton,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Reinforcement learning: An introduction by richards’ sutton,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.622682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.867946Z digest=sha256:b971d4bdf233e58f9baf9e0605010470974dc3a6f8c5044bba7a93484549a3a0

Observation 63e51f31-3ce3-41dd-a0f5-29f0c4a04cee · outbound

This paper cites Average-reward off- policy policy evaluation with function approximation,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Average-reward off- policy policy evaluation with function approximation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.606153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.873578Z digest=sha256:c3a4dd30eac2bb5958993882839da8ab00a76425eaab1dcbccc4a8b3c904b065

Observation 9be1150f-6663-4a75-9b9c-539ee364bde6 · outbound

This paper cites Off-policy average reward actor-critic with deterministic policy search,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Off-policy average reward actor-critic with deterministic policy search,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.589049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.878783Z digest=sha256:c097aba100216e662ded21ec67d22d694585c288e215944469cee8a2a6f6f313

Observation 153af67c-6a51-4bce-a795-01b689aa1b11 · outbound

This paper cites Power control for a network of access points,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Power control for a network of access points,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.573092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.883865Z digest=sha256:e48b9f78fe32fcda9538d8beefb23db60d1af82ab3283b96b01aa1066e552226

Observation 444c08a9-f2dd-430c-992a-1da342f5233d · outbound

This paper cites Base station employing shared resources among antenna units,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Base station employing shared resources among antenna units,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.556489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.888817Z digest=sha256:3d0bf05eeccc4537135ad5db51fc1111b14ecc25f72f3eff5be5e89d3e0ce6c4

Observation 832225f1-a78d-4760-b37c-37bcf3a5e033 · outbound

This paper cites Methods and apparatus for power management in a wireless communication system,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Methods and apparatus for power management in a wireless communication system,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.538941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.893915Z digest=sha256:d2064af4cbf9ef0c8278eba09e8aa3ba8b5d3350931ba823afb3a8357bdb878a

Observation 9b85fed7-2f67-4e5f-a05b-86b4a65ceb9e · outbound

This paper cites Generalized global bandit and its application in cellular coverage optimization,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Generalized global bandit and its application in cellular coverage optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.522036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.899554Z digest=sha256:45d4ec03e8723ac7de16e7edde40749d0b43ae83ef4c8fc80e82bc93e481ee5a

Observation 3048fc99-bf64-4f0d-87f0-7383e2230e45 · outbound

This paper cites A non-stationary online learning approach to mobility management,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management A non-stationary online learning approach to mobility management,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.505499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.904069Z digest=sha256:39df664555eb121ead3d615cd391a8a41e5d12208ed31d12f340ade5c7ea0f21

Observation 8b685e6b-9a91-481d-822e-971cf0f84b39 · outbound

This paper cites A Deep Q-Learning Method for Downlink Power Allocation in Multi-Cell Networks.

Average Reward Reinforcement Learning for Wireless Radio Resource Management A Deep Q-Learning Method for Downlink Power Allocation in Multi-Cell Networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:02.908832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:02.908832Z digest=sha256:3434557804daaf5bef0dd1f0b6abb4ce7b42a143005ea18355abf8bb468b2d51

Observation 8f0986a6-d45d-476f-b06d-17ed7a5f424d · outbound

This paper cites Power allocation in multi-user cellular networks with deep Q learning approach,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Power allocation in multi-user cellular networks with deep Q learning approach,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.487857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.914367Z digest=sha256:0154bf64866389c068527b4ede6ed289aa0f7655ebb1c96f410b272f1e3e080d

Observation d86cf87a-3445-470b-a252-e99e5069f05f · outbound

This paper cites Joint power control and channel allocation for interference mitigation based on reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Joint power control and channel allocation for interference mitigation based on reinforcement learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.469648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.919206Z digest=sha256:69b811df942d4e0198cca149676d757f3422246d22ea74eb6702384e53f60879

Observation e896803b-aba4-4573-b966-ed5796e76b39 · outbound

This paper cites Deep actor-critic learning for distributed power control in wireless mobile networks,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Deep actor-critic learning for distributed power control in wireless mobile networks,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.447250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.923840Z digest=sha256:6863d6df05afb2460ad94eacf8b843e7684ad1d0c0662b9425a1e369291d1232

Observation a9e7afcb-a20d-4767-806d-f055c673c843 · outbound

This paper cites Deep reinforcement learning based wire- less network optimization: A comparative study,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Deep reinforcement learning based wire- less network optimization: A comparative study,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.429992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.928335Z digest=sha256:67c9ff8d929063ae14adbebe425e1032b37c0f9a4a97493676a014f4280bb4a7

Observation 4a8fb1ba-fd12-4558-8c98-551e9435361f · outbound

This paper cites ColO- RAN: Developing Machine Learning-based xApps for Open RAN Closed-loop Control on Programmable Experimental Platforms,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management ColO- RAN: Developing Machine Learning-based xApps for Open RAN Closed-loop Control on Programmable Experimental Platforms,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.412476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.933233Z digest=sha256:8be31e862d4447ea2e791b4ec19d93cb05be406f6e44c3a3a263c355c8fb7d5e

Observation 0efcb88b-4d34-4582-8df1-cf93c067722c · outbound

This paper cites FlexRAN: A flexible and programmable platform for software- defined radio access networks,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management FlexRAN: A flexible and programmable platform for software- defined radio access networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.396691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.938283Z digest=sha256:88e64d6532673099afc3426cd405dd8385ab9d60544d6bda5c6b3dc687c9f9c5

Observation 04d1d7ad-6a01-41cf-9b35-9849bb3fec9e · outbound

This paper cites Deep reinforcement learning for joint spectrum and power allocation in cellular networks,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Deep reinforcement learning for joint spectrum and power allocation in cellular networks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.376074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.943681Z digest=sha256:7f078909be015fd96f48cf4824c05292199d7c1d4534e1515039885ad1f3ca33

Observation 8c8a43ac-33d9-4d35-aeda-e0329a01963f · outbound

This paper cites Multi-agent deep reinforcement learning for dynamic power allocation in wireless networks,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Multi-agent deep reinforcement learning for dynamic power allocation in wireless networks,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.359596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.948386Z digest=sha256:cf680e7e54bace83858a818bea6886355e055edfe96d936f6e608be0125e9de4

Observation abbb5159-d0a4-4f5a-a14c-376c7d667aa0 · outbound

This paper cites Multi- agent reinforcement learning for wireless user scheduling: Performance, scalablility, and generalization,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Multi- agent reinforcement learning for wireless user scheduling: Performance, scalablility, and generalization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.340416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.953229Z digest=sha256:31714489ba8621d817857777f095f1cf72af68ffa343f7bc91f2121dea794a6d

Observation 4786d838-3339-4a54-a55a-9d0ad12459ac · outbound

This paper cites Resource management in wireless networks via multi-agent deep reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Resource management in wireless networks via multi-agent deep reinforcement learning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:02.958639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:02.958639Z digest=sha256:6e3a098e95083f8b0ad4440fc1d5c468eaae58fad59c461b7de391f8ceca110f

Observation 9f6fe38c-60b8-4321-a8e6-407c26949b68 · outbound

This paper cites Distributed MARL for scheduling in conflict graphs,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Distributed MARL for scheduling in conflict graphs,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.311863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.963929Z digest=sha256:416e27c81b32dcd90da1467eb13f0ae6e596d6d485ee88c84e5f10b43d962ed5

Observation d2de51e2-2bdd-4ffa-9194-8accdc97981a · outbound

This paper cites Offline reinforcement learning for wireless network optimization with mixture datasets,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Offline reinforcement learning for wireless network optimization with mixture datasets,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.295227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.970809Z digest=sha256:4a5b9150237ee019ebd095490e8974d7900dd69d09f9733878971607cd9eab35

Observation 3daa3c11-e5c9-4b28-bc45-5af3768177b2 · outbound

This paper cites Advancing RAN slicing with offline reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Advancing RAN slicing with offline reinforcement learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.279124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.976751Z digest=sha256:1d23fbba8785eabdafc6fbc6cb776309ed4e612938e4cdabdfe72f82cae493db

Observation 0fc23e0a-4c89-4b71-abaf-532fb6834614 · outbound

This paper cites Mean-variance policy iteration for risk-averse reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Mean-variance policy iteration for risk-averse reinforcement learning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.261064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.982057Z digest=sha256:eda54bd3a3ef7a221e5cdf926491ad35e6c9499f2dbfdaba214426241f1f7de2

Observation e0aed05d-01fa-49d7-8e9f-d6d13de8f23a · outbound

This paper cites Averaged-dqn: Variance reduc- tion and stabilization for deep reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Averaged-dqn: Variance reduc- tion and stabilization for deep reinforcement learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.243701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.987694Z digest=sha256:247be3301d6f14068c6a9eb15ef953ed6c84563afa3786619a40381f6b83eee6

Observation 38faea7d-0567-4d26-9138-308f0ede6ca3 · outbound

This paper cites Average-Reward Reinforcement Learning with Trust Region Methods.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Average-Reward Reinforcement Learning with Trust Region Methods

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:02.992888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:02.992888Z digest=sha256:c3bf6d0663adea7a7e937a0834ffe71c176529a0427715f8eba595de2ceaf485

Observation 79263734-8de1-48bf-9c0d-79864f8f6157 · outbound

This paper cites The NS-3 network simulator,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management The NS-3 network simulator,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.224731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:02.997779Z digest=sha256:59d6667b68479d8bd31b227382baa0cad24f232634264a1342f7374878329751

Observation bf85f0c0-a91e-49a3-b59b-75d2d16428ae · outbound

This paper cites Network slicing architecture,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Network slicing architecture,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.206772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:03.002821Z digest=sha256:c474a81cdd059a318be5913809c02256afd12e807cc6079291988c3578c78fb1

Observation a0c4bc3e-559f-45c2-8445-4840f43886d2 · outbound

This paper cites NetworkGym: Democratizing Network AI via Sim-aaS,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management NetworkGym: Democratizing Network AI via Sim-aaS,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.188710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:03.007765Z digest=sha256:1121918f2a5284e9482a6480d8e069b18fe1d16cf30344d6196c9ab42acb1072

Observation 77cd05a6-0482-4fbd-8745-f0df02bb77af · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Soft Actor-Critic Algorithms and Applications

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:03.012480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:58:03.012480Z digest=sha256:69a320d6aac7c07553c39493e5e4c64c273b9ffbf0c5526109f3134b0d0ba89a

Observation 75f746b8-3394-412e-9bdd-861ba0a638f8 · outbound

This paper cites A Deeper Look at Discounting Mismatch in Actor-Critic Algorithms.

Average Reward Reinforcement Learning for Wireless Radio Resource Management A Deeper Look at Discounting Mismatch in Actor-Critic Algorithms

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:58:03.078406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:03.017617Z digest=sha256:9c86444f6a1d0acb5778a9b7a2fe9762ec0a711e926f702cd0c459bd50614c06

Observation 8a8d116c-0e34-452f-9949-38b760fd7858 · outbound

This paper cites Revisiting the minimalist approach to offline reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Revisiting the minimalist approach to offline reinforcement learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.171997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:03.022995Z digest=sha256:3c00924fb0a320619a62d7dea51652fd86b0bee836c7e50b765db7ffea4cc2e2

Observation 9bc0867d-2531-42e9-b963-a953fdb3ed79 · outbound

This paper cites Supported policy optimization for offline reinforcement learning,.

Average Reward Reinforcement Learning for Wireless Radio Resource Management Supported policy optimization for offline reinforcement learning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:03.154181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:58:03.027593Z digest=sha256:552f024a28ff9f86574aaa53aa634c9c7f52ebbeb0f3bdf45a1f893cdcaf07cf

Pith citing papers

No inbound Pith citation observations are available.