Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:58:03.027593Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2501.06700.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:58:03.027593Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fe418e7b-8936-457b-be7e-cf0d22d8cbb8 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Reinforcement learning: An introduction by richards’ sutton,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 63e51f31-3ce3-41dd-a0f5-29f0c4a04cee · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Average-reward off- policy policy evaluation with function approximation,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9be1150f-6663-4a75-9b9c-539ee364bde6 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Off-policy average reward actor-critic with deterministic policy search,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 153af67c-6a51-4bce-a795-01b689aa1b11 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Power control for a network of access points,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 444c08a9-f2dd-430c-992a-1da342f5233d · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Base station employing shared resources among antenna units,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 832225f1-a78d-4760-b37c-37bcf3a5e033 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Methods and apparatus for power management in a wireless communication system,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9b85fed7-2f67-4e5f-a05b-86b4a65ceb9e · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Generalized global bandit and its application in cellular coverage optimization,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3048fc99-bf64-4f0d-87f0-7383e2230e45 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management A non-stationary online learning approach to mobility management,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8b685e6b-9a91-481d-822e-971cf0f84b39 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management A Deep Q-Learning Method for Downlink Power Allocation in Multi-Cell Networks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f0986a6-d45d-476f-b06d-17ed7a5f424d · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Power allocation in multi-user cellular networks with deep Q learning approach,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d86cf87a-3445-470b-a252-e99e5069f05f · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Joint power control and channel allocation for interference mitigation based on reinforcement learning,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e896803b-aba4-4573-b966-ed5796e76b39 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Deep actor-critic learning for distributed power control in wireless mobile networks,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a9e7afcb-a20d-4767-806d-f055c673c843 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Deep reinforcement learning based wire- less network optimization: A comparative study,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4a8fb1ba-fd12-4558-8c98-551e9435361f · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management ColO- RAN: Developing Machine Learning-based xApps for Open RAN Closed-loop Control on Programmable Experimental Platforms,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0efcb88b-4d34-4582-8df1-cf93c067722c · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management FlexRAN: A flexible and programmable platform for software- defined radio access networks,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 04d1d7ad-6a01-41cf-9b35-9849bb3fec9e · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Deep reinforcement learning for joint spectrum and power allocation in cellular networks,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8c8a43ac-33d9-4d35-aeda-e0329a01963f · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Multi-agent deep reinforcement learning for dynamic power allocation in wireless networks,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation abbb5159-d0a4-4f5a-a14c-376c7d667aa0 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Multi- agent reinforcement learning for wireless user scheduling: Performance, scalablility, and generalization,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4786d838-3339-4a54-a55a-9d0ad12459ac · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Resource management in wireless networks via multi-agent deep reinforcement learning,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f6fe38c-60b8-4321-a8e6-407c26949b68 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Distributed MARL for scheduling in conflict graphs,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d2de51e2-2bdd-4ffa-9194-8accdc97981a · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Offline reinforcement learning for wireless network optimization with mixture datasets,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3daa3c11-e5c9-4b28-bc45-5af3768177b2 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Advancing RAN slicing with offline reinforcement learning,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0fc23e0a-4c89-4b71-abaf-532fb6834614 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Mean-variance policy iteration for risk-averse reinforcement learning,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e0aed05d-01fa-49d7-8e9f-d6d13de8f23a · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Averaged-dqn: Variance reduc- tion and stabilization for deep reinforcement learning,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 38faea7d-0567-4d26-9138-308f0ede6ca3 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Average-Reward Reinforcement Learning with Trust Region Methods
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79263734-8de1-48bf-9c0d-79864f8f6157 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management The NS-3 network simulator,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bf85f0c0-a91e-49a3-b59b-75d2d16428ae · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Network slicing architecture,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a0c4bc3e-559f-45c2-8445-4840f43886d2 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management NetworkGym: Democratizing Network AI via Sim-aaS,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 77cd05a6-0482-4fbd-8745-f0df02bb77af · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Soft Actor-Critic Algorithms and Applications
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75f746b8-3394-412e-9bdd-861ba0a638f8 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management A Deeper Look at Discounting Mismatch in Actor-Critic Algorithms
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8a8d116c-0e34-452f-9949-38b760fd7858 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Revisiting the minimalist approach to offline reinforcement learning,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9bc0867d-2531-42e9-b963-a953fdb3ed79 · outbound
Average Reward Reinforcement Learning for Wireless Radio Resource Management Supported policy optimization for offline reinforcement learning,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
No inbound Pith citation observations are available.