Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:04:46.991531Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2501.12620.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:04:46.991531Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7f2a4f0d-9d03-48c4-b149-e1c74b208cc8 · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning For the model update phase, the computational overhead is Oupdate = Oforward + Obackward, (15) where Oforward = Obs1 ∗ B ∗ Nbatches ∗ Nupdate epochs (16) and Obackward = Oforward ∗
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 65ed9f9b-8195-4b85-944d-24baca67a081 · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a89d4666-b8f9-41cd-91a4-69b343c95dbe · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning Dropout: a simple way to prevent neural networks from overfitting
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d40834fd-d09a-4ba2-bdc1-9e44facb1add · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning Rewarding episodic visitation discrepancy for exploration in reinforcement learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 333312ed-983b-4693-880c-71b6b4fedfbf · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning 39 Adaptive Data Exploitation in Deep Reinforcement Learning G
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fe75fc48-d6e6-479d-87b7-6fa62bc6a6fe · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning For each Procgen environment, Table 2 lists the best augmentation method of DrAC as reported in (Raileanu & Fergus, 2021)
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb566a84-2e95-438d-85f2-fc2e5af83878 · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 259c1f85-2aa2-489f-ab05-dcbad524c1d4 · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning In this part, we use the official implementation (Raileanu et al.,
Reference 255
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9c485d65-63f3-4807-99df-0f1f09e34cf8 · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning Master- ing visual continuous control: Improved data-augmented reinforcement learning
Reference 1988
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 87514cc3-0225-44c3-b39f-58edfedd4a4a · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning Badia, A
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ca56ebb-5817-4acc-81da-c4e7c5225dc4 · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a66473-883c-4354-9c20-b71e6da1e9b8 · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning Prioritized Experience Replay
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca8056b-83fc-42b3-a864-a3be104dc8ee · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning The policy loss is defined as: Lπ(θ) = −Eτ ∼π [min (ρt(θ)At, clip (ρt(θ), 1 − ϵ, 1 + ϵ) At)] , (7) where ρt(θ) = πθ(at|st) πθold (at|st) , (8) and ϵ is a clipping range coefficient
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6e537524-2daf-403f-baa6-79ea88b20ef2 · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning W., Hilton, J., Klimov, O., and Schulman, J
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bfe265ff-ab1a-4d37-8bf0-a4fc32e9d83c · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning and Bai, Y
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a98bf03c-5a33-4f71-9165-891d7b25f8fe · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning Then we test the PPO agent with three ADEPT algorithms
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85907215-7784-43fc-b120-2d3353e04b07 · outbound
Adaptive Data Exploitation in Deep Reinforcement Learning Lever- aging procedural generation to benchmark reinforcement learning
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.