Pith. sign in

Paper Citation Record · LEDGER

Adaptive Data Exploitation in Deep Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2501.12620.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12620 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:04:46.991531Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f2a4f0d-9d03-48c4-b149-e1c74b208cc8 · outbound

This paper cites For the model update phase, the computational overhead is Oupdate = Oforward + Obackward, (15) where Oforward = Obs1 ∗ B ∗ Nbatches ∗ Nupdate epochs (16) and Obackward = Oforward ∗.

Adaptive Data Exploitation in Deep Reinforcement Learning For the model update phase, the computational overhead is Oupdate = Oforward + Obackward, (15) where Oforward = Obs1 ∗ B ∗ Nbatches ∗ Nupdate epochs (16) and Obackward = Oforward ∗

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:47.098382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.986726Z digest=sha256:22cfe93c984d5a2a64572b0049c33f0e19c9e6860ffb2aa0ff366a189f84b5d8

Observation 65ed9f9b-8195-4b85-944d-24baca67a081 · outbound

This paper cites an unresolved cited work.

Adaptive Data Exploitation in Deep Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T17:04:47.080523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.991531Z digest=sha256:6f4ac570cce6ef56c43b85cf278e672513dfd42e666c9b8403fc9ad6a89edf5a

Observation a89d4666-b8f9-41cd-91a4-69b343c95dbe · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.

Adaptive Data Exploitation in Deep Reinforcement Learning Dropout: a simple way to prevent neural networks from overfitting

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:46.939324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:04:46.939324Z digest=sha256:bc23e031c1a0ae6453338b6d8dc2e9d2bc7fbc62784fc9c4f2aa27bdec341fde

Observation d40834fd-d09a-4ba2-bdc1-9e44facb1add · outbound

This paper cites Rewarding episodic visitation discrepancy for exploration in reinforcement learning.

Adaptive Data Exploitation in Deep Reinforcement Learning Rewarding episodic visitation discrepancy for exploration in reinforcement learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:47.218309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.949106Z digest=sha256:d22428715d25490c63f7600b664c853774b0faf8cde6e9daca4f1a2943c94fa8

Observation 333312ed-983b-4693-880c-71b6b4fedfbf · outbound

This paper cites 39 Adaptive Data Exploitation in Deep Reinforcement Learning G.

Adaptive Data Exploitation in Deep Reinforcement Learning 39 Adaptive Data Exploitation in Deep Reinforcement Learning G

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:47.116608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.981908Z digest=sha256:03725e1a1be6d9455e5e454af4dab459950fb178d095592a90573987e515193f

Observation fe75fc48-d6e6-479d-87b7-6fa62bc6a6fe · outbound

This paper cites For each Procgen environment, Table 2 lists the best augmentation method of DrAC as reported in (Raileanu & Fergus, 2021).

Adaptive Data Exploitation in Deep Reinforcement Learning For each Procgen environment, Table 2 lists the best augmentation method of DrAC as reported in (Raileanu & Fergus, 2021)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:47.167648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.964423Z digest=sha256:cc40331bc37f5381b440a5552880a986241e2b10d019dbd43f2ddd439f302e3c

Observation fb566a84-2e95-438d-85f2-fc2e5af83878 · outbound

This paper cites an unresolved cited work.

Adaptive Data Exploitation in Deep Reinforcement Learning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-10T17:04:47.150079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.969282Z digest=sha256:e5f3bc8468bbda4bbaef9cdfb8538e4f72f49a58f3fa8d2009e3f8aec3f9fca7

Observation 259c1f85-2aa2-489f-ab05-dcbad524c1d4 · outbound

This paper cites In this part, we use the official implementation (Raileanu et al.,.

Adaptive Data Exploitation in Deep Reinforcement Learning In this part, we use the official implementation (Raileanu et al.,

Reference 255

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:47.184069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.959568Z digest=sha256:2bc657435dab9f380f2c3208640bd18a83388800b580c3c5513c45f3ae2c9b2c

Observation 9c485d65-63f3-4807-99df-0f1f09e34cf8 · outbound

This paper cites Master- ing visual continuous control: Improved data-augmented reinforcement learning.

Adaptive Data Exploitation in Deep Reinforcement Learning Master- ing visual continuous control: Improved data-augmented reinforcement learning

Reference 1988

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:47.233746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.944177Z digest=sha256:e771dbf3069fc778cfa094167c9a1335a71dbfdffa73bef13ae36d1ef67c9b2d

Observation 87514cc3-0225-44c3-b39f-58edfedd4a4a · outbound

This paper cites Badia, A.

Adaptive Data Exploitation in Deep Reinforcement Learning Badia, A

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:46.906488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:04:46.906488Z digest=sha256:1bbc95dbfbf95fbdc1d47fa1fe32cb8531ef68452ea35575510f96ab9201d931

Observation 6ca56ebb-5817-4acc-81da-c4e7c5225dc4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Adaptive Data Exploitation in Deep Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:46.934094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:04:46.934094Z digest=sha256:2245a104cfdc6960587c6720ddc8e62eb936305983f33d1c4d70d7aaa1db7e13

Observation 95a66473-883c-4354-9c20-b71e6da1e9b8 · outbound

This paper cites Prioritized Experience Replay.

Adaptive Data Exploitation in Deep Reinforcement Learning Prioritized Experience Replay

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:46.928969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:04:46.928969Z digest=sha256:a670f2ae50b2bc5a87bc4b76eb972e0b666124850e6ef29e15490530b0868e42

Observation 1ca8056b-83fc-42b3-a864-a3be104dc8ee · outbound

This paper cites The policy loss is defined as: Lπ(θ) = −Eτ ∼π [min (ρt(θ)At, clip (ρt(θ), 1 − ϵ, 1 + ϵ) At)] , (7) where ρt(θ) = πθ(at|st) πθold (at|st) , (8) and ϵ is a clipping range coefficient.

Adaptive Data Exploitation in Deep Reinforcement Learning The policy loss is defined as: Lπ(θ) = −Eτ ∼π [min (ρt(θ)At, clip (ρt(θ), 1 − ϵ, 1 + ϵ) At)] , (7) where ρt(θ) = πθ(at|st) πθold (at|st) , (8) and ϵ is a clipping range coefficient

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:47.202730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.953806Z digest=sha256:4f008084f25c7f6d0d13046fdef5d9ed233745d1c6cb7ef61a11417ff32c1947

Observation 6e537524-2daf-403f-baa6-79ea88b20ef2 · outbound

This paper cites W., Hilton, J., Klimov, O., and Schulman, J.

Adaptive Data Exploitation in Deep Reinforcement Learning W., Hilton, J., Klimov, O., and Schulman, J

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:47.276136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.918613Z digest=sha256:606a5429bb9986b0a648288f6b170430208d9d1feaa91c195d520a0d133488fb

Observation bfe265ff-ab1a-4d37-8bf0-a4fc32e9d83c · outbound

This paper cites and Bai, Y.

Adaptive Data Exploitation in Deep Reinforcement Learning and Bai, Y

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:47.260283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.923653Z digest=sha256:7648f9cdff11c54320583c4ce30a4f532dba4f244af76242aef7cd5ad76be9d3

Observation a98bf03c-5a33-4f71-9165-891d7b25f8fe · outbound

This paper cites Then we test the PPO agent with three ADEPT algorithms.

Adaptive Data Exploitation in Deep Reinforcement Learning Then we test the PPO agent with three ADEPT algorithms

Reference 2022

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T17:04:47.133881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.975605Z digest=sha256:5765aaa61cfd0a924354af670d74a5cd830e6f2a102db9162d2d3da7e622a240

Observation 85907215-7784-43fc-b120-2d3353e04b07 · outbound

This paper cites Lever- aging procedural generation to benchmark reinforcement learning.

Adaptive Data Exploitation in Deep Reinforcement Learning Lever- aging procedural generation to benchmark reinforcement learning

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:47.291854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T17:04:46.913255Z digest=sha256:2b39dc47e85a7e0c20701c53cae9c0a8f40fc2fe1553178fa10725c5555d7b55

Pith citing papers

No inbound Pith citation observations are available.