Pith. sign in

Paper Citation Record · LEDGER

Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2101.05982.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2101.05982 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:23:15.758249Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T00:05:09.694789Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2bc50aab-90da-403e-8d65-f027559c3200 · inbound

Hadamax Encoding: Elevating Performance in Model-Free Atari cites this paper.

Hadamax Encoding: Elevating Performance in Model-Free Atari Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:15.758249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:15.758249Z digest=sha256:a6081bb55eb3a63a9bff7539220c53ce9dc589c07ab41efd0dbc4dcd3701f59e

Observation ef582d84-4c4b-44e8-9edd-05a8791a980d · inbound

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only cites this paper.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:44.897231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:44.897231Z digest=sha256:d1969f44a6a66b2714a78b19c444c563ede989be27064f2bd07c320d17e59531

Observation 6d8f962e-0ab7-4fa2-b9c5-655623c282df · inbound

Universal Value-Function Uncertainties cites this paper.

Universal Value-Function Uncertainties Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:54.433515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:49:54.433515Z digest=sha256:76bce8faa6289207842d318ab89e3e5c96fa4287f529e26c241fd8cacdaf5eb0

Observation 2ee69c96-805f-45cb-82de-044df882b3db · inbound

Safe Planning and Policy Optimization via World Model Learning cites this paper.

Safe Planning and Policy Optimization via World Model Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:38:20.418907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:38:20.418907Z digest=sha256:bf0d53b37aec9349a365eb32e60d138a66d6e93a6b0fd5277fa0e1f45f79a670

Observation b58b6687-ec6e-4d18-b8aa-6fc75da5b71c · inbound

The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning cites this paper.

The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:32:50.412668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:32:50.412668Z digest=sha256:e81ba092603bb8e06b35616dd46413db825fc556caf4e819a1ba9fbd9a2e14d6

Observation 726edaa6-805d-4d80-9f87-9537bbeb5c61 · inbound

StaQ it! Growing neural networks for Policy Mirror Descent cites this paper.

StaQ it! Growing neural networks for Policy Mirror Descent Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:01.863748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:01.863748Z digest=sha256:84d013ad99d595321c36b20dfdb691a62022fdab9090a5d4c1a6f0cd0b4de595

Observation a1146ce5-7daf-4a84-a69c-1cb359dba087 · inbound

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control cites this paper.

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:07.736250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:07.736250Z digest=sha256:4573862881409186353251bb58bcac2dcb2b40ee673953446d828a7f8a88fd2e

Observation 1ea33fd0-e326-43fa-8320-6b28f72fc17d · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:12:05.288471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:5999107eee96f118b9b04f6e5e4775961dedf49d9cf6c3d1e41980e62b10e294

Observation 2931c806-2933-4b2c-994c-89cf5963f460 · inbound

Exploring the robustness of TractOracle methods in RL-based tractography cites this paper.

Exploring the robustness of TractOracle methods in RL-based tractography Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:13:56.033564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:13:56.033564Z digest=sha256:4a04099d50f60a4bc4e92134548971e20ca64536e6147c820576357dcabb082f

Observation 3a31fd7d-cfa2-471d-95dc-8e5cef1c18ba · inbound

From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning cites this paper.

From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:40:34.101065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T00:36:32.681470Z digest=sha256:bf84cecc317b1e0d20e5a1adb54e2d17cc51fadddd1b2fe3f428894dd8926cad

Observation 001dbc78-7a2b-4cff-a309-75d678605e99 · inbound

From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning cites this paper.

From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:04:19.084955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T19:02:38.240098Z digest=sha256:d325ad666bca69cc43be700c3a2619df9265ef7714be8e108c2d6e3647144f76

Observation 98f63cc1-5242-4820-86e2-1075fe6e3200 · inbound

What Matters for Simulation to Online Reinforcement Learning on Real Robots cites this paper.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.350061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.350061Z digest=sha256:5da969ae7b95a004454526ae862e6754baa7e75442673834124dff0d5d506d67

Observation 87c728a3-0fd5-4f24-a6d8-3c060a8719c8 · inbound

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity cites this paper.

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:14:46.870973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:11:04.222842Z digest=sha256:b0ed659bca334567d884e51777a88de6e9244a8c322edd60d21b81e521731e07

Observation 41a496fc-aae2-426e-9983-0253162fa8d8 · inbound

RL Token: Bootstrapping Online RL with Vision-Language-Action Models cites this paper.

RL Token: Bootstrapping Online RL with Vision-Language-Action Models Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:09.203649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:56:34.978806Z digest=sha256:fc9a0c085874d91ff6ffb9330bca5154ecd8daac813030a1ab3022d4e014527d

Observation 5199169e-efc2-4347-9902-71254aaadce7 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 139

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:10:42.429745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:216dffd1376079642c3d12598f5019be755afa1de495b003ef3f2ab89683563c

Observation 2c6efab4-06a8-4f45-a19c-dc765a6f0c1a · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:05:09.696207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:3f96845f11a55717a3f12074b99f7359e0e0bd12e529ef3c50073e1751827070

Observation e874d28b-ad6c-4613-a732-8fe2791b8b0b · inbound

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling cites this paper.

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:08:59.689210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T03:05:36.871497Z digest=sha256:5cba2863b000db4237b2429cc481a621fbe0f3ff107513687f3eb182e79579f9

Observation ea3cfe6a-0723-4654-b752-a326c9d11182 · inbound

Implicit Safety Alignment from Crowd Preferences cites this paper.

Implicit Safety Alignment from Crowd Preferences Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:31:16.890011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-22T08:28:41.652865Z digest=sha256:be6fdd5be2c0c9033572ac3953dbed877af4b4bc40af8e900ae884386fb8d087