Pith. sign in

Paper Citation Record · LEDGER

Value-Based Deep RL Scales Predictably

As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 5 inbound Pith citation observations for arXiv:2502.04327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04327 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:50:34.155640Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:52.599976Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T23:09:12.714863Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a94a2006-4cda-4f14-981a-d5d2dda951c6 · outbound

This paper cites write newline.

Value-Based Deep RL Scales Predictably write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:33.979227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:33.979227Z digest=sha256:7ed622452aa424aab8720e9c04cb3737810b57b6f697eec83e44d5ff62885c4f

Observation ab0650aa-07fb-499b-9796-cde2ef210b24 · outbound

This paper cites GPT -4 technical report.

Value-Based Deep RL Scales Predictably GPT -4 technical report

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.673980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:33.983471Z digest=sha256:d9f642fb9855a7133e7d76a1b0f819b54e00655e0b14e8b8e96bb95e1181c8dd

Observation 46769471-4517-4647-ad00-03ab6025acd0 · outbound

This paper cites The isotonic regression problem and its dual.

Value-Based Deep RL Scales Predictably The isotonic regression problem and its dual

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.664945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:33.986597Z digest=sha256:873671c26bffda4c194723cd232013b83c5b7744c96679eb035a6ddb6593d15c

Observation d089a7e0-efd8-4d33-834b-0c01e5af1a3a · outbound

This paper cites Pattern Recognition and Machine Learning.

Value-Based Deep RL Scales Predictably Pattern Recognition and Machine Learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.655527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:33.990214Z digest=sha256:f7cdc61953d76ecb2b6eaa3d9adcd29adab1d035d8f029d62def0455195d0523

Observation 2a04ab3d-9f6b-422c-a0dc-1c5360bd21ee · outbound

This paper cites OpenAI Gym , 2016.

Value-Based Deep RL Scales Predictably OpenAI Gym , 2016

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.646732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:33.993346Z digest=sha256:773804b31a76ebccd747834c4500c9de61c7d092af72d50f42fe6c5384db8a45

Observation 7eb90764-31b0-40c4-b454-096967c9bbc5 · outbound

This paper cites Video generation models as world simulators.

Value-Based Deep RL Scales Predictably Video generation models as world simulators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:33.996307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:33.996307Z digest=sha256:eec593bddedc2824bb8cf3258709cc56f8d68dbbd47bcaa52e778eae21907be3

Observation 852f239f-9b6c-4765-8c63-6b52b164378d · outbound

This paper cites Randomized ensembled double Q -learning: Learning fast without a model.

Value-Based Deep RL Scales Predictably Randomized ensembled double Q -learning: Learning fast without a model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.632177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:33.999739Z digest=sha256:7fc11e2298e98847653dfc062ee74c1592c4695a0f13ce3298504fbef1c885e7

Observation 60159157-8014-4dbf-88d1-249abf341791 · outbound

This paper cites The value-improvement path: Towards better representations for reinforcement learning.

Value-Based Deep RL Scales Predictably The value-improvement path: Towards better representations for reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.623340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.002855Z digest=sha256:61511cfc8c8500eccca73c663ee78558996472a285cd905bdb76cfbbf9fe4f72

Observation 3f4817d9-e105-4c53-9a66-0ca15e87e663 · outbound

This paper cites Sample-efficient reinforcement learning by breaking the replay ratio barrier.

Value-Based Deep RL Scales Predictably Sample-efficient reinforcement learning by breaking the replay ratio barrier

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.614096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.005965Z digest=sha256:d5882ebbd4ba67a9050414623811117ef6e76b7cc814dd5568ee56b3ec563cba

Observation b2c7db4c-7743-4e74-9ef8-232c268208c6 · outbound

This paper cites The Llama 3 herd of models.

Value-Based Deep RL Scales Predictably The Llama 3 herd of models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.605256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.009020Z digest=sha256:0cc69073874efc7444111aa8e9aedb0706ddb15a3d68794bd2f78fdaede7bb69

Observation 8cc2861c-5e2b-4a33-b119-9338a2ee829a · outbound

This paper cites IMPALA : Scalable distributed deep- RL with importance weighted actor-learner architectures.

Value-Based Deep RL Scales Predictably IMPALA : Scalable distributed deep- RL with importance weighted actor-learner architectures

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.596658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.011583Z digest=sha256:47834d88141c5891882fd1e2caa032dcb6c6645e99fdabd2ddfc2cd905bcb2b8

Observation d5250699-4756-407d-8bbb-dc4baada40de · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Value-Based Deep RL Scales Predictably Language models scale reliably with over-training and on downstream tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.587849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.014741Z digest=sha256:441207adbea1408c555e361f25ef370ad020a682db6858a1ee7786acf1f4e801

Observation 8ccf7e35-19a6-4457-8bca-a0d23914458b · outbound

This paper cites Simplifying deep temporal difference learning.

Value-Based Deep RL Scales Predictably Simplifying deep temporal difference learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.578847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.017377Z digest=sha256:37a5be32621b1b373dba06d9860746ae62d4e895a89a569d6eb05119d7d3a232

Observation 65dda2fc-cb8a-4df9-831e-6e5e1a2a50b8 · outbound

This paper cites Scaling laws for reward model overoptimization.

Value-Based Deep RL Scales Predictably Scaling laws for reward model overoptimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.020093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.020093Z digest=sha256:b0eeb0d64a41d1f9fae10f072a61ce0014b5b4b3f445cdc75029ad0c4f4d1210

Observation 6868a284-8f1c-40c2-b8f4-abef9ae465be · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Value-Based Deep RL Scales Predictably Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.564758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.022820Z digest=sha256:b8062dde2795819aedfd47bc993c1f459f709ba6f16f52b7025354dd9f15471d

Observation a9aabe5d-25c1-4a34-a9d0-7b6e22123557 · outbound

This paper cites Mastering diverse domains through world models.

Value-Based Deep RL Scales Predictably Mastering diverse domains through world models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.556484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.025588Z digest=sha256:5d7d8fe0551b90898cd652687d84af5aca323b38905c57be98c68a50c981d0d0

Observation f93fd152-ee35-4bb2-8b61-7b5df227e8e4 · outbound

This paper cites Scaling laws for single-agent reinforcement learning.

Value-Based Deep RL Scales Predictably Scaling laws for single-agent reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.547912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.028174Z digest=sha256:2e9696cf322aac12629e818aa6fe0e81d13bb434d88cf2f38142b6e128d58385

Observation 38c12941-3766-4884-abc6-5dd8d4ade8a0 · outbound

This paper cites Training compute-optimal large language models.

Value-Based Deep RL Scales Predictably Training compute-optimal large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.539244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.031060Z digest=sha256:260d7c1dba5caec5f94299e87da034af512f8efd0916dc013d5b6569c685586a

Observation d9a6a78b-3850-451b-a8f8-8d996de70a2f · outbound

This paper cites When to trust your model: Model-based policy optimization.

Value-Based Deep RL Scales Predictably When to trust your model: Model-based policy optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.530725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.033876Z digest=sha256:40979facb542b56c19e2348cac9f373d150df1a6141598b694d3166aff6cd5b8

Observation 445f83e8-9890-4e60-9939-858eafe69d29 · outbound

This paper cites an unresolved cited work.

Value-Based Deep RL Scales Predictably Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:50:34.522130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.037606Z digest=sha256:ce4c7b91d3830ca5f6e5f3eabff7dcc53fe0caca812449f09a0b37647feffe9f

Observation 834dc572-89e9-4b65-8a16-51917ffbe58c · outbound

This paper cites Scaling laws for neural language models.

Value-Based Deep RL Scales Predictably Scaling laws for neural language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.513941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.040612Z digest=sha256:c0551b3cc04186595e4959537d0db28f1d2a84824ee7828e7ccf395f5a02e080

Observation 4318d582-b797-4984-95ff-7d4fa607f897 · outbound

This paper cites One weird trick for parallelizing convolutional neural networks.

Value-Based Deep RL Scales Predictably One weird trick for parallelizing convolutional neural networks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.505067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.043514Z digest=sha256:5029d3f673024c7f7f90f2dba877c0fa450934bbf2c055224ea06d40f658822c

Observation 45af0646-657e-42e6-81ce-8e5607ea19d4 · outbound

This paper cites Implicit under-parameterization inhibits data-efficient deep reinforcement learning.

Value-Based Deep RL Scales Predictably Implicit under-parameterization inhibits data-efficient deep reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.496003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.046258Z digest=sha256:fa46f5c2c156f066f33cc8ac4ce095eb5009428fd574d38ad02548abe5741971

Observation 6d8cd011-d773-4ed7-a2bf-4ba6b7b85091 · outbound

This paper cites DR3 : Value-based deep reinforcement learning requires explicit regularization.

Value-Based Deep RL Scales Predictably DR3 : Value-based deep reinforcement learning requires explicit regularization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.486942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.049833Z digest=sha256:f6aa36a610b1452ed0194f64067b2a43cef000360706b30878f468ad79383e43

Observation 2d7d541c-d0e3-4144-b23a-b5ac511ae60d · outbound

This paper cites Offline Q -learning on diverse multi-task data both scales and generalizes.

Value-Based Deep RL Scales Predictably Offline Q -learning on diverse multi-task data both scales and generalizes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.477670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.052642Z digest=sha256:fd4c58ed6224db12cead8a2a1faa3944c2f2a1a1fc9ce1b69fdd462d565eb18c

Observation a77f9a6a-8821-42b1-8285-193d73da3323 · outbound

This paper cites Plastic: Improving input and label plasticity for sample efficient reinforcement learning.

Value-Based Deep RL Scales Predictably Plastic: Improving input and label plasticity for sample efficient reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.468705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.055443Z digest=sha256:18384dcbc6ac68e47910504eb3be8a9ce7c06fea33a7148c0f3c35e71c0bbaa6

Observation 42100522-0b73-4742-a361-b3e43167bb73 · outbound

This paper cites SimBa : Simplicity bias for scaling up parameters in deep reinforcement learning.

Value-Based Deep RL Scales Predictably SimBa : Simplicity bias for scaling up parameters in deep reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.459855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.058595Z digest=sha256:25b8e415071b0420e3b4631fed0ceb723da846a58d7dfd894cab9c7596b67e56

Observation 3aaa09b1-f44e-4579-bef2-a6aa6218ac8a · outbound

This paper cites Offline reinforcement learning: Tutorial, review, and perspectives on open problems.

Value-Based Deep RL Scales Predictably Offline reinforcement learning: Tutorial, review, and perspectives on open problems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.061981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.061981Z digest=sha256:da7ffb120e8749fea3632f562a0d558ffaf6aa141e43762770815fc910c17e2e

Observation 6f7126d6-f226-4751-9d6f-f5391b2d76eb · outbound

This paper cites Efficient deep reinforcement learning requires regulating overfitting.

Value-Based Deep RL Scales Predictably Efficient deep reinforcement learning requires regulating overfitting

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.445968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.064627Z digest=sha256:a766c026e52d687e7623ba33a456c3604f0307ffbdc8b915fa02186fd4bc34f3

Observation 8123446f-328a-45db-ac53-6d0b50893b5b · outbound

This paper cites Parallel Q -learning: Scaling off-policy reinforcement learning under massively parallel simulation.

Value-Based Deep RL Scales Predictably Parallel Q -learning: Scaling off-policy reinforcement learning under massively parallel simulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.436750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.067376Z digest=sha256:0ee3cb2fdad90c3e138e458199f7311e1aed39538b5bd335670250f17fff17fe

Observation da99da43-bab3-4570-aadd-b9f5ac7bc6b6 · outbound

This paper cites Continuous control with deep reinforcement learning.

Value-Based Deep RL Scales Predictably Continuous control with deep reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.427523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.070061Z digest=sha256:e903b2a58abc82b2cdb08e8d5708a412fc9c7b4763b9279dd1b8e774a4e4d861

Observation 84f79e94-a75d-49cf-b86f-7c976fd631aa · outbound

This paper cites Scaling laws for fine-grained mixture of experts.

Value-Based Deep RL Scales Predictably Scaling laws for fine-grained mixture of experts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.418696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.072762Z digest=sha256:4ab3257cdc5ff0c6b4104bed5637fc34920aa712f218685afd1489d8722b701f

Observation 95c8a962-f99a-4380-b627-72f3c86d5438 · outbound

This paper cites Understanding plasticity in neural networks.

Value-Based Deep RL Scales Predictably Understanding plasticity in neural networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.409854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.075389Z digest=sha256:3bd84e0262992564c795faa818ef6ffd4b7e7e47b4bcaa7833b91a96b51b98ca

Observation 420e74c0-e1dd-4c9d-bd63-330696925885 · outbound

This paper cites Isaac Gym : High performance GPU -based physics simulation for robot learning.

Value-Based Deep RL Scales Predictably Isaac Gym : High performance GPU -based physics simulation for robot learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.401201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.078156Z digest=sha256:582b67c975d1ef0502cede0d15499ff09e5b67a1c1cf5eedc53aafbe3464a3dd

Observation 0d76c8ba-933a-4a8e-8993-da19d8bfa3d0 · outbound

This paper cites An empirical model of large-batch training.

Value-Based Deep RL Scales Predictably An empirical model of large-batch training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.392251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.080820Z digest=sha256:66068971f76c3ca40b0b2196d4d1122087ba9a09e355015c9c8054d81c1f3f09

Observation 48466706-5f65-4e81-9dc2-308b1bd52c5e · outbound

This paper cites Human-level control through deep reinforcement learning.

Value-Based Deep RL Scales Predictably Human-level control through deep reinforcement learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.083679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.083679Z digest=sha256:199d6d508c20ae07285b4b30f2978a1f10134af6a3501a79eda8b671657cdc73

Observation f3dbbb1e-7b6f-4a92-bbf1-8f16247f1a65 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Value-Based Deep RL Scales Predictably Asynchronous methods for deep reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.378538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.086479Z digest=sha256:31cd3f3639f16623d29821b3eacf2a72d0dd94392723e9e53874e86cb087d1c5

Observation ee438f56-4ef8-409c-904c-0dbf3daf1055 · outbound

This paper cites Scaling data-constrained language models.

Value-Based Deep RL Scales Predictably Scaling data-constrained language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.369315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.089206Z digest=sha256:b5379d08c089da0aebfee171199ce2d6e7611c1166dcb027a7d9a6546e39baa5

Observation 0c005f43-a340-46aa-a5df-b1260383fd20 · outbound

This paper cites Overestimation, overfitting, and plasticity in actor-critic: The bitter lesson of reinforcement learning.

Value-Based Deep RL Scales Predictably Overestimation, overfitting, and plasticity in actor-critic: The bitter lesson of reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.360790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.092049Z digest=sha256:c38a1ee65bf360ae4d641bc1bd071a9456f6370785433af55a20801c209fd109

Observation 3da14635-a3a1-4516-bf12-8fb94a1243ee · outbound

This paper cites Bigger, regularized, optimistic: Scaling for compute and sample-efficient continuous control.

Value-Based Deep RL Scales Predictably Bigger, regularized, optimistic: Scaling for compute and sample-efficient continuous control

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.351722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.094759Z digest=sha256:7427b960b501b20d38fc8ea7d05c742978e1da06c7dc251bb79d4ad2cf371195

Observation 15f7c4a0-458e-488e-b1be-109c74191160 · outbound

This paper cites The primacy bias in deep reinforcement learning.

Value-Based Deep RL Scales Predictably The primacy bias in deep reinforcement learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.342821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.097737Z digest=sha256:016800a0068946156a5b8658104ea26716f69bf8ebcb5c085a65307ddf5bf273

Observation 1b2963ce-248d-4e98-8dd4-4600f46c7c7b · outbound

This paper cites Is value learning really the main bottleneck in offline RL ? Advances in Neural Information Processing Systems, 2024.

Value-Based Deep RL Scales Predictably Is value learning really the main bottleneck in offline RL ? Advances in Neural Information Processing Systems, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.333902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.100435Z digest=sha256:d621af4d23e6d28e8da301a908bdf91f5d50597b75ec756d41bac1de757a5fc8

Observation a1946095-49fa-4c8a-bbef-b856df7aba23 · outbound

This paper cites Hierarchical text-conditional image generation with CLIP latents.

Value-Based Deep RL Scales Predictably Hierarchical text-conditional image generation with CLIP latents

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.325067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.103094Z digest=sha256:01de3d50037b411b86884d6bc0fc3c8dcf257b625c0e8ef4712315d071a52278

Observation 2e36fc28-0661-4ad9-bcd8-81ada1a772bb · outbound

This paper cites Mastering Atari , Go , chess and Shogi by planning with a learned model.

Value-Based Deep RL Scales Predictably Mastering Atari , Go , chess and Shogi by planning with a learned model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.316118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.106006Z digest=sha256:3c939d086ba239e350e4d293dc355ab283e64f1435f7dd30ea52b7ca0de94969

Observation 9f4e15e0-c640-431a-9fb8-700b030b5344 · outbound

This paper cites Proximal policy optimization algorithms.

Value-Based Deep RL Scales Predictably Proximal policy optimization algorithms

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.108691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.108691Z digest=sha256:fde07f38e86339b8f2b528dfd650b866c80a337b653f8395aaa04b892f5fcbb4

Observation 449bf428-f076-4b1c-94f1-bb0404809cd8 · outbound

This paper cites Bigger, better, faster: Human-level Atari with human-level efficiency.

Value-Based Deep RL Scales Predictably Bigger, better, faster: Human-level Atari with human-level efficiency

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.302091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.111324Z digest=sha256:8baa730df1eeed33c862319a43666a13987ffc1eb4c198a4c07e6b48db044bd9

Observation 5a430338-03f3-4ad6-b4c5-a34cd3733eec · outbound

This paper cites Mastering the game of Go with deep neural networks and tree search.

Value-Based Deep RL Scales Predictably Mastering the game of Go with deep neural networks and tree search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.113924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.113924Z digest=sha256:a3e6cdaeba05f4a0413999b58fcb5ae935cc31a1793bc905c73bfbdb19941bd2

Observation 3eaa4bac-7c82-40f6-b3df-9b940bd3a4d9 · outbound

This paper cites SAPG : Split and aggregate policy gradients.

Value-Based Deep RL Scales Predictably SAPG : Split and aggregate policy gradients

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.288447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.116676Z digest=sha256:36f87e2a2f5d847aaf4da892e7d23ba6f80537481e0a7a0d4e828c6ddc84f67d

Observation 13a14c9a-0a6f-42c0-a627-5d31545b7602 · outbound

This paper cites The dormant neuron phenomenon in deep reinforcement learning.

Value-Based Deep RL Scales Predictably The dormant neuron phenomenon in deep reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.279987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.119692Z digest=sha256:5d1f336ee440f551eeeb3dad2c0d5d101857638bcc35435bd9897d1c4006fbcf

Observation d880c06a-a724-438d-887c-f8e3e5ce3455 · outbound

This paper cites Offline actor-critic reinforcement learning scales to large models.

Value-Based Deep RL Scales Predictably Offline actor-critic reinforcement learning scales to large models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.271721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.122488Z digest=sha256:5515a0fd297891b79f808401a42f25bbb713e073955646039cccb46b97bb6cd5

Observation c808fa9a-2d26-4506-a96a-d5b8cf981156 · outbound

This paper cites Reinforcement Learning: An Introduction.

Value-Based Deep RL Scales Predictably Reinforcement Learning: An Introduction

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.263108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.125185Z digest=sha256:cbde416c9efd9c60b3a294c79ae4fd5a89f8641648981bb57f75242084143cc6

Observation fde0c6e6-6d02-4df5-af67-dbcdad5eef6b · outbound

This paper cites DeepMind control suite.

Value-Based Deep RL Scales Predictably DeepMind control suite

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.254228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.128913Z digest=sha256:ed51fe47f6c0adf3af1930067289e2ad02969cf9ee4f9b28d6e6a4e483c5bba2

Observation 4f5daa81-71cf-47e6-9a55-6a1aa11b9cef · outbound

This paper cites Gemini : A family of highly capable multimodal models.

Value-Based Deep RL Scales Predictably Gemini : A family of highly capable multimodal models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.245209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.131743Z digest=sha256:86a8873c3450e78d6dd8d3aa93fb409cc7474abe45333b0c8e0104e257380118

Observation a61064de-eb67-4b09-9730-3f2f821f42e5 · outbound

This paper cites dm\_control: Software and tasks for continuous control.

Value-Based Deep RL Scales Predictably dm\_control: Software and tasks for continuous control

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.236045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.134688Z digest=sha256:54de7644fcf1e62e7b898fd59114ff7d9f38506ac98bcad250235df5941ed258

Observation 2e34445f-b5a9-4d6f-afdb-e3684d360775 · outbound

This paper cites SciPy 1.0: Fundamental algorithms for scientific computing in Python.

Value-Based Deep RL Scales Predictably SciPy 1.0: Fundamental algorithms for scientific computing in Python

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.227606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.137501Z digest=sha256:975a1a2cfae95194ddc945c202b779a1f6c071b6b93216c45bf465f139160093

Observation 4eff357b-0b1b-4dae-9a4b-5508821a0fdc · outbound

This paper cites DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization.

Value-Based Deep RL Scales Predictably DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.140276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.140276Z digest=sha256:c5c6498849f610c962588f8f967bd849b012e2a378d477d1f34172d3a3bc2837

Observation de0631a8-5a78-4ab3-8c53-81235939c286 · outbound

This paper cites Tensor programs V : Tuning large neural networks via zero-shot hyperparameter transfer.

Value-Based Deep RL Scales Predictably Tensor programs V : Tuning large neural networks via zero-shot hyperparameter transfer

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.218727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.143693Z digest=sha256:ae0d8e66d16d14f4b319eb280c66bc0a86bbc5dd6f27ac3732e345e44e264634

Observation 3c1fa771-e5cd-4d0a-979f-d37b620ecc1a · outbound

This paper cites How to leverage unlabeled data in offline reinforcement learning.

Value-Based Deep RL Scales Predictably How to leverage unlabeled data in offline reinforcement learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:50:34.209081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:50:34.146473Z digest=sha256:a182372884f046f2e4743078e617a79737f780626d9881ac31eb1557ee29b0e8

Observation b2c92c87-3905-4123-b6a1-b0bbf2073a2d · outbound

This paper cites @esa (Ref.

Value-Based Deep RL Scales Predictably @esa (Ref

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.149185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.149185Z digest=sha256:ea6cbd026c0ee909f1eac5fbbcb843f6137b77aad0f83a7de683f8b6c192bedd

Observation a06c7c6a-6d48-47bf-a2db-61889dbf1be1 · outbound

This paper cites an unresolved cited work.

Value-Based Deep RL Scales Predictably Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.152692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.152692Z digest=sha256:076eb00a38607d47dae0f528e73dc2c35006dd37ab54af7785958b430a8e0661

Observation 5295faeb-a8ea-4c43-854e-c51d9f0ff23c · outbound

This paper cites an unresolved cited work.

Value-Based Deep RL Scales Predictably Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T22:50:34.155640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:50:34.155640Z digest=sha256:25cc3ff9d2ebdf112a997a674a43d3662363055fd94f25d90fcb3230e49a3579

Pith citing papers

Observation c54498a5-69f8-4676-9eeb-e9dc48e64daa · inbound

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners cites this paper.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Value-Based Deep RL Scales Predictably

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.599976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.599976Z digest=sha256:6713667ff901c049d198b6a3d8fd685b1d6f99e8c1ef5620412830041344682b

Observation e7d746db-8c02-4bce-8d81-72aab31f19d0 · inbound

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies cites this paper.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Value-Based Deep RL Scales Predictably

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:06.797394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:06.797394Z digest=sha256:85c7d0ead948fff39b66fae2c9de0696a106940438735ca3c43c416d99fc8f2c

Observation b6387637-3a7d-4c8d-ab63-6aa13395de49 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Value-Based Deep RL Scales Predictably

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.580539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T05:33:19.038889Z digest=sha256:b2cf7f3c9a39a4a39ce60e1ce6b494e6fd7dbdcd5168650dfa3d39e94a392a5a

Observation cadfc5c9-2956-4f39-ade1-1c0075c43668 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Value-Based Deep RL Scales Predictably

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:32:24.810267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T06:27:38.643667Z digest=sha256:ffe6dc9db8bb2e9f64510cb2649ebac07bdc1fefd82cf7c5a83b753d142eb612

Observation 19d61ca9-241b-4520-8130-36a7f28b563c · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Value-Based Deep RL Scales Predictably

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:09:12.718570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T23:04:12.943222Z digest=sha256:883020961301cf817ee368980a587b8feda689446421899c54516d099146a08c