Pith. sign in

Paper Citation Record · LEDGER

Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2403.00514.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.00514 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:51.075651Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T13:01:23.635839Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 19548d37-4bec-43bc-af16-66f226479d82 · inbound

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners cites this paper.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.075651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.075651Z digest=sha256:faf220973c0dd6f154cef95055ef88c48cc65056d5e62f3154d113986be572cb

Observation 8fd96c6b-db8e-4849-a3f8-57a82ec57793 · inbound

Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments cites this paper.

Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:10:03.917362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:10:03.917362Z digest=sha256:f30370b931dda8e8c3646585e105c69a71d71de7bf2dea99fb0596d06c23c74f

Observation 77e2d3d5-0248-4d0c-98b1-276b66674e6a · inbound

Ensemble Elastic DQN: A Step Dependent Ensemble Approach for Reducing Overestimation in Deep Value-Based Reinforcement Learning cites this paper.

Ensemble Elastic DQN: A Step Dependent Ensemble Approach for Reducing Overestimation in Deep Value-Based Reinforcement Learning Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:49.734756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:18:49.734756Z digest=sha256:e3fa7d9fe7f94ad9c26235e461cebcbff065a01db1c4a75489bca94295dc9ed0

Observation e48555d3-1a4f-4afc-9fb0-98007b4db202 · inbound

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control cites this paper.

A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:07.829896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:07.829896Z digest=sha256:02af2ede5096974a0300569ba4baecb5ad85c3fec55e566eeef16027f9596550

Observation e7e28391-58e3-45e9-8de3-f0ae950b80ef · inbound

On the Effect of Regularization in Policy Mirror Descent cites this paper.

On the Effect of Regularization in Policy Mirror Descent Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:46.737730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:18:46.737730Z digest=sha256:8a6cc25aa02a2d5f4a707e4720eaf48b1415c94a92aba64aafaca9a11e402015

Observation 3b7ccadc-2df1-43f7-987f-f5286c24b362 · inbound

Balancing Expressivity and Robustness: Constrained Rational Activations for Reinforcement Learning cites this paper.

Balancing Expressivity and Robustness: Constrained Rational Activations for Reinforcement Learning Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:35.644714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:55:35.644714Z digest=sha256:9c5e6a1364696b381b103f10564f1191c13d33778b79522dde1694fefd5b4c46

Observation 5fc359b0-f3cd-4b89-b0ec-b6118a569801 · inbound

Activation Function Design Sustains Plasticity in Continual Learning cites this paper.

Activation Function Design Sustains Plasticity in Continual Learning Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:23.638555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T13:00:27.749673Z digest=sha256:eb5ff46b38cf22459948a69db256af9d4d91f9f351d452193494597a1dc97caf

Observation 90130f14-8157-4070-aee3-50ebfdb0a4d1 · inbound

Forager: a lightweight testbed for continual learning with partial observability in RL cites this paper.

Forager: a lightweight testbed for continual learning with partial observability in RL Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:46:53.453377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:16:47.462346Z digest=sha256:6c3a645a6f20d496bf5a8af25f28b2b847ed3fa360de44b2ec5b95eaac7f1e16

Observation d6fde913-fe6b-47d7-b25d-6c1a436a441f · inbound

Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning cites this paper.

Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:51:17.033858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:51:17.033858Z digest=sha256:8e15be38087c494b9ecf5ab82dea5d42d9cd71f9f292c19c922d4a1d928cf37d

Observation 6af8d223-f4fb-493e-b5ab-f11e853aa1d2 · inbound

Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning cites this paper.

Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-07-31T04:05:50.150983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T04:05:50.150983Z digest=sha256:fb87499fe749ac2434e0939ca429249f12fc944798dd5259949f494a8648e15f