Pith. sign in

Paper Citation Record · LEDGER

Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2403.03950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.03950 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:32:44.345803Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 68eb5b5f-310f-4f76-8ad4-8c2f50cec93c · inbound

Chronos: Learning the Language of Time Series cites this paper.

Chronos: Learning the Language of Time Series Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T08:27:23.487475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T08:27:23.298009Z digest=sha256:c68a2004056f9644b45e1ee0b227ef467834ab75d87538856ff7d1c12ab31134

Observation 829d6796-a05e-4a0b-8d27-9a4f3b1b4acf · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:04:10.485375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:7f31c1c809d11dc47d62e8e13aece663106e7c496ac3539dea06b15e79ba3897

Observation 57709a0f-590d-45ac-9b61-08887d3f6b20 · inbound

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior cites this paper.

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.424708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T02:38:26.196000Z digest=sha256:08ee5aca419208e7384500242225543e4b2bd31a89cbaae0bc9f31a75fc7c1be

Observation fa4b0738-8482-4e9b-94ea-f186d525c781 · inbound

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior cites this paper.

Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:55:33.154530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T07:53:48.604436Z digest=sha256:e070a42ce5c2279c5bd88eb932af0d9c3e14025e8437507b155a5ae1b5a0b105

Observation 5d9a0d03-ad46-430e-a542-e611bd7328eb · inbound

D2 Actor Critic: Diffusion Actor Meets Distributional Critic cites this paper.

D2 Actor Critic: Diffusion Actor Meets Distributional Critic Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T07:35:29.400283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T07:31:27.330951Z digest=sha256:a77fe7bb2447d46e92c6a497ea5e55cc134f64b1a3c60b6409770ea3c4ad161c

Observation 544e6d5b-cf72-4159-be1f-56330955559d · inbound

Value Flows cites this paper.

Value Flows Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:28.459658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:28.459658Z digest=sha256:f44882f1b1d5a4a6e383572287a8d8c666810913500485f12b57a41242e37df4

Observation 2e14a882-24c4-442f-8c64-ba290c397f12 · inbound

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning cites this paper.

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 843

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:49.758326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:49.758326Z digest=sha256:ec8d24d784cd4cf9f2623ad0b7b30b0a2d3c439abec34f83af4ed3b48d1d7f07

Observation 94690d6b-e86f-467c-b64f-00310c11c066 · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:37:13.913128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T03:36:09.272019Z digest=sha256:4fc6af1729e24d420ea6b5bd2f6d82e284cffbeec611bb6b9a8e166a732ee8b4

Observation 94defd6d-5c96-4f7e-98bd-d0d42e4ee0a4 · inbound

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows cites this paper.

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T02:53:15.129472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:53:15.129472Z digest=sha256:7e6b748ebc1d011d251ab4e73b681d8142afd5bed08aea1374313aa6d6ac8f2c

Observation d3fc7f7f-8fe7-4ec1-a38f-88e78e34dfe7 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:36:17.834646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:b0e2fd7ab9eb610573ddbfb5b98e752d471ddbfc4599e8fa5e6d8d6e294b7f05

Observation 3bc30c35-a270-4039-aac2-5cb11aaa1904 · inbound

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning cites this paper.

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:46:27.856199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:55:22.577012Z digest=sha256:bb26089f1871605d5518c6dfc24f34a377f922dbfbc8a4377f25b0d1744b88b4

Observation b1523808-7122-45ad-97e8-60040f337875 · inbound

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity cites this paper.

Distributional Value Estimation Without Target Networks for Robust Quality-Diversity Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:14:46.858264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T00:11:04.222842Z digest=sha256:b361697904f524ce81faac3d61c748f62e48c07500d9dcabb23a0be84952204a

Observation d4073d5f-6457-45fb-bc1c-42240b54a772 · inbound

Hierarchical Behaviour Spaces cites this paper.

Hierarchical Behaviour Spaces Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:01:12.805850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:33:48.527621Z digest=sha256:80b32730986cfad364fcc413e727143922c5dc7e575f1e9081a3c9aa717096bb

Observation 83ddd3f9-3c3d-4670-a4d4-9280249664c5 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:00.701398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:b5e4faac8db26b8ed44e59ab532b8ad271889723752a20e3baefcc074f875768

Observation ac178a96-50c7-432e-87f1-73fda8433b49 · inbound

Survival Reinforcement Learning: Toward Scalable Self-Supervised RL cites this paper.

Survival Reinforcement Learning: Toward Scalable Self-Supervised RL Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:22:46.545028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T23:22:27.775787Z digest=sha256:90a27ac45930b9229bcc07e39a42055d3fd980e02f7644cc6011caa61894dff0

Observation 3ae4e93a-643f-4d7d-9153-0cd30820f91e · inbound

Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning cites this paper.

Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:31.233703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:12:00.111067Z digest=sha256:2b56af65985ec775b90e4a186a4bff91601d4a0d70b599aae505b01393d39e05

Observation 5a0a0036-5fe6-4d6a-86ca-3f9a6224d9b1 · inbound

Superhuman AI for Generals.io Using Self-Play Reinforcement Learning cites this paper.

Superhuman AI for Generals.io Using Self-Play Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:45.477702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T09:03:49.551408Z digest=sha256:c2dee8b901ce2be90a6571d735b584ac842f23b5e3426f8cf24fb0d77e86beb0

Observation e00edf4c-2546-4721-8b4c-d7689ed94a6a · inbound

World Value Models for Robotic Manipulation cites this paper.

World Value Models for Robotic Manipulation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:40:00.236783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T23:33:50.854261Z digest=sha256:9f6583c07284680e593110a900a31a03e55c0a49c43a136d9670cec6062d102a

Observation 35f039f2-de60-4110-8433-5dad67dea679 · inbound

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation cites this paper.

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:03:05.132513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T03:55:47.911210Z digest=sha256:17dc4ffeab492496b0e358c6b8466e55881ec58c574fa9690ca57f92c2f251d6

Observation 6640e43d-9c1e-471c-b3fb-589ad26aba89 · inbound

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation cites this paper.

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T11:25:18.198611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:25:18.198611Z digest=sha256:1e510ff986119b48c001139c4ba018fc5afb2a2f014c290429d9d3cca3a6a690

Observation 55b73d7d-5c57-4140-a3a8-1e9cf64de5e8 · inbound

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation cites this paper.

WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T07:17:21.581451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:17:21.581451Z digest=sha256:0a0bc76bcdfc175431cb2ad2d10558bc33f0a5c2b0207568e1bab727241d7ac0

Observation 703baa5e-1d4f-4cce-8642-aff5faab71a4 · inbound

Relative Value Learning cites this paper.

Relative Value Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T08:32:00.021033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:32:00.021033Z digest=sha256:51069880e75b9520dec7da73b30f288e5981b77b0f5e0461d0e48101f92528db

Observation 7d547984-a1a0-4dc0-995e-8ca771626df5 · inbound

ReBRAC-v2: The Return of the King cites this paper.

ReBRAC-v2: The Return of the King Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T00:32:44.345803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:32:44.345803Z digest=sha256:f4c3fd45e764c3fa6a12e2df25da0cb562676a70976d08b759ea71592087e051