Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2403.03950.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:32:44.345803Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 68eb5b5f-310f-4f76-8ad4-8c2f50cec93c · inbound
Chronos: Learning the Language of Time Series Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 829d6796-a05e-4a0b-8d27-9a4f3b1b4acf · inbound
Training Language Models to Self-Correct via Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 57709a0f-590d-45ac-9b61-08887d3f6b20 · inbound
Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fa4b0738-8482-4e9b-94ea-f186d525c781 · inbound
Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d9a0d03-ad46-430e-a542-e611bd7328eb · inbound
D2 Actor Critic: Diffusion Actor Meets Distributional Critic Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 544e6d5b-cf72-4159-be1f-56330955559d · inbound
Value Flows Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e14a882-24c4-442f-8c64-ba290c397f12 · inbound
RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 843
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94690d6b-e86f-467c-b64f-00310c11c066 · inbound
SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 94defd6d-5c96-4f7e-98bd-d0d42e4ee0a4 · inbound
SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3fc7f7f-8fe7-4ec1-a38f-88e78e34dfe7 · inbound
What Does Flow Matching Bring To TD Learning? Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3bc30c35-a270-4039-aac2-5cb11aaa1904 · inbound
Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b1523808-7122-45ad-97e8-60040f337875 · inbound
Distributional Value Estimation Without Target Networks for Robust Quality-Diversity Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d4073d5f-6457-45fb-bc1c-42240b54a772 · inbound
Hierarchical Behaviour Spaces Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83ddd3f9-3c3d-4670-a4d4-9280249664c5 · inbound
Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac178a96-50c7-432e-87f1-73fda8433b49 · inbound
Survival Reinforcement Learning: Toward Scalable Self-Supervised RL Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ae4e93a-643f-4d7d-9153-0cd30820f91e · inbound
Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a0a0036-5fe6-4d6a-86ca-3f9a6224d9b1 · inbound
Superhuman AI for Generals.io Using Self-Play Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e00edf4c-2546-4721-8b4c-d7689ed94a6a · inbound
World Value Models for Robotic Manipulation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 35f039f2-de60-4110-8433-5dad67dea679 · inbound
WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6640e43d-9c1e-471c-b3fb-589ad26aba89 · inbound
WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55b73d7d-5c57-4140-a3a8-1e9cf64de5e8 · inbound
WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 703baa5e-1d4f-4cce-8642-aff5faab71a4 · inbound
Relative Value Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d547984-a1a0-4dc0-995e-8ca771626df5 · inbound
ReBRAC-v2: The Return of the King Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.