Pith. sign in

Paper Citation Record · LEDGER

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning

As of 6 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 1 inbound Pith citation observation for arXiv:2601.23075.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.23075 v2

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:26:50.470669Z

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T02:55:22.577012Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T12:46:27.898232Z

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96a2db2e-3d93-434d-9532-4c0d803f1999 · outbound

This paper cites mean).Let πc(a|s) =N(µ,Σ) with fixed diagonal Σ = Diag(σ2) and parameter µ∈R m.

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning mean).Let πc(a|s) =N(µ,Σ) with fixed diagonal Σ = Diag(σ2) and parameter µ∈R m

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:50.403340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:50.403340Z digest=sha256:89b413fd1dc84fdb71f6fcf57e29e8b11aec1e1652d71805b32590a093bc9bad

Observation 3c458954-539e-488f-a63b-79c5db8e1de3 · outbound

This paper cites logits).Consider one action dimension i with logits zi ∈R K and softmax probabilities pi = softmax(zi)∈∆ K−1.

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning logits).Consider one action dimension i with logits zi ∈R K and softmax probabilities pi = softmax(zi)∈∆ K−1

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:50.470669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:50.470669Z digest=sha256:8212eff848d0d0d5e0624b5bf652b04cb70a3a9802dc55722c88af7e79940acb

Observation 789c6ba2-c50b-4d84-b762-c08c2fed956e · outbound

This paper cites SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning.

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning

Reference 456

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:49.906939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:49.906939Z digest=sha256:7b6ada03b4665d5876a46dad2091d1a8c3f7ec9c85e71e09dd5231d252123ab4

Observation 2e14a882-24c4-442f-8c64-ba290c397f12 · outbound

This paper cites Stop Regressing: Training Value Functions via Classification for Scalable Deep RL.

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 843

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:49.758326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:49.758326Z digest=sha256:ec8d24d784cd4cf9f2623ad0b7b30b0a2d3c439abec34f83af4ed3b48d1d7f07

Observation 0b865dce-dccd-424b-8539-56b73b144627 · outbound

This paper cites Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners.

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 1937

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:50.060301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:50.060301Z digest=sha256:160147271b5ba32ebe360e0e23d26e420101c49e326be31d8f837b50ffd29bda

Observation 4e0da62e-0158-4e00-bb8b-2e0aa0c9eeb8 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:50.223064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:50.223064Z digest=sha256:7c07c47f9d9a98f6482a1064c7ffaa10fcb01c86b41ba5b7539f1900789248cc

Pith citing papers

Observation 4cd33387-6be2-4374-aaf0-7d4ee76a0ec4 · inbound

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning cites this paper.

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-25T01:17:49.776384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:55:22.577012Z digest=sha256:c9713386d90984236255f85b683b5ad471cfd2ba50af87191399fff49b08ad0f