Pith. sign in

Paper Citation Record · LEDGER

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 3 inbound Pith citation observations for arXiv:2505.20268.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20268 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 3 of 3 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T22:06:24.412259Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:39:03.174419Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0b754f7a-d39b-4cc8-9df6-fcd65096b85b · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.276562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:10b53132fef107f876a30f39553f6d389b09e07a417f33dad0ae823c8bb8fafe

Observation c0716695-dd2f-478e-96b2-7f034332ea94 · inbound

Towards Differentially Private Reinforcement Learning with General Function Approximation cites this paper.

Towards Differentially Private Reinforcement Learning with General Function Approximation Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:31:00.541325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:15:59.533563Z digest=sha256:c4c2d910170dea3a4a62fc05562ba019b4dd78a0cf8f6ab97dbd7e3f1e33d60d

Observation 0ca2e471-63fd-473e-96bb-80f75fc32575 · inbound

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? cites this paper.

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:03.177721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T22:06:24.412259Z digest=sha256:1b44205e28c4cdeabaa7932125c9177a7ccc5d56fcefbba9bf3104aa9f8528f8