Pith. sign in

Paper Citation Record · LEDGER

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2502.06781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06781 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:40:24.891449Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 40a1032b-0f5d-4e52-8a77-687f5ff1de13 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.345899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:83df67eed44eb2bfac35913e406a6ab791505b5f57b1347a60cdad862240f756

Observation da937472-cdc1-490a-beb9-4b79cd4fbc57 · inbound

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving cites this paper.

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T06:40:24.891449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:40:24.891449Z digest=sha256:bfafb18d264885c7479c8a72a86b091a08a1ce0db5d6462dcf2f8bca16fad379

Observation cc15766a-f859-4eaa-9b3b-82060a08db9c · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.981554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.981554Z digest=sha256:7771cf116d56c36a267967f2475d5836e204a075b4735c69887f30fa8249a360

Observation 6e7d4cec-4529-43b9-bac8-67b44a0416a9 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 224

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.534551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:c36272ee177c0d44a77c9a62de5402772386400f0a27759a0bb4558d9372cf35

Observation 7bc710db-20a8-4c95-a0e0-c252eb3cb39f · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 165

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:00:56.109782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:c2954c6216be272060860185e5e59d763a411dfc60a42511bb265b0144a50ee3