Pith. sign in

Paper Citation Record · LEDGER

Offline Reinforcement Learning for LLM Multi-Step Reasoning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2412.16145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16145 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:30.615635Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 40d8a2c3-061f-4d2d-ab59-133b144a810e · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.615635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.615635Z digest=sha256:0cca8c04f15176c864e19eba635cac36a4e16982d228caac163159569b4574ac

Observation 3c787b02-2301-4e54-b8de-08a800880ea4 · inbound

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning cites this paper.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.280984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.280984Z digest=sha256:6a77662adab1bafa73ee2cceca82108ec7b9b12523df7edf44eb1c043c01e822

Observation 1ebdf16a-4c76-47fc-970b-4ff90b9ad66e · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:37.039893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:37.039893Z digest=sha256:5d0df7924247334b7547194c068e951e434a12a748e09b9a01830286d0d9cd97

Observation a9ca7058-d9b0-4a0d-ba6d-d67f668e1201 · inbound

Think Clearly: Improving Reasoning via Redundant Token Pruning cites this paper.

Think Clearly: Improving Reasoning via Redundant Token Pruning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.742674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.742674Z digest=sha256:7c23fc91bc6ed9077541a674c0bd0fc3b5290b295f9bd46d06f317410a1d0e40

Observation 7e0152b2-7e3e-4348-be31-8bd04f407e02 · inbound

Reinforcement Learning in hyperbolic space for multi-step reasoning cites this paper.

Reinforcement Learning in hyperbolic space for multi-step reasoning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:27.704354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:27.704354Z digest=sha256:89713ad0728c1d9d5fee9001b2ad84581cd84846c6b6daf8e7aadaa1a182d268

Observation a2612bc7-e4b2-45f7-82e7-425310cab47c · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.334653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:b8779f5a1c236e934717249959ccb2d9e5ef65043e8ecadf476385bdf98c5f8a

Observation 2c55574d-7af5-489b-80b5-00c7710b88e2 · inbound

Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya cites this paper.

Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:00:21.063340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T21:58:40.973772Z digest=sha256:f93166421af8ea8040cf281b71b35f4b70f36ec7f3239853dcf2f4ba3e86a248

Observation 6440f138-b617-4260-9f74-7f4622f8fb3b · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.653343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:91edd2355e821b35cb8ea8220dbaeda8fc2c22b53d0e879a6ab768f32e5f624e

Observation 6583c324-0b76-40b0-b47f-248c6f4e67b0 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 240

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T13:00:56.038173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:78a108373306b226297537117adad0f5117492cb4b1185c3afe8d3deae984c8b

Observation 4f89117f-84fb-4cf7-b597-b2bc031fceba · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 212

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.096925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:91e2fbd6f4c53d6f9bd8d82cb78b6a449a0d6694f7f0b8f7d1ff3b33d96ffe5a

Observation 14a4ff1a-7006-4230-be33-2d08996d25f1 · inbound

LeAct: Learning to Reason from Expert Actions cites this paper.

LeAct: Learning to Reason from Expert Actions Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T06:32:23.624774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:32:23.624774Z digest=sha256:85380a50cb5a24a56531e6efb71c9d8a8b064ee6b41179c716c195d64c29b072

Observation 8b8a284a-2592-458f-a08e-8596a8c151ca · inbound

CRAFT: Learn the Schema, Execute the Plan cites this paper.

CRAFT: Learn the Schema, Execute the Plan Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T10:18:39.144202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:18:39.144202Z digest=sha256:6f8bddde2a049f581747a16257f3a9fd748bdb02a660349d00230dd9a9c2148b