Pith. sign in

Paper Citation Record · LEDGER

Offline Reinforcement Learning for LLM Multi-Step Reasoning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2412.16145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16145 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:20:13.837391Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5f1be69d-e707-456a-8242-d80ec7109be4 · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.837391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.837391Z digest=sha256:a7492fa604188251d8f7279d87806c2b306a964eea4ac2c3bbcd130626ec9c7c

Observation 3b599f16-a305-4fc8-9721-66e3ce66ec3c · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.436669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.436669Z digest=sha256:1b4a2c18c0f6e52ba6e6ffb509006d67a6999326f44139bccdfd9784c10aebfe

Observation 40d8a2c3-061f-4d2d-ab59-133b144a810e · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.615635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.615635Z digest=sha256:13aea5cbe4e7545d367ff44cae86e59d4c79f8cc6064573c2ce38677a3cd2a7d

Observation 3c787b02-2301-4e54-b8de-08a800880ea4 · inbound

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning cites this paper.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.280984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.280984Z digest=sha256:023080d1d0c8dfb8cacdebf180a7219eecaeba7defd7691ada33460f18007a58

Observation 1ebdf16a-4c76-47fc-970b-4ff90b9ad66e · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:37.039893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:37.039893Z digest=sha256:bdf3e08c27754feae0c44b269a7bda560245d7524a0b4af8688286a73b80a532

Observation a9ca7058-d9b0-4a0d-ba6d-d67f668e1201 · inbound

Think Clearly: Improving Reasoning via Redundant Token Pruning cites this paper.

Think Clearly: Improving Reasoning via Redundant Token Pruning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.742674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.742674Z digest=sha256:f4b98a1b0b3befaa738bf85864f90e438c19e546be4db445c495026190737279

Observation 7e0152b2-7e3e-4348-be31-8bd04f407e02 · inbound

Reinforcement Learning in hyperbolic space for multi-step reasoning cites this paper.

Reinforcement Learning in hyperbolic space for multi-step reasoning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:27.704354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:27.704354Z digest=sha256:89713ad0728c1d9d5fee9001b2ad84581cd84846c6b6daf8e7aadaa1a182d268

Observation a2612bc7-e4b2-45f7-82e7-425310cab47c · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.334653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:366f816ad3db907544ad112446da91e844aebf69323e224a947ecf845af4bbad

Observation 2c55574d-7af5-489b-80b5-00c7710b88e2 · inbound

Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya cites this paper.

Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:00:21.063340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T21:58:40.973772Z digest=sha256:80fb3b3c0f07e6c1df8c1249f66a434404e58314ba023434d8e434028c7f12a0

Observation 6440f138-b617-4260-9f74-7f4622f8fb3b · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.653343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:62ac5d85cef86984ec2ddcbf7976d62e6178fc4003fddb2008d6e5a524b920f4

Observation 6583c324-0b76-40b0-b47f-248c6f4e67b0 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 240

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T13:00:56.038173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:b01d0a334eab5535dcad60339ed0db3c67adba1579bb0f795896d48dc2e544e4

Observation 4f89117f-84fb-4cf7-b597-b2bc031fceba · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 212

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.096925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:9c77fdbb66f89f717ee155c0a17b13ce6205506132cad8034da4c2233482c7e8

Observation 14a4ff1a-7006-4230-be33-2d08996d25f1 · inbound

LeAct: Learning to Reason from Expert Actions cites this paper.

LeAct: Learning to Reason from Expert Actions Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T06:32:23.624774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:32:23.624774Z digest=sha256:85380a50cb5a24a56531e6efb71c9d8a8b064ee6b41179c716c195d64c29b072

Observation 8b8a284a-2592-458f-a08e-8596a8c151ca · inbound

CRAFT: Learn the Schema, Execute the Plan cites this paper.

CRAFT: Learn the Schema, Execute the Plan Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T10:18:39.144202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:18:39.144202Z digest=sha256:85925a40be86272e90a3e9dcf73014fc17e53d71e26c2bc3c70b9e723d8860d4