Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2412.16145.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:30.615635Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 40d8a2c3-061f-4d2d-ab59-133b144a810e · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c787b02-2301-4e54-b8de-08a800880ea4 · inbound
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ebdf16a-4c76-47fc-970b-4ff90b9ad66e · inbound
A Technical Survey of Reinforcement Learning Techniques for Large Language Models Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 132
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9ca7058-d9b0-4a0d-ba6d-d67f668e1201 · inbound
Think Clearly: Improving Reasoning via Redundant Token Pruning Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e0152b2-7e3e-4348-be31-8bd04f407e02 · inbound
Reinforcement Learning in hyperbolic space for multi-step reasoning Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2612bc7-e4b2-45f7-82e7-425310cab47c · inbound
On the optimization dynamics of RLVR: Gradient gap and step size thresholds Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c55574d-7af5-489b-80b5-00c7710b88e2 · inbound
Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6440f138-b617-4260-9f74-7f4622f8fb3b · inbound
Trust Region On-Policy Distillation Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 257
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6583c324-0b76-40b0-b47f-248c6f4e67b0 · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 240
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f89117f-84fb-4cf7-b597-b2bc031fceba · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 212
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14a4ff1a-7006-4230-be33-2d08996d25f1 · inbound
LeAct: Learning to Reason from Expert Actions Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b8a284a-2592-458f-a08e-8596a8c151ca · inbound
CRAFT: Learn the Schema, Execute the Plan Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.