Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:56:22.548936Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2412.18279.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:56:22.548936Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:28:58.286475Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T22:20:49.395329Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 47f27bd7-943a-4824-8900-f8c17c71916e · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Kakade, Jason D
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ef117a35-1d9e-497b-8b5b-f445757d78ad · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0c0ec4a2-7d79-49e2-8452-275762ab279b · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Enhancing textual textbook question answering with large language models and retrieval augmented generation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c47b805-0b1f-487b-a9b9-bba92ad26948 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Program Synthesis with Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c4c37a-24b1-48ac-8771-73cdd995ceb5 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1765773-f787-4b03-a31c-4fa9e89d699c · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 99cd411f-1c7d-46c9-a656-002b7db112ae · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Evaluating Large Language Models Trained on Code
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5337f7a0-db0f-4db6-810a-0c2116f91ab6 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Teaching Large Language Models to Self-Debug
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c27dda19-e679-4c18-9591-2be4b11cb8c1 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Christiano, Jan Leike, Tom B
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 55f5ec46-89e8-42f3-b854-8b03b886ca60 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Training Verifiers to Solve Math Word Problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2d7ee51-ec5b-4c3e-90dd-60c61f6c5f97 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization RAFT: reward ranked finetuning for generative foundation model alignment
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5840476f-026e-4cfc-886f-316930502ecf · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Addressing function approximation error in actor- critic methods
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d79406e0-fc3d-4b9e-b4a1-e2be2aeeb154 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization On Designing Effective RL Reward at Training Time for LLM Reasoning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60a9534b-3660-473d-bc99-93e66693ff89 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Step-level Value Preference Optimization for Mathematical Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09f37b5b-361d-44c2-9c7f-82f495636f97 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 80ef1a1f-0d61-4c7b-b409-ce1a98a8d3b2 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization PSYDIAL: Personality-based Synthetic Dialogue Generation using Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95cf2eee-ae1f-4735-9b12-c788d09b6494 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eba2264f-08e8-43ee-b19d-ac7f41ad7f62 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Measuring mathematical problem solving with the MATH dataset
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce6134c-d52e-4ad5-9ed0-bca3103a838c · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e019d461-189d-42f9-a85d-d39da7808a20 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc6fcc9b-15a1-4f2f-a100-8df140d98083 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 076dad93-3150-45f2-97f7-fa92d32b0265 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 634a185a-7fa5-4cf9-9d14-268eb4571d06 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Deep Reinforcement Learning and the Deadly Triad
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7919dcf5-3e18-4426-88e6-46a6140037a1 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Ramasesh, AmbroseSlone, CemAnil, ImanolSchlag, TheoGutman-Solo, YuhuaiWu, BehnamNeyshabur, Guy Gur-Ari, and Vedant Misra
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2ae719b7-e628-427b-b802-9ddf63d8f4d6 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization TACO: Topics in Algorithmic COde generation dataset
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b19d65-81c5-4dd2-b439-c1dfedd28cb4 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Remax: A simple, effective, and efficient reinforcement learning method for aligning large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2b1f2dc6-3753-4552-a39a-932993fef074 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization On the Linear Convergence of Policy Gradient under Hadamard Parameterization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a9b96a0d-0917-453d-a1c3-81d6901e0bda · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization On the Convergence of Projected Policy Gradient for Any Constant Step Sizes
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53de0872-a891-4fab-9ada-976fdf1488b8 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Elementary Analysis of Policy Gradient Methods
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e31126b9-e4bd-4a0d-96cd-7dd6e9fbb037 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 852b9206-1902-4753-8ec6-d59310ff5c86 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Leveraging non-uniformity in first-order non-convex optimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ec528c7f-09e1-4b22-8040-7c2b3297a825 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization On the global convergence rates of softmax policy gradient methods
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 09fe81e4-c924-4628-9635-d61c9acecac8 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Skywork-o1 open series.https://huggingface.co/Skywork, November 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 89f9c63b-54dd-48ed-bdb1-11d45a586f9d · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Manning, Stefano Ermon, and Chelsea Finn
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 621e3870-a4e6-4683-b07d-6e202d0c3c6b · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Offline Regularised Reinforcement Learning for Large Language Models Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5175a9c0-ea33-4ca2-bb68-d8fda1743e83 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Van de Wiele, Vlad Mnih, Nicolas Heess, and Jost Tobias Springenberg
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7148fab6-237d-4722-9a0a-58ff2335d79c · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 667f61e7-6ec8-47ff-aa8e-eea2ea2da1c8 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Proximal Policy Optimization Algorithms
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c945365-0fb6-44fb-aec4-c5771848d22d · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af85cb5-331a-4a79-978b-a870a480247d · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Sutton, David McAllester, Satinder Singh, and Yishay Mansour
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5af334ff-287e-476f-921f-dd0f1a6ca420 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Mathscale: Scaling instruction tuning for mathematical reasoning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 337596f2-d442-4880-899e-b11c7c5a4c75 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Qwen2.5: A party of foundation models, September 2024
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94800a6a-203e-4372-a873-cb8be31f73a1 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization LLaMA: Open and Efficient Foundation Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eabccda1-291c-40b6-a596-f6d342763f57 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1cf4d63-f665-4fa2-bdaf-958ccaaaa325 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5ea1bca7-7164-4177-8a73-8898b12f94bb · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2d4b108a-5b06-4fc9-a75e-23e4fbb235d0 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Brown, and Ken Goldberg
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 37573d63-20d1-49e6-94eb-0d990791ef1a · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Qwen2 Technical Report
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08d47d4b-e6a4-4e4a-81fc-127386ad2ec4 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e2e0133-219c-409f-ad96-6553b1331efa · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e69b9c43-c12c-4c06-9502-eb7e9d24e244 · outbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13723c2e-df2d-4094-92f3-2598b454f8e0 · inbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b1cd13-a158-4d77-8c29-25dfaa2c2a60 · inbound
Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.