Pith. sign in

Paper Citation Record · LEDGER

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2412.18279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18279 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:56:22.548936Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:28:58.286475Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T22:20:49.395329Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47f27bd7-943a-4824-8900-f8c17c71916e · outbound

This paper cites Kakade, Jason D.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Kakade, Jason D

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.242182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.318078Z digest=sha256:43b04422d41c6605a173f374609ed183f0a76094defaa0e041b80dc2c8092fe7

Observation ef117a35-1d9e-497b-8b5b-f445757d78ad · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.230453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.323056Z digest=sha256:278b361bdfc34d242c9b4c4f6b0365c0e294fe695dd76e075fabff4a940fa5a0

Observation 0c0ec4a2-7d79-49e2-8452-275762ab279b · outbound

This paper cites Enhancing textual textbook question answering with large language models and retrieval augmented generation.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Enhancing textual textbook question answering with large language models and retrieval augmented generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.327850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.327850Z digest=sha256:a6c9952d1d3c31466dc194737ee74065f294a02654c39e2c40c673649bc073b8

Observation 0c47b805-0b1f-487b-a9b9-bba92ad26948 · outbound

This paper cites Program Synthesis with Large Language Models.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Program Synthesis with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.336962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.336962Z digest=sha256:22671a429d5881acef0a0b7a26c81bb9f23fd01f9a9e973503a6dac50dcc821c

Observation b8c4c37a-24b1-48ac-8771-73cdd995ceb5 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.340941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.340941Z digest=sha256:3a037e3717d3606e64a7f7719daebebb8ee6dad1583bda477240cbf250719056

Observation a1765773-f787-4b03-a31c-4fa9e89d699c · outbound

This paper cites an unresolved cited work.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:56:23.217534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.345546Z digest=sha256:d24564606cf34122d76d06aa6eb326fb5ef34999f18eab4621ec3eee81b6477a

Observation 99cd411f-1c7d-46c9-a656-002b7db112ae · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.349602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.349602Z digest=sha256:31c799b4609af5e41554c665f062ad1494f147085e8f2170b6de9909ffb9fedf

Observation 5337f7a0-db0f-4db6-810a-0c2116f91ab6 · outbound

This paper cites Teaching Large Language Models to Self-Debug.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Teaching Large Language Models to Self-Debug

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.353461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.353461Z digest=sha256:168a1c19faa7c71a4c0b37ccea479c484d5498fe9c939c67a1e7e60c5b3ca9e2

Observation c27dda19-e679-4c18-9591-2be4b11cb8c1 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Christiano, Jan Leike, Tom B

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.204999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.357548Z digest=sha256:8879d40af4837daffa21cb290200d08a8e2afd25029b272242e1822cafc5d251

Observation 55f5ec46-89e8-42f3-b854-8b03b886ca60 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.361804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.361804Z digest=sha256:320ef509f1f3245461a902411ff8663dda45dc4e9507b270360027abf40f56cc

Observation c2d7ee51-ec5b-4c3e-90dd-60c61f6c5f97 · outbound

This paper cites RAFT: reward ranked finetuning for generative foundation model alignment.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization RAFT: reward ranked finetuning for generative foundation model alignment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.192380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.365974Z digest=sha256:9a76e70897613c1ffc74366c18f5856846cf90f9d381e7d3fa4519a94b403222

Observation 5840476f-026e-4cfc-886f-316930502ecf · outbound

This paper cites Addressing function approximation error in actor- critic methods.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Addressing function approximation error in actor- critic methods

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.179569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.369853Z digest=sha256:be4082ffa858eeb541656838474e09e4d2cb1e819c35cf411d3728877310599f

Observation d79406e0-fc3d-4b9e-b4a1-e2be2aeeb154 · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.373566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.373566Z digest=sha256:75620bcb8f24164e9f7cf0b6b923f9b8d2d3010e253b92c854453929f7971e42

Observation 60a9534b-3660-473d-bc99-93e66693ff89 · outbound

This paper cites Step-level Value Preference Optimization for Mathematical Reasoning.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Step-level Value Preference Optimization for Mathematical Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.377628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.377628Z digest=sha256:b96607cc2e6ec1c5fcc2e7fba6e92053797cac656c786c9d7c9fe8a0c1ff3f9c

Observation 09f37b5b-361d-44c2-9c7f-82f495636f97 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.166167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.382182Z digest=sha256:47d5a36b0ae89a308658b42c4f012a2a350becaaaf8b360919c0764daf43e871

Observation 80ef1a1f-0d61-4c7b-b409-ce1a98a8d3b2 · outbound

This paper cites PSYDIAL: Personality-based Synthetic Dialogue Generation using Large Language Models.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization PSYDIAL: Personality-based Synthetic Dialogue Generation using Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.386035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.386035Z digest=sha256:537c580c87b5ea62a5fe521830086b539780dd504a006fe751743983394db01c

Observation 95cf2eee-ae1f-4735-9b12-c788d09b6494 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.152947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.390301Z digest=sha256:4db0ce301619b146d39680721540fce01e02897dd63dfc2e753cb59e0f9f0831

Observation eba2264f-08e8-43ee-b19d-ac7f41ad7f62 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Measuring mathematical problem solving with the MATH dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.395440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.395440Z digest=sha256:4f870a40c204ba2ffabffef2b2d25bedd7d74439b08177989dd956d1f171e267

Observation bce6134c-d52e-4ad5-9ed0-bca3103a838c · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.400256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.400256Z digest=sha256:0f55deb3cfb3805138b83c6ffcdf7492600b2a82bdf7f9039ba91030f299a7c2

Observation e019d461-189d-42f9-a85d-d39da7808a20 · outbound

This paper cites Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.405941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.405941Z digest=sha256:83e9e0acf553f678a2bb86ef94c68ccd989c154af40b6b33d4830084c3fe6365

Observation bc6fcc9b-15a1-4f2f-a100-8df140d98083 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.410383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.410383Z digest=sha256:08cf1b17085b4bb44ab33130df145b53f8ecf6d143171aa8531caef711df1b8c

Observation 076dad93-3150-45f2-97f7-fa92d32b0265 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.417074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.417074Z digest=sha256:6268ae18876c950d0e3cf998c0f2ad903a05d085f2ac8499bf8f53207ae46e67

Observation 634a185a-7fa5-4cf9-9d14-268eb4571d06 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Deep Reinforcement Learning and the Deadly Triad

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.421332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.421332Z digest=sha256:0e74771c8958110572f26ad97916f77415581364f6d550001636233d6c0a7c4a

Observation 7919dcf5-3e18-4426-88e6-46a6140037a1 · outbound

This paper cites Ramasesh, AmbroseSlone, CemAnil, ImanolSchlag, TheoGutman-Solo, YuhuaiWu, BehnamNeyshabur, Guy Gur-Ari, and Vedant Misra.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Ramasesh, AmbroseSlone, CemAnil, ImanolSchlag, TheoGutman-Solo, YuhuaiWu, BehnamNeyshabur, Guy Gur-Ari, and Vedant Misra

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.132145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.426437Z digest=sha256:16cb8c444063c40768961c11211fef009780a39cd9586ead1e3b80b1c11b7a22

Observation 2ae719b7-e628-427b-b802-9ddf63d8f4d6 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization TACO: Topics in Algorithmic COde generation dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.434443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.434443Z digest=sha256:bafa2124b5f796752d4c6e994e310a60cf21a09db221b770e22266edff1bd514

Observation 76b19d65-81c5-4dd2-b439-c1dfedd28cb4 · outbound

This paper cites Remax: A simple, effective, and efficient reinforcement learning method for aligning large language models.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Remax: A simple, effective, and efficient reinforcement learning method for aligning large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.113005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.439683Z digest=sha256:4f50658e0f3694702bb402e4f3efc25d3763e9e0343950e75a7d07ab6385c9f4

Observation 2b1f2dc6-3753-4552-a39a-932993fef074 · outbound

This paper cites On the Linear Convergence of Policy Gradient under Hadamard Parameterization.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization On the Linear Convergence of Policy Gradient under Hadamard Parameterization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:56:22.741664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.444288Z digest=sha256:4fda2f7afb2fe43848c8e98a6c34cbcc9605ceb527aaeca4a25aae8c8255f41f

Observation a9b96a0d-0917-453d-a1c3-81d6901e0bda · outbound

This paper cites On the Convergence of Projected Policy Gradient for Any Constant Step Sizes.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization On the Convergence of Projected Policy Gradient for Any Constant Step Sizes

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.449036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.449036Z digest=sha256:f4bcec4fcaf57d57a9972b1b64cfa32b38f5b27f85170cff467e70a2aceb9cd8

Observation 53de0872-a891-4fab-9ada-976fdf1488b8 · outbound

This paper cites Elementary Analysis of Policy Gradient Methods.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Elementary Analysis of Policy Gradient Methods

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.453068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.453068Z digest=sha256:f9a0e84ea052caabf11461b6e2cf96ad010f62971b52ef6711eeb66fa0f2ae4f

Observation e31126b9-e4bd-4a0d-96cd-7dd6e9fbb037 · outbound

This paper cites Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.457607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.457607Z digest=sha256:5e2b3fd087a05007e9c34888d5fa9e55d59bb2df8b1ae76eeb2cbc34af82bfbd

Observation 852b9206-1902-4753-8ec6-d59310ff5c86 · outbound

This paper cites Leveraging non-uniformity in first-order non-convex optimization.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Leveraging non-uniformity in first-order non-convex optimization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.099144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.462266Z digest=sha256:3cad5016730c37f96daa5af7c7baabfb6cab732f379c0ee69cae6ea4997c645f

Observation ec528c7f-09e1-4b22-8040-7c2b3297a825 · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization On the global convergence rates of softmax policy gradient methods

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.080227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.467246Z digest=sha256:80b17b356339e0131d748abb58df3841a71189f2e70f735b150375288ccfad41

Observation 09fe81e4-c924-4628-9635-d61c9acecac8 · outbound

This paper cites Skywork-o1 open series.https://huggingface.co/Skywork, November 2024.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Skywork-o1 open series.https://huggingface.co/Skywork, November 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.066134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.472118Z digest=sha256:c2d777c80801919c602cfd71fc07ca2687dee9b36c441b2fb68d23f97fe88eb5

Observation 89f9c63b-54dd-48ed-bdb1-11d45a586f9d · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Manning, Stefano Ermon, and Chelsea Finn

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.054411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.476435Z digest=sha256:3815a26ca7a02e91f39f6c6508f938cd58e259878c9e8bc2c9422e5b412427c1

Observation 621e3870-a4e6-4683-b07d-6e202d0c3c6b · outbound

This paper cites Offline Regularised Reinforcement Learning for Large Language Models Alignment.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Offline Regularised Reinforcement Learning for Large Language Models Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.480620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.480620Z digest=sha256:8fdac493645f23118a179c9b9543fe0ded8e5de3a078c750c102e3d3b008cc97

Observation 5175a9c0-ea33-4ca2-bb68-d8fda1743e83 · outbound

This paper cites Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Van de Wiele, Vlad Mnih, Nicolas Heess, and Jost Tobias Springenberg.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Van de Wiele, Vlad Mnih, Nicolas Heess, and Jost Tobias Springenberg

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.042722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.485473Z digest=sha256:5435d6155ebd4d1fc4be4c15683d788ebf87d24fc7cdcc8914a124bb4c5f12f7

Observation 7148fab6-237d-4722-9a0a-58ff2335d79c · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.491649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.491649Z digest=sha256:9b0a6c3ed16dc9f03bc9c53923d695e4768e7f9027316c444bb097a1c6568bdc

Observation 667f61e7-6ec8-47ff-aa8e-eea2ea2da1c8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.496036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.496036Z digest=sha256:2bf775d5aba4002ccda8baf741f960d4cd5b8f4834c56736b51a6191aef142ec

Observation 4c945365-0fb6-44fb-aec4-c5771848d22d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.500768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.500768Z digest=sha256:d43f989095240693e2325a8798da8001ff2a1c2e25d56d104a1b4e8924058bcf

Observation 4af85cb5-331a-4a79-978b-a870a480247d · outbound

This paper cites Sutton, David McAllester, Satinder Singh, and Yishay Mansour.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Sutton, David McAllester, Satinder Singh, and Yishay Mansour

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.028279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.504835Z digest=sha256:75aeb0a13df3c5f0ce2cfd4473029c4c6b2affa6fefc4e2f6e22d941e484b019

Observation 5af334ff-287e-476f-921f-dd0f1a6ca420 · outbound

This paper cites Mathscale: Scaling instruction tuning for mathematical reasoning.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Mathscale: Scaling instruction tuning for mathematical reasoning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:23.014407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.508855Z digest=sha256:d9be553937622b3e78a92bc5b4522ba6a3a98ed5c7f0b2128b4d1e86d2856040

Observation 337596f2-d442-4880-899e-b11c7c5a4c75 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Qwen2.5: A party of foundation models, September 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.513048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.513048Z digest=sha256:42773461bda0ab4737582809c92c57059e0b0d89501307c7e8ebad2ed46b1dfc

Observation 94800a6a-203e-4372-a873-cb8be31f73a1 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.517239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.517239Z digest=sha256:905a81bdb88bc26b365c41e089f39804374cfa391ca240a2df28522a8c947032

Observation eabccda1-291c-40b6-a596-f6d342763f57 · outbound

This paper cites Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.521338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.521338Z digest=sha256:412f75f30aafbf4b015e42fde22dfe7c6556bef7efc7a3ac31e687ace5c2c0a3

Observation e1cf4d63-f665-4fa2-bdaf-958ccaaaa325 · outbound

This paper cites Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:22.991257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.525884Z digest=sha256:966b2689bc339c3c50d8ed1e0362ec083a0f1ac92cedf98274ce5d70a9158047

Observation 5ea1bca7-7164-4177-8a73-8898b12f94bb · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:22.976182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.529361Z digest=sha256:7087b6c1889af00a3a1813f53e879637ae597714731d27335e0ebe9b21a76c2c

Observation 2d4b108a-5b06-4fc9-a75e-23e4fbb235d0 · outbound

This paper cites Brown, and Ken Goldberg.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Brown, and Ken Goldberg

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:22.962709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.533140Z digest=sha256:69ead11f897e05d1bd5dbb679a359b79f12cb9e93a9014cb2b1870ddc691bb4d

Observation 37573d63-20d1-49e6-94eb-0d990791ef1a · outbound

This paper cites Qwen2 Technical Report.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Qwen2 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.536918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.536918Z digest=sha256:395a62c81fb3ec1b5b1404872ea51beae5b2fff4e7f09e25a30d765b3d217edf

Observation 08d47d4b-e6a4-4e4a-81fc-127386ad2ec4 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.540940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.540940Z digest=sha256:0c9f4ed2bfaccba9d514337d46f83e521f2ee2226835fb0cbb334aea3618e130

Observation 8e2e0133-219c-409f-ad96-6553b1331efa · outbound

This paper cites Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:56:22.949709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:56:22.545285Z digest=sha256:68ccd90dc930c050d639d557fc9469592a330175bc6eaa27419b3057605ed031

Observation e69b9c43-c12c-4c06-9502-eb7e9d24e244 · outbound

This paper cites Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.548936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.548936Z digest=sha256:f4bd76028fb180babe1d5205ca2805672289e471cc62aa4f67c670e284d41b35

Pith citing papers

Observation 13723c2e-df2d-4094-92f3-2598b454f8e0 · inbound

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents cites this paper.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.286475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.286475Z digest=sha256:0baefb00b6f79309122fb333a6c834123171fae68663caa6f2134eaca5b438a6

Observation e8b1cd13-a158-4d77-8c29-25dfaa2c2a60 · inbound

Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning cites this paper.

Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:49.396977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:55:17.077281Z digest=sha256:fbf5452fcf579adb0ce8482ab8cf7f5d25113054363c49c2dacfc381e801cfe0