Pith. sign in

Paper Citation Record · LEDGER

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

As of 21 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 92 inbound Pith citation observations for arXiv:2406.18629.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.18629 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T23:58:29.040819Z

measured 130 of 130 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 92 of 92 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:47:11.169539Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact35
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6cc6da4a-7f20-472c-93af-ce641a20af4b · outbound

This paper cites GPT-4 Technical Report.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.068590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:c55dc1b0f4c855d588a6a865cc59a1ca003b449925a02373398c3063aa6caebe

Observation 5f81d01a-12e0-41f1-9710-d4a4fd8518b2 · outbound

This paper cites Llemma: An Open Language Model For Mathematics.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Llemma: An Open Language Model For Mathematics

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:17:47.048289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:681bcbd3028787767c9dd4c51dfeff6842d4f2de485e9f2877d5579182e7da03

Observation 8911546b-1d35-42c6-abb5-1068da101f26 · outbound

This paper cites Qwen Technical Report.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Qwen Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.085943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:766f904152a1e38f9c12f11bce5603c771e87374495b5a461797494e9286713c

Observation 95df9dbc-b51e-4716-84d4-a1ff874671d7 · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs AlphaMath Almost Zero: Process Supervision without Process

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.092722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:30f5382502d8b6b181331cb485ac9a354d20f26a9731420a89d98b49b970736b

Observation 036242d0-3f62-4f0a-9799-50689a486566 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.098112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:f8174961b78c43fa147da7a38e17a72a05fc9e525ba9572c80a312f81761d998

Observation 0d4146c7-3b79-4bed-a7f4-c373ae647493 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:19:36.299641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:f5910963e291400e3412231e4abfc68143c2b664273832288ca9e558de18d128

Observation 4fdd0d13-c785-45f0-b0d0-cd0d1c5cbc16 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.109676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:a1f23b53fa7055cba1b124e983ecc534b7702cd443df5b72daad1226d1fa62d8

Observation e81133a8-39b6-4626-b85e-6c14d1e98abd · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs ORPO: Monolithic Preference Optimization without Reference Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.115187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:c6b61d6d5af7b838646523ee9c7d54b222a87ece9dc411690c54f770cd406b9c

Observation aa8c5b32-a312-4b29-ac0b-f19dd3af5479 · outbound

This paper cites Common 7B Language Models Already Possess Strong Math Capabilities.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Common 7B Language Models Already Possess Strong Math Capabilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.120167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:cbbc9265085272a115bffdc35d6f643a1c2fc7e00d9fc4b526d72d6032158895

Observation 489efe0a-d477-4c8b-94b2-47078a1d17a9 · outbound

This paper cites MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.125069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:2484960ef2cf8bf1e2139f9d67753a019680b8c17b388cf36d2274774245069a

Observation c0eaa2a6-8da8-4363-8916-3cc0dc939799 · outbound

This paper cites Let's Verify Step by Step.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Let's Verify Step by Step

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.129806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:a822a098c14bb04a0350fee238e7478708f471a4c97bfae37809b41c422587c6

Observation 4b075053-92a5-4a37-8f89-3f69aefd8a91 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Rho-1: Not All Tokens Are What You Need

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.135044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:c8bf3dbd835bca336cb9423a5f8fe56fa837728609d51f08be18198dea2dfa86

Observation bb851368-450b-43be-a9a0-49a5e3f9781d · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.139816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:1f6728038788db6a796bf535c4ba77db9620a771e13744f31774ce95bcb9d1ca

Observation fb61548a-ef33-449e-a772-18361de229fc · outbound

This paper cites Augmenting Math Word Problems via Iterative Question Composing.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Augmenting Math Word Problems via Iterative Question Composing

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.144173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:34ad173f4797d996decf4baa16b6fcafd214b776388f94264a8407b94441fc5c

Observation 9707572d-5283-4e31-93c5-8370a9918894 · outbound

This paper cites Improving Large Language Model Fine-tuning for Solving Math Problems.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Improving Large Language Model Fine-tuning for Solving Math Problems

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.148772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:3d26559329cb7e62ba12d7c54a0b4be6ac0895b0658099e763b54c7fbdd88e15

Observation 7ca1941d-7477-4e79-a4dd-3ff3a5abb369 · outbound

This paper cites MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.153407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:3fd9128f37b63cf236c7ace3717db625ddf0325855725bd3995c36610d0d3043

Observation b50c459f-c26c-4ef3-86ac-73c93556487d · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.158159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:2cc8856318cc4ea2cd32b14dee853f05699d226a74a29b651399dd6e908b9140

Observation bd279192-e885-45f5-9df8-41ea13196fb9 · outbound

This paper cites Orca-Math: Unlocking the potential of SLMs in Grade School Math.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Orca-Math: Unlocking the potential of SLMs in Grade School Math

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.162986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:b674c089e2e6dbd79b305d20d4a697cd2a28000ebf96ee6ca5bfd73557f1f14d

Observation 0d86d06d-2721-4ab5-af11-1b427c1e8f62 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.167115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:cf205391425263e63dae11bfd5e228d3b3399d25e6191318362372776cb0a8de

Observation 2b98df74-23c5-4bb5-8021-7ac74090b29d · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Code Llama: Open Foundation Models for Code

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.171477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:12c31563f5e90b3a1953c141cc0fffd4c71081351af9e25bf9a73bf04dfa8e69

Observation e8ed9e34-afb2-4b82-8dd0-7d5de9925bf6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.175855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:b6d31bd0d5ea6c9119104709266a6cfd021d69c39ae6b62db5a4b0d8d2f2b298

Observation 6de8ae2e-03b9-44ec-a1c8-60d64d7d8cbd · outbound

This paper cites MathScale: Scaling Instruction Tuning for Mathematical Reasoning.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MathScale: Scaling Instruction Tuning for Mathematical Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.180880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:77ef2a542cd540fed169b1bb435aa87c3ba4ceb4061bbbe402d13890bae94fcf

Observation 889dc090-5a34-4059-9d08-6d37ed4e0238 · outbound

This paper cites Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.185916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:79435f48e41f220c54ae85a23d4d48772c0337a4682168d15c9dc52c61cc9af6

Observation d17fe6da-b6cc-4608-972f-a4cf8b53a4e7 · outbound

This paper cites OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.190696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:e7fa38150b5ac08fd7fee2acb7b1ab4a663efe849a7fc3b866e68729cbf3a820

Observation 1a383ce5-5ad3-408e-8c76-85838ec2650c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.195217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:4219250a28b5be0e22ece5a8314d6ca90c6c3fb4d75fb7e716a9cdfb19494631

Observation ec1fd29b-da13-44ce-961b-f320537be3cc · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Zephyr: Direct Distillation of LM Alignment

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.200511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:e07dbeb71c59251a91d523fe21f1f67e4391189f07a1721a3937c4df12d99713

Observation d0abb843-beab-40b5-a7d1-4f466f8d5266 · outbound

This paper cites MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.205867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:d2b1fac133074f8d14a1bda7815d2149b4eb52a88bac5b5179c8df2d65f9a0dd

Observation b9903deb-c0fe-4bc6-b8b1-85f6d9478e51 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.210896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:ba6845503d6369bb4344ff0aa6417f7d0cf2f583a72aa89a661257ba65101ee3

Observation b0635685-0424-427f-954b-7dae3584314e · outbound

This paper cites ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.215547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:04d9dd6bdbe19a9605f320634c2b01a44e53331de1f844b23c6ffd0404592c7e

Observation 26791200-3458-4d05-8142-c370e8dc9306 · outbound

This paper cites InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.220895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:1de257d62e4db97c4466ead689153ad8e83d7c9fd91d13b350b84d8c0d322427

Observation 9e3522f4-4a64-419a-8bb9-b4f87a7d88a2 · outbound

This paper cites Answering Questions by Meta-Reasoning over Multiple Chains of Thought.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Answering Questions by Meta-Reasoning over Multiple Chains of Thought

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.225691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:754fa3f84a5353ce03aff3624fe213e6f792588044376a213a55f369220a1d04

Observation 859f9129-73d9-43d8-84d4-4772f4109909 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.230072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:7d830df9f938f7741c33ba240ae47e52c20b66dad187ef20b2093e21557e2b80

Observation 48280907-40fd-4801-94bb-c33f2fdc1998 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.234743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:640fa39deddc6a0d145804b1e25118ff0125fdc7b0823218f0273145633a24f0

Observation 94c1f6e1-2a00-4a6c-9985-8f2832498306 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.240596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:c1e215ea7d77c7a3a7d32bfed52650196130b2ce60432dc4fa8ecfc928b4956f

Observation 923df388-c29d-4125-bfce-3ad63a44236e · outbound

This paper cites MAmmoTH2: Scaling Instructions from the Web.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MAmmoTH2: Scaling Instructions from the Web

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.245776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:4af193f907b218080439ac52e2e075656dfbd8fe3a45f797c0d40268bd8a879a

Observation 17f98e61-12f4-4375-90f7-b28e009959fa · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.251197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:9ec21731ab182a74086123e3622f35a71a4a27871ceb23e8f3d5e08ccfabfec9

Observation 80e66f49-39ce-4d1d-82d9-3d72e932817a · outbound

This paper cites JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.256685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:7a87022bd329dc2aeb842b69fefbdbbb71931d30745e012ef6047fa21b21b52e

Observation b76230c7-eec0-420b-a315-402e3895a086 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.261388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:13b0e0e34b76fadb294313f7d69aa2e2a551fc8994fb94743aaae53a2a861724

Pith citing papers

Observation fa4bf688-ed37-43a3-8817-050c10a3402a · inbound

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents cites this paper.

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T09:42:04.287817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T09:41:59.979595Z digest=sha256:60a9077e2343174542f83fda77e8ac2f93f9290243d30fc0f8805bb88e77f222

Observation 5b0a0db2-e7bf-4f99-b7c4-43cfa6aff57d · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:3310e8791caed53a6e0acf849ac4fd3f6f4b28f20a82a3673d09472ea7cdf91e

Observation bf05dcc1-26d7-43c1-8876-d01f3cb28f33 · inbound

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment cites this paper.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.925085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.925085Z digest=sha256:f4d597857c05fc44e9147338a97126b5541e3f6a61434aac4b9c36c2d2d018a5

Observation c6ee085e-1b7d-49ac-a12d-16e73d18f24b · inbound

Preference Optimization for Reasoning with Pseudo Feedback cites this paper.

Preference Optimization for Reasoning with Pseudo Feedback Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.606869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.606869Z digest=sha256:f50ba6f6d097acc2040a5abb4ea7988aefcebe7fbf663b36d4238e86ae0548d3

Observation 90bd06fd-699c-4981-a096-0ddeb21a5be2 · inbound

Mars-PO: Multi-Agent Reasoning System Preference Optimization cites this paper.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.883872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.883872Z digest=sha256:99be4aa5acc4ccb7b7dab9b0354ba624a7df5958d12ea6719c7a740e010b7047

Observation 232c31f4-a495-4526-85c3-39f190e2e66e · inbound

Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability cites this paper.

Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T05:43:31.521431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:43:31.521431Z digest=sha256:4de3f3874796260e006197c02976845eeefce7308c2c5688de58bfa1d8f43a53

Observation 37ad7b07-5d02-43ae-a588-8ce714176bf7 · inbound

A Systematic Examination of Preference Learning through the Lens of Instruction-Following cites this paper.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.695300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.695300Z digest=sha256:3f5fa7d3da321243a20c1451e8d95fd4536465ed9902b9c7f25c62707f0aec51

Observation d6784637-1a2b-4c90-840f-60079620f06e · inbound

A Systematic Examination of Preference Learning through the Lens of Instruction-Following cites this paper.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.698420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.698420Z digest=sha256:829c03634e731bf03a98b2a50e3ec312e45a02cdad2dd4c9ddbc13544ff281af

Observation 076dad93-3150-45f2-97f7-fa92d32b0265 · inbound

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization cites this paper.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.417074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.417074Z digest=sha256:6268ae18876c950d0e3cf998c0f2ad903a05d085f2ac8499bf8f53207ae46e67

Observation ca2d97d6-65c3-449a-a8e3-ee73b60c5396 · inbound

Plug-and-Play Training Framework for Preference Optimization cites this paper.

Plug-and-Play Training Framework for Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.108960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.108960Z digest=sha256:98d0f29c063b27e770066f221cc19ee6bb9fcfbfb17c2da79a3c4d8103b4dfe7

Observation 93ce969c-ecec-4fb8-aeb2-2c5092d15dd2 · inbound

Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search cites this paper.

Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:07.176830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:37:07.176830Z digest=sha256:cdc765963727d2b3c7d9ec21c8b677c804c73157778a12e6e6434e570689d1ae

Observation c88df6c2-b598-4e56-bd5f-faeed1545a19 · inbound

SDPO: Segment-Level Direct Preference Optimization for Social Agents cites this paper.

SDPO: Segment-Level Direct Preference Optimization for Social Agents Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:28:56.735753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:28:56.735753Z digest=sha256:506c4d4b1fb525f36bde891ed1f2dfc372a0bfb3fff5db53708606d2789a3a9c

Observation 1ae5de45-65db-4af0-9cce-29c3e93d268a · inbound

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning cites this paper.

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:13.255347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:13.255347Z digest=sha256:22380d9d76acf094d24aa0164631a23feb29c235010f22a884fa1760887f81e2

Observation be8edc1b-2d84-466a-aeca-f54a9f5890cb · inbound

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step cites this paper.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.315309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.315309Z digest=sha256:4d906d2021e73864f6b3504aa4ab6fd996a99d94225d976216e46742e227b630

Observation e7fa5bfa-b430-4b78-a825-3df5b72de4c3 · inbound

Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages cites this paper.

Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:49:27.751399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:49:27.751399Z digest=sha256:c475d8371bf93f137e91dd21d60a25251c7554c69c7d9d866f99338df31ded35

Observation 30af0243-2239-4962-b88d-f4ad59925137 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.721346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.721346Z digest=sha256:c776a37c1aa3531bead63648a279510cb479e21aca9b78f84b62e33d0d6ef12e

Observation 91d6dd8e-0e46-4c22-ad3b-e4a5371bbc10 · inbound

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information cites this paper.

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T13:25:52.021917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:25:52.021917Z digest=sha256:c69ad905c901195467432201b010a75a58a569c65bed6de0435353210f01001e

Observation 1bce7f0a-c59f-4953-aeb5-9971afb79efa · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:6ce4b24da7d2f401c0415b361c31f41e83b27d3d82c5715d97dcfc284c463ccb

Observation 4f4307c2-cea9-4983-a12e-5a00a1d79f35 · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.393381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.393381Z digest=sha256:c9f996c4ef5a8160bacf8733b9b12a39dff5091399fba0061ffd6033ebd7b0c7

Observation 290e26b9-8ff6-43c0-ada4-d2130980bf9c · inbound

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models cites this paper.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.321272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.321272Z digest=sha256:dc6c9a3c7b7313c69962a96d71dd812108a1c941bf8c27082fd170e110c2a492

Observation 48c1bd92-9ef3-44f5-81fb-ac6f08efd1f9 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 181

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:da18bcf935403bd49bedc31a8ab41114e6c4cf62ba82131106a6d5f3a7840323

Observation 794dfe46-4d9f-44d7-bc45-92f56829af41 · inbound

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs cites this paper.

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T18:42:39.023650Z digest=sha256:aabccc28dd136045319c2f140f531f05d9c0f670f74f76ee517de3177dc75047

Observation 17b3d4ed-48a9-4b14-bb00-3b7b7de3309e · inbound

A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents cites this paper.

A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:47:11.169539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:47:11.169539Z digest=sha256:35adf20ced864b57a6c765ab2d131f6757591e9de6c7dd8ecbd0a1434e4f9901

Observation 92f99346-541d-4eac-a7e4-949de8ccb1cd · inbound

Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization cites this paper.

Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:54.456080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:54.456080Z digest=sha256:7d1b8f02609f2522df53b7c1f6900b69b4af6ac0b11f11533df807e04fcabf44

Observation 3ce94a61-710f-44ce-b31b-e7a69c357353 · inbound

Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think cites this paper.

Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:26:00.389667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:26:00.389667Z digest=sha256:66a3fd7cf12dc304de050ea2e28d23004730ce521f70870d332abea7c3c073e6

Observation b43040ec-c1c5-4539-a775-d04965fa8068 · inbound

Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning cites this paper.

Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:04:19.938135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:04:19.938135Z digest=sha256:4a2c3ae3f4d5220081340b63ed978572bc068a02327e4db3f31471c4ef6f35f5

Observation 7c81debc-3edb-4cd7-b1ab-dfbd6175ec03 · inbound

DeepCritic: Deliberate Critique with Large Language Models cites this paper.

DeepCritic: Deliberate Critique with Large Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:07.407517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:42:07.407517Z digest=sha256:4b31718cf9d7524e5ea4c409140062102ac3c3b6c3c757a43e8e9307de630f51

Observation 8660729c-8e83-49b1-95ca-e77ece1c9d78 · inbound

A Survey on Progress in LLM Alignment from the Perspective of Reward Design cites this paper.

A Survey on Progress in LLM Alignment from the Perspective of Reward Design Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:06.963137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:52:06.963137Z digest=sha256:2086789a97a7cb5ec688f852674286be171bc086e3e1b3c762d12cd58e7c0c86

Observation 868ba187-310e-4489-819b-af9a1aa9cb1c · inbound

Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving cites this paper.

Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:31:49.452715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:31:49.452715Z digest=sha256:1cca5b9f1c53efd35096c59e86ea6e39fa9ce20d684838fe80f5658fd8bcca86

Observation 558b4f26-ef03-4180-8019-5dba93cc9ef2 · inbound

Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience cites this paper.

Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T23:31:12.275468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:31:12.275468Z digest=sha256:68ed09206569176d74a3e5f84968d7cf9990bcadcc25d5ccd5f81bac7fc875d1

Observation 760fa94c-8fe6-4107-ba6b-d26cc0667127 · inbound

CRPE: Expanding The Reasoning Capability of Large Language Model for Code Generation cites this paper.

CRPE: Expanding The Reasoning Capability of Large Language Model for Code Generation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:22:19.134458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:22:19.134458Z digest=sha256:2afe81d5a054f054773c14bbd339d3868b33e9de93e03d98b7547f77b3456a9c

Observation 0257b275-83f4-4750-b72e-b9cb18fffc55 · inbound

MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision cites this paper.

MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:23.284087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:23.284087Z digest=sha256:dd16ecab854f83206579aa92b51aa839198008ba78902e546cc6a25041455e2a

Observation c5476062-0026-4499-be03-a071d6d1d788 · inbound

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning cites this paper.

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:59.458706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:22:59.458706Z digest=sha256:5fcefe1c786e7f397696f34be3187291f429ca997702056d296d436f899e8c03

Observation 6883d3c8-a050-40a8-8301-f9d80190edc0 · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:58.907511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:58.907511Z digest=sha256:24d1529683f6c2775dc01e9b7859a8bd4b24d91c2a8683c574ca14ef388fc050

Observation 93167904-f699-4575-8d5c-3297e29e1fcb · inbound

Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries cites this paper.

Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:54.505939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:59:54.505939Z digest=sha256:9911b16a9c3d9314cd9af6f9f461e27084df9fc51695a55aabc9bb6bd6baba6a

Observation 1122bdd9-2470-4da0-a66b-0cc656f55190 · inbound

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning cites this paper.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.944799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.944799Z digest=sha256:b7998fd743ebe362b590b0b6279d6879eeebe990cd4ea9c3c308c8b86bc9447a

Observation c01072d5-3cee-4766-bfb9-26bbf4b826a7 · inbound

Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models cites this paper.

Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:19.624700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:19.624700Z digest=sha256:61b418da0c01502101202e629d0b84a8eb6825c3e4acbea8a52e6005c93bcbe6

Observation da8f3bc5-b4f0-4027-9da9-3c0cee68f900 · inbound

APT: Improving Specialist LLM Performance with Weakness Case Acquisition and Iterative Preference Training cites this paper.

APT: Improving Specialist LLM Performance with Weakness Case Acquisition and Iterative Preference Training Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:11.100055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:07:11.100055Z digest=sha256:ce980bbc09de61d219d0495e4371e41f1fd35ca0b5c14021074d81c9ea1500e2

Observation 96a52f76-7061-4763-814e-a502ca788c4c · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.300557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.300557Z digest=sha256:08ada25cc1d8818a7e424f38b9c14fe57fa3054baad23dc51dd786842614be6a

Observation d19fd706-3403-47a2-ad12-cb865ce56a1b · inbound

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization cites this paper.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:43.248348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:43.248348Z digest=sha256:5bfec22070f20e97fb39037322df38248e93400fb9673f6072cf2610fa536a29

Observation b40b769e-4c06-48e5-9d9d-e10fc34fc93c · inbound

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning cites this paper.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.795071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.795071Z digest=sha256:b10c83d9156b396861fd0a5c277b18af9101147ee4c557a2747e78107107de36

Observation d9653dec-3380-4c78-8316-8b21d5210f57 · inbound

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation cites this paper.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.558586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.558586Z digest=sha256:e224a7a7cd96b9e7efb65260b6049434a3bc2c0eff79b53b236d7010687becdb

Observation 5a898ff4-7998-4ad4-9dd9-40d056379cdf · inbound

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations cites this paper.

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.162976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-19T06:09:26.269452Z digest=sha256:e8d7a3fa6b0aae0e19bae8f8b41ce6ee1ff007dd4be7138807e3127153a1f187

Observation edf792c8-05a7-415b-8318-369e1b8a0856 · inbound

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization cites this paper.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.880112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.880112Z digest=sha256:78b720ca9c55951989dbd9f458b1c671bb7fc81c1f0c8ab5f77d4e63b6d2bdd3

Observation 00092c23-b383-49c4-b201-bfa48500b59f · inbound

Mitigating Object Hallucinations via Sentence-Level Early Intervention cites this paper.

Mitigating Object Hallucinations via Sentence-Level Early Intervention Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:35:32.638911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T08:31:24.173135Z digest=sha256:d77438d5981396f4954ab172a53fad38fbb34e73d68a7e57522ad9e76559483a

Observation 17cbbd3f-3e34-4169-a709-8d750ec9ca1e · inbound

Unlearning of Knowledge Graph Embedding via Preference Optimization cites this paper.

Unlearning of Knowledge Graph Embedding via Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:03.267995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:03.267995Z digest=sha256:633fa30e77308e87513ba5dbd906b8ac2df4003f85e406fed360d8a4934cef9d

Observation c3f08595-d6ec-47d1-b174-aa944314cf68 · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.477083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.477083Z digest=sha256:14218db52cc18cce8432a92b42db3274924ac79a177b1f2d35bbb840ad26d309

Observation 1e3cd6c3-9c29-4b23-8a4a-9af3c335d1af · inbound

Sample-efficient LLM Optimization with Reset Replay cites this paper.

Sample-efficient LLM Optimization with Reset Replay Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:38:08.044450Z digest=sha256:d512235c266ccab229b57ca8d7b32f84b1d8c6db99afacfab9bab5656223e413

Observation 564cad11-567b-4dad-a399-c7f3cfefca93 · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:405543a1990f694b6f7e8d30b1f44f8ed7054656c92ed03ea82afc6824490f59

Observation 62f755ec-653d-4aec-ae1b-d63ac4f84676 · inbound

Alignment with Fill-In-the-Middle for Enhancing Code Generation cites this paper.

Alignment with Fill-In-the-Middle for Enhancing Code Generation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 626

Resolution
unresolved
no resolver link, observed 2026-08-15T16:54:20.341720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:54:20.341720Z digest=sha256:5ca40efb5abe29b85cbb1938339162df40f6f7c14f9f5efaac5385b9a4352fda

Observation 808fbd4d-c3e0-4f1f-88bc-b8bff91c6cc2 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:10dd42eb90fdcb0c897546e96cce0b6b31a39cafcaf56e8343e2edce3ef0ee4d

Observation 10d2c1ef-7518-47a2-9077-539b32c76d44 · inbound

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making cites this paper.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.865781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.865781Z digest=sha256:bcaa5937dacb5e2d01ed74b5fe9940fbbabe89ead59aba376899a7b0147e83b1

Observation 65f2fb61-d0af-416e-b177-2e1bac71d1e6 · inbound

GPO: Learning from Critical Steps to Improve LLM Reasoning cites this paper.

GPO: Learning from Critical Steps to Improve LLM Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T15:55:56.278478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:55:56.278478Z digest=sha256:d5ebf38116610e45ee8a764638e1d69bef7f25a1bbfd3d6af5a38795ff89321d

Observation 49446334-2c1d-4385-a9e1-a519bf67dd97 · inbound

Future Policy Approximation for Offline Reinforcement Learning in LLM Reasoning cites this paper.

Future Policy Approximation for Offline Reinforcement Learning in LLM Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T14:41:24.891052Z digest=sha256:60c0b48265c934bfab28339e7013a31ea6e3a6f36a20338fa703776dcb400016

Observation eec23df6-7e9f-4b05-95c8-180768cf78a1 · inbound

SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T09:33:57.037949Z digest=sha256:dfec0c086a05a8e624e774f2a64ba18d448670632d2efac0af318164c75ff4a6

Observation c015ebd5-0bb5-4d99-86fc-21eb80bd959a · inbound

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.480368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.480368Z digest=sha256:14610bd8c967a2fa04c60239cfd2a1ad099f606b8448e16c5c264fb8164279e7

Observation 1c37e923-e5d4-4838-b5d0-fba9af9e4bdc · inbound

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling cites this paper.

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T03:15:57.744706Z digest=sha256:f4c6a7c4d9416800eb10d204c1c4545802df4f4c153fc7696daf18ed1838f625

Observation fb12e888-3d24-44dd-b835-dc5ce12a995f · inbound

Hard Negative Sample-Augmented DPO Post-Training for Small Language Models cites this paper.

Hard Negative Sample-Augmented DPO Post-Training for Small Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T21:57:16.279587Z digest=sha256:7229f71f42152b8ce8f174896884ca88319895b8d6189093cc9751df1f5770b9

Observation 3aee8528-579a-4fed-8205-11dc1dea540c · inbound

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation cites this paper.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.025781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.025781Z digest=sha256:3dfab1cddd23cf1a56a516dfd93ed92ca58b42e7a72adffbbc78eb2c45c4eba7

Observation a7b9b8bd-b4c5-4748-82ce-a4918d724865 · inbound

Curr-RLCER:Curriculum Reinforcement Learning For Coherence Explainable Recommendation cites this paper.

Curr-RLCER:Curriculum Reinforcement Learning For Coherence Explainable Recommendation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:43:11.384787Z digest=sha256:e3f79eca85485055b43eb9b32f6a03fb6a72212ef34175c1930381ebf900075c

Observation cddd1db6-ec81-4b3c-a7e5-c64dcf347b56 · inbound

Decomposing the Delta: What Do Models Actually Learn from Preference Pairs? cites this paper.

Decomposing the Delta: What Do Models Actually Learn from Preference Pairs? Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:22:00.563479Z digest=sha256:58d3258eafc05aaa648de90243c2de08f120ee3031b113a7947f36f3b7666d28

Observation f443fea7-b8cd-4791-b487-5ca6d634b5ac · inbound

On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization cites this paper.

On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T19:35:59.832868Z digest=sha256:55d5022e05ef34eb65f87613b2eef2f2bd220ffe5cc94c933c21466805d96e50

Observation 6c56e097-7310-4955-acda-b5063cb3d4bd · inbound

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding cites this paper.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-09T16:29:05.186607Z digest=sha256:1630d2f8ff1e5dc68ffc713f20347f92c26d7d054fad27fdd98ce0d7c4b61052

Observation 4cdde1c5-a6fa-4195-bf6e-12b1b00d862e · inbound

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction cites this paper.

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:22:15.608348Z digest=sha256:220740122ed03d1910cd0a893dfc334e7d11984e158993c1213a2e22343a5e85

Observation e145583c-698b-40c5-8677-aaac3491cc14 · inbound

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models cites this paper.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:705f05136b9d5767a3b025ec4fcf9f0f17a7b8291f609d0427ed5f73086663bb

Observation 8970bf92-f524-4ea3-b4d6-39a3f68b7ccf · inbound

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge cites this paper.

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:19:00.238602Z digest=sha256:a55ef96f58f443560b2f650db7640bb5c08fc8b68ab193dbf24c0eec9c2a2a96

Observation f459a32c-788c-4e36-b98d-4479abf8e8f0 · inbound

YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning cites this paper.

YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T05:29:42.381280Z digest=sha256:f4df38c0168cdf4dee152f789b94cb31ad99a161722fbf98d535f4535b004b71

Observation d0349acd-d333-4845-8385-e9e4ec6bbcf1 · inbound

Enhancing LLM Metacognition via Cognitive Pairwise Training cites this paper.

Enhancing LLM Metacognition via Cognitive Pairwise Training Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:02:33.922491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T19:01:18.153145Z digest=sha256:7961434452208922c35104cbd9e509f3eff25f1636787214f75b415f27bc658c

Observation b7776c39-e0cf-4b81-9ecd-8433735cd0b0 · inbound

DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling cites this paper.

DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T16:57:10.030422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T22:16:01.457276Z digest=sha256:ea471d757291d0be2a04cc44b9b8b6f7384a5a1b8c83662fef4ac891df1bdb43

Observation d6322ad5-456d-46c5-b859-1e7907ebcf31 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 178

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:25.924065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T18:39:44.696961Z digest=sha256:d67be08080d721b2facf23d1672c925c5af24e93e517a1bb551eed046add6c4a

Observation 2ff953ea-7321-44d2-9486-31e62b458d5e · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-02T12:05:18.089147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:05:18.089147Z digest=sha256:d2575fec3d0bf8c4cfffaf219d24350ac76084ad68c9248dfd492f88a35efab5

Observation 765c1947-6e7a-40f9-a237-21b4d3aeec03 · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:27:36.668377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:0ec246ce5c4e22fa96a10118ef214fd235395dd31b8c27bab0a3e8f1196c6cb5

Observation 0138b81b-b3bd-49ee-90ad-45c393ff9032 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:37:49.416543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:0e6996758527c789d5df2f7c5dbe88b326bc1bb83f7ccf7484b723700a0a8eb8

Observation d88db0e8-6e3c-4970-8b32-4f9170e4df84 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:27.879246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:27.879246Z digest=sha256:256f3af4502b207ac2d29b325a111db0e29578883fdf63bfc9a35cb60913c30f

Observation 165fabcd-4440-4baf-a1cd-42f140a0680d · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:08:33.402998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T06:39:34.199607Z digest=sha256:79f56327545f7248d5c901318b101f447534c0ed186c9cf7ae26847205452a02

Observation 1a30c586-887c-4550-8011-78f8f7c1b6de · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:19.127147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:19.127147Z digest=sha256:6319631e8218f7a5f259c2ff3fb8a60b678cbdd7d6c864b0e191f873a93ad5d0

Observation fa580899-0bc5-4646-a750-1e0487fdcf0a · inbound

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence cites this paper.

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:38:44.174836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T03:57:19.028507Z digest=sha256:48727e27cf918659e85e09c44b85f4fd5f378703797abe863b9e6407db5a318d

Observation 6d9ae9eb-77f8-48ae-80d4-99f0a7491913 · inbound

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents cites this paper.

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:40:07.843363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-25T19:55:51.114244Z digest=sha256:5a27c8ff80ff9753a3c945f300a9a03791903719759543d37ce1bfb7d4b953a2

Observation d4408eb1-6fd1-4e65-a041-39d1f3336854 · inbound

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM cites this paper.

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:33:28.303470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T13:27:45.796650Z digest=sha256:2ac7385fef8fd3c5059b15e8d93a3a41024642d775d697d01c30006a39bcb641

Observation dbbdc301-8aa1-432d-8792-d5fdf14ce555 · inbound

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models cites this paper.

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:09:50.060500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T05:32:59.640335Z digest=sha256:9a81db6763bef498d58ee31cb8f0c8b7a03d24b6b1488747ad298b33c2b36d37

Observation 1187e579-c9d9-4db7-bdca-02c9bfc88a57 · inbound

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models cites this paper.

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T17:23:45.531776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T05:19:07.877346Z digest=sha256:29a4c3dd327b001201163dcf6da6a790f7db2e17a72b3eebf733fb0c4611ed91

Observation bcd9bee0-47a9-4340-8155-476652639379 · inbound

Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement cites this paper.

Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T08:04:28.377800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T07:55:28.254309Z digest=sha256:b0c0ce4473ba746ee26540bfd4b2077ed179eda605e5765d4ac4a8ce23fdbae0

Observation cfa2c755-5b1e-4f8a-9995-6e32ecd56d6b · inbound

StrucTab: A Structured Optimization Framework for Table Parsing cites this paper.

StrucTab: A Structured Optimization Framework for Table Parsing Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:24:18.922918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T06:22:21.694355Z digest=sha256:7c625f6818c07ba98e4a170f0d79021538de60aa1bd5643857c620ceaf310f56

Observation 08ae0c01-cea2-4452-afda-ac733a5a2c7d · inbound

Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning cites this paper.

Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T02:23:04.515705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:23:04.515705Z digest=sha256:b5d42fe153a5147061f1a50f3d85762854a5fc814d966001d2f6e5e7997fff34

Observation 52443909-5f8f-4eee-b2b7-9b01f26e98cb · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:ed80ed8bcdc87ca328c986b832a70d7b871f67162d48d91041fd794e6c51cc62

Observation 85614d0d-d12e-484e-976e-8ffddea0d4f4 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:42.479453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:42.479453Z digest=sha256:9acccdd827cddafc828e100ada7e8ca594326225e7e1c750374bcd0295a86992

Observation 87b2b2cf-224f-4ae4-9ad6-dde87fe12866 · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:8e2c88e1041fc735e4d560dccde9af602958ff0db2a670b28b6685276bbefaa1

Observation 6ad3e2a1-1796-4da4-ae1e-a62d7563a082 · inbound

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding cites this paper.

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T13:42:24.987811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:42:24.987811Z digest=sha256:a2846dec2de8912bac77187b363330e8028daae4380d1b05ab7cd3550c173ead

Observation 14ed6ea2-811c-45d0-84b0-51b5553374de · inbound

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling cites this paper.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.591812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.591812Z digest=sha256:1da7fad89cb73d8ec699bc654c4d4c1ba9edb499b1f3e9fd9f5deacd7a299667

Observation 0e5494c2-1c61-41df-8ec9-96186cc38ff7 · inbound

PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation cites this paper.

PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T14:56:46.932676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:56:46.932676Z digest=sha256:938494a856440ca8c179681c78c23b98c2782463ede9cc77e26a30e2c6e74b0b

Observation e050620a-34d5-46d5-b4d6-852f9bc765c1 · inbound

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization cites this paper.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.114561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.114561Z digest=sha256:bd7e6e37e7c26ba7c84e51d58ef35e02617fcce7ce95360b272b5df1b4ee7e75

Observation 2ba0f0b6-a43d-400e-84d0-9177250c6a24 · inbound

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs cites this paper.

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T00:39:40.267792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:39:40.267792Z digest=sha256:2f3cb33138ccaca47506d1f63eea494c8273a938947f2da36f65a4fd756b958c