Pith. sign in

Paper Citation Record · LEDGER

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

As of 13 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 75 inbound Pith citation observations for arXiv:2406.18629.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.18629 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T23:58:29.040819Z

measured 113 of 113 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 75 of 75 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:20:13.925085Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact35
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6cc6da4a-7f20-472c-93af-ce641a20af4b · outbound

This paper cites GPT-4 Technical Report.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.068590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:2d73088c4c8a4a4b3060d35e85e2e27582a5b35d157ea65907a01be407372d29

Observation 5f81d01a-12e0-41f1-9710-d4a4fd8518b2 · outbound

This paper cites Llemma: An Open Language Model For Mathematics.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Llemma: An Open Language Model For Mathematics

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:17:47.048289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:38b25c04beea26a1408d02fe0d6943aaf192c55ffc6ca4d582d62a409d09656e

Observation 8911546b-1d35-42c6-abb5-1068da101f26 · outbound

This paper cites Qwen Technical Report.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Qwen Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.085943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:aa36ce5fde2439b99b048c0d1c4b4cc0322c24c1a64b0caeb7514e302cfc3ae7

Observation 95df9dbc-b51e-4716-84d4-a1ff874671d7 · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs AlphaMath Almost Zero: Process Supervision without Process

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.092722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:b08361d010b5c629aca3ff8ef771727dacca22012d962ed3a510f5331873d4a7

Observation 036242d0-3f62-4f0a-9799-50689a486566 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.098112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:b44ac16a86237444b5ad8585d039f2906215fff8b6bdd4309330fd5c71909545

Observation 0d4146c7-3b79-4bed-a7f4-c373ae647493 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:19:36.299641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:8928af2db64d758c660d1a0131bf80c1d0157035371bd00e6fd22537d8ad243b

Observation 4fdd0d13-c785-45f0-b0d0-cd0d1c5cbc16 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.109676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:73b7a2ec54905d454c42bf9df368fc624b68fdf5d039859da39ad3b193c2d225

Observation e81133a8-39b6-4626-b85e-6c14d1e98abd · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs ORPO: Monolithic Preference Optimization without Reference Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.115187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:660b383cfa9832f4ba98acb3c6a1ec7941f77df2760d5a2c5c2a821c097bfdd4

Observation aa8c5b32-a312-4b29-ac0b-f19dd3af5479 · outbound

This paper cites Common 7B Language Models Already Possess Strong Math Capabilities.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Common 7B Language Models Already Possess Strong Math Capabilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.120167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:123d947cbdd3f9166b2beb512ef7a4bfb31fab54674571ff0e8fc0eefbeffb49

Observation 489efe0a-d477-4c8b-94b2-47078a1d17a9 · outbound

This paper cites MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.125069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:e63d97b19346c421f68ba11fe9534a4fd69f12ad3d50f3318606001244c50428

Observation c0eaa2a6-8da8-4363-8916-3cc0dc939799 · outbound

This paper cites Let's Verify Step by Step.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Let's Verify Step by Step

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.129806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:b8c8a00ff4b3a50b0edee6f35679cd9a58ce220d738cc86a5b691bfc1a713a30

Observation 4b075053-92a5-4a37-8f89-3f69aefd8a91 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Rho-1: Not All Tokens Are What You Need

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.135044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:51897bd22da77515ad3fabd01d627cd94b07ec8050e31eaa285cfcecd885d1a7

Observation bb851368-450b-43be-a9a0-49a5e3f9781d · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.139816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:2ece3b7cb6bbb6820176d37e7623080338b9ed69472e0dd42c24cc56a142c6f8

Observation fb61548a-ef33-449e-a772-18361de229fc · outbound

This paper cites Augmenting Math Word Problems via Iterative Question Composing.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Augmenting Math Word Problems via Iterative Question Composing

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.144173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:5574980466908cef1a0431f2d17064839161511ca214099727cf17676bfcc4f4

Observation 9707572d-5283-4e31-93c5-8370a9918894 · outbound

This paper cites Improving Large Language Model Fine-tuning for Solving Math Problems.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Improving Large Language Model Fine-tuning for Solving Math Problems

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.148772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:afdc5aae1f347846ff198bb7b89cdcb7b2ddd534a4d8dea26240a40e050a9202

Observation 7ca1941d-7477-4e79-a4dd-3ff3a5abb369 · outbound

This paper cites MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.153407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:0868dacbfffe118eb5af491a11188e26682237ba4f67432045df2e30d94f1e67

Observation b50c459f-c26c-4ef3-86ac-73c93556487d · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.158159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:2335a503bf07456973e735762b809af5baf82ab8056c5391c0187a03c70116da

Observation bd279192-e885-45f5-9df8-41ea13196fb9 · outbound

This paper cites Orca-Math: Unlocking the potential of SLMs in Grade School Math.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Orca-Math: Unlocking the potential of SLMs in Grade School Math

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.162986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:62a02569947465dd85606fc1d416b99ba636e27d86c1d4580bf80337be744bd5

Observation 0d86d06d-2721-4ab5-af11-1b427c1e8f62 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.167115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:4d0ea308a42121fcc1690fb3571ffed721c39db78472929cc96dcb9a9b6e42d7

Observation 2b98df74-23c5-4bb5-8021-7ac74090b29d · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Code Llama: Open Foundation Models for Code

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.171477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:5ddde613bb32e565e609f81c12aa0cfba3676aad359f6ada3123c10b89abb73b

Observation e8ed9e34-afb2-4b82-8dd0-7d5de9925bf6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.175855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:80545cbc443e5ad35391e4f8fdb9114d95772a3538f77f2323a4426c650f1475

Observation 6de8ae2e-03b9-44ec-a1c8-60d64d7d8cbd · outbound

This paper cites MathScale: Scaling Instruction Tuning for Mathematical Reasoning.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MathScale: Scaling Instruction Tuning for Mathematical Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.180880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:c86e01e9ca1ff3abf52b555240193f8a64b3411828a57d2e25dd0a236d8c5cd3

Observation 889dc090-5a34-4059-9d08-6d37ed4e0238 · outbound

This paper cites Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.185916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:326244bd478822b05e75cab65668195b9c88ec7b2d984e08f1360e444c8951a2

Observation d17fe6da-b6cc-4608-972f-a4cf8b53a4e7 · outbound

This paper cites OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.190696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:f9399f98f4df92ef4ca2ca872782be451221589eb7e729475a07c2d8277f616c

Observation 1a383ce5-5ad3-408e-8c76-85838ec2650c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.195217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:b113ca49cce4d3864516c5a5af4f9afbb612f372baca281b3d26ba7b119e3825

Observation ec1fd29b-da13-44ce-961b-f320537be3cc · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Zephyr: Direct Distillation of LM Alignment

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.200511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:0b971d01f56941c2bed49a148667e9db97370976a2725d3caed32edb61ad1f10

Observation d0abb843-beab-40b5-a7d1-4f466f8d5266 · outbound

This paper cites MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.205867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:46353e65342efa6bd619832fd7f9f50849dae99926d0b33f2537cccc44b30367

Observation b9903deb-c0fe-4bc6-b8b1-85f6d9478e51 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.210896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:63a9fbfc3c163710aeb66d5eab0ac6a60d737456908a99fa7c46390d95dbe595

Observation b0635685-0424-427f-954b-7dae3584314e · outbound

This paper cites ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.215547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:ecdcf410d849714b03da3d669685795c1f693547d78a97519d580162a27edf2a

Observation 26791200-3458-4d05-8142-c370e8dc9306 · outbound

This paper cites InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.220895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:298a12fe632712ec26aad5e5d67a57f3730fcb0fb9593d39146fb72c960b5db3

Observation 9e3522f4-4a64-419a-8bb9-b4f87a7d88a2 · outbound

This paper cites Answering Questions by Meta-Reasoning over Multiple Chains of Thought.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Answering Questions by Meta-Reasoning over Multiple Chains of Thought

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.225691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:61fcb5a9606c06303d003428184c2cfc15b0e74560a9ac37f6c9c005ed8c6095

Observation 859f9129-73d9-43d8-84d4-4772f4109909 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.230072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:6bafdf6de5d942a00c8f4cd6f6412b621da12284f29f6be800c69126811a8a32

Observation 48280907-40fd-4801-94bb-c33f2fdc1998 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.234743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:f85d12105480acd1359808d0bb3515f8aec60b215aedb78ebb16a23dca39a572

Observation 94c1f6e1-2a00-4a6c-9985-8f2832498306 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.240596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:e2b160ac30d98c7a33af473ceec6ef7884d03b1634b1d9a1852465e805210f71

Observation 923df388-c29d-4125-bfce-3ad63a44236e · outbound

This paper cites MAmmoTH2: Scaling Instructions from the Web.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MAmmoTH2: Scaling Instructions from the Web

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.245776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:c65307ea97835f8d061ca7467a5348bfbb96c8dac5b3078d9d97462fc19680fe

Observation 17f98e61-12f4-4375-90f7-b28e009959fa · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.251197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:04f80ca754fc8ff32894d925d4af7e79164cb35fa7864e87c2d1c3c09690199f

Observation 80e66f49-39ce-4d1d-82d9-3d72e932817a · outbound

This paper cites JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.256685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:3843dc13313ce18e391af4bf7d76a63ec0c9f7f2c014e0a57af421dc3bbd335a

Observation b76230c7-eec0-420b-a315-402e3895a086 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:58:29.261388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:58:29.040819Z digest=sha256:eb25d08592f5cc4895a49c5a3cb43cc09f5b0c028a760747a663c94056e529df

Pith citing papers

Observation fa4bf688-ed37-43a3-8817-050c10a3402a · inbound

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents cites this paper.

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T09:42:04.287817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-20T09:41:59.979595Z digest=sha256:5902776ea50272e712113df56d40f4f243fe34332c5fda357fa55466eed7a33f

Observation 5b0a0db2-e7bf-4f99-b7c4-43cfa6aff57d · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:1758fdf4e20bf6294bb165f0db028337287aec09b463a7777889e8c6c86a4a26

Observation bf05dcc1-26d7-43c1-8876-d01f3cb28f33 · inbound

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment cites this paper.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.925085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.925085Z digest=sha256:98aa7a687c3b4ebb2a841c8d54720f0b54fe721913d25805899af343364647ea

Observation c6ee085e-1b7d-49ac-a12d-16e73d18f24b · inbound

Preference Optimization for Reasoning with Pseudo Feedback cites this paper.

Preference Optimization for Reasoning with Pseudo Feedback Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.606869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.606869Z digest=sha256:0404156d89f0a76a50946003ed01b676a95bb315646cace01b2fa8e3f89535f8

Observation 90bd06fd-699c-4981-a096-0ddeb21a5be2 · inbound

Mars-PO: Multi-Agent Reasoning System Preference Optimization cites this paper.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.883872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.883872Z digest=sha256:f043bafb283853dfb96854d1957a51620bda1226acda7f8b702598244274ebb8

Observation 232c31f4-a495-4526-85c3-39f190e2e66e · inbound

Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability cites this paper.

Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T05:43:31.521431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:43:31.521431Z digest=sha256:d12bb1ab23896a3ab0719ba1ff0d18f4e39710e0041284151c7dbdb93620b751

Observation 37ad7b07-5d02-43ae-a588-8ce714176bf7 · inbound

A Systematic Examination of Preference Learning through the Lens of Instruction-Following cites this paper.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.695300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.695300Z digest=sha256:9dfd1be086e5987d67203a1dd082c7624c51101dfc7096c28d6802cd05013c54

Observation d6784637-1a2b-4c90-840f-60079620f06e · inbound

A Systematic Examination of Preference Learning through the Lens of Instruction-Following cites this paper.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.698420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.698420Z digest=sha256:891a21da10365c42fdd75e211e19391b3a1281eb11109957e2b40a10cdde5860

Observation 076dad93-3150-45f2-97f7-fa92d32b0265 · inbound

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization cites this paper.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.417074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.417074Z digest=sha256:856c16db43da50e17639056c5a53c372e5ce9336640ea0178ccd3bb245671e57

Observation ca2d97d6-65c3-449a-a8e3-ee73b60c5396 · inbound

Plug-and-Play Training Framework for Preference Optimization cites this paper.

Plug-and-Play Training Framework for Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.108960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.108960Z digest=sha256:38066dbd99476c4b5ebb8cf2736a5aacffa7e687acfb0c0f1714839cd0ae7446

Observation 93ce969c-ecec-4fb8-aeb2-2c5092d15dd2 · inbound

Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search cites this paper.

Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:07.176830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:37:07.176830Z digest=sha256:2068b9956ec5a451731b76925872bcd67b83be60c1b862bf912eed7bef59699d

Observation c88df6c2-b598-4e56-bd5f-faeed1545a19 · inbound

SDPO: Segment-Level Direct Preference Optimization for Social Agents cites this paper.

SDPO: Segment-Level Direct Preference Optimization for Social Agents Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:28:56.735753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:28:56.735753Z digest=sha256:e955a266dca8d512d29da8cad23662e70cb9d677d6925a10f10da734b4c21d80

Observation 1ae5de45-65db-4af0-9cce-29c3e93d268a · inbound

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning cites this paper.

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:13.255347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:13.255347Z digest=sha256:8bde207ed5aefae586a42aee803b582ded7def5c721410c643ed62893f238a4e

Observation be8edc1b-2d84-466a-aeca-f54a9f5890cb · inbound

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step cites this paper.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.315309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.315309Z digest=sha256:e8b88d99c4edc59c6c1f57d6eb801a3c5fbeb28f815c61dcc2cb247ad2deed5c

Observation e7fa5bfa-b430-4b78-a825-3df5b72de4c3 · inbound

Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages cites this paper.

Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:49:27.751399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:49:27.751399Z digest=sha256:25c567d6d42de9e2f256054429a7b7cbb05097e7ae5a9e2b0ceaac34316837bb

Observation 30af0243-2239-4962-b88d-f4ad59925137 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.721346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.721346Z digest=sha256:c776a37c1aa3531bead63648a279510cb479e21aca9b78f84b62e33d0d6ef12e

Observation 91d6dd8e-0e46-4c22-ad3b-e4a5371bbc10 · inbound

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information cites this paper.

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T13:25:52.021917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:25:52.021917Z digest=sha256:c69ad905c901195467432201b010a75a58a569c65bed6de0435353210f01001e

Observation 1bce7f0a-c59f-4953-aeb5-9971afb79efa · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:a81de1d2d5ca0ed86cd995236204a9e3bbd0c6f880d595cf39f8a971d8640a3f

Observation 4f4307c2-cea9-4983-a12e-5a00a1d79f35 · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.393381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.393381Z digest=sha256:999bd9f821407341d15f239f6b69df4fa0c1c68fdaebfe8649d0c79a2953266c

Observation 290e26b9-8ff6-43c0-ada4-d2130980bf9c · inbound

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models cites this paper.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.321272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.321272Z digest=sha256:9838f4942fc79d340b720899bfc2ccf8541bd699d3fdbd424643dadf6d5bcbf5

Observation 48c1bd92-9ef3-44f5-81fb-ac6f08efd1f9 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 181

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:fbf27b8405344320506a0b48f94aa10af0232c59aa5f147f688f749fe78cffe1

Observation 794dfe46-4d9f-44d7-bc45-92f56829af41 · inbound

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs cites this paper.

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T18:42:39.023650Z digest=sha256:4ec9b1fb613fb2e6b39e18193db240a4b2d9826945fd6e20fe232dd3b127e57e

Observation c5476062-0026-4499-be03-a071d6d1d788 · inbound

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning cites this paper.

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:59.458706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:22:59.458706Z digest=sha256:adc95ac464a49036febda535369a104bcb7456cd4b9706eb06940a48a2df3889

Observation 6883d3c8-a050-40a8-8301-f9d80190edc0 · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:58.907511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:58.907511Z digest=sha256:ca8c6910f21c3566793f8d9c0cff64a970f773702b073a67f8e48829dbeb1dcd

Observation 93167904-f699-4575-8d5c-3297e29e1fcb · inbound

Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries cites this paper.

Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:54.505939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:59:54.505939Z digest=sha256:a6da9355c0bfa27526435c93e285052fe376a81e2844587445ef803fa0abdf9a

Observation 1122bdd9-2470-4da0-a66b-0cc656f55190 · inbound

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning cites this paper.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.944799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.944799Z digest=sha256:dcd2a535a6d4d85b6499a60e0e03b828f5fdd116a7ea3487debfb1034b4c89f3

Observation c01072d5-3cee-4766-bfb9-26bbf4b826a7 · inbound

Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models cites this paper.

Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:19.624700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:19.624700Z digest=sha256:7a45a36df2c8c85fb7cc545f412a512e9a8ebd4314f6cffddd1c2b968b6f8d36

Observation da8f3bc5-b4f0-4027-9da9-3c0cee68f900 · inbound

APT: Improving Specialist LLM Performance with Weakness Case Acquisition and Iterative Preference Training cites this paper.

APT: Improving Specialist LLM Performance with Weakness Case Acquisition and Iterative Preference Training Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:11.100055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:07:11.100055Z digest=sha256:29164b5548251e8931c4946ebf5f514664a7bd7305da04db4a55aa039fb3ac2d

Observation 96a52f76-7061-4763-814e-a502ca788c4c · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.300557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.300557Z digest=sha256:4d06c53fda02bf1c3a3234ab7882d0abf5d23c965b1739622500300ac69ec909

Observation d19fd706-3403-47a2-ad12-cb865ce56a1b · inbound

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization cites this paper.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:43.248348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:43.248348Z digest=sha256:a4be2fae1aae8c44f7c35cc314894d1a9b921ada171321cf36c4f5c0ac5ef91a

Observation b40b769e-4c06-48e5-9d9d-e10fc34fc93c · inbound

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning cites this paper.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.795071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.795071Z digest=sha256:7ab64c4033acb83167e512f7dd5b1ddf7166565b198d1f1eefc356f41d2ff78a

Observation d9653dec-3380-4c78-8316-8b21d5210f57 · inbound

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation cites this paper.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.558586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.558586Z digest=sha256:16b9bd2d1250e352744ce78c2a8d47de532a6b0b6927defa71f59a694033461e

Observation 5a898ff4-7998-4ad4-9dd9-40d056379cdf · inbound

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations cites this paper.

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.162976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-19T06:09:26.269452Z digest=sha256:d0615d1fdcee109831a89826eae6679577e24528501ca22dc3863db9be69bba0

Observation edf792c8-05a7-415b-8318-369e1b8a0856 · inbound

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization cites this paper.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.880112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.880112Z digest=sha256:7357e6c0165a42ce25162d27fa838d0751de2e99c9e6047df4a12c193bdfd44f

Observation 00092c23-b383-49c4-b201-bfa48500b59f · inbound

Mitigating Object Hallucinations via Sentence-Level Early Intervention cites this paper.

Mitigating Object Hallucinations via Sentence-Level Early Intervention Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:35:32.638911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-25T08:31:24.173135Z digest=sha256:69dded392984ab95993476a7ac039b99d99217574b60c50d729f7d6835309b29

Observation c3f08595-d6ec-47d1-b174-aa944314cf68 · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.477083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.477083Z digest=sha256:cccf60975dd2fc9b29d2e0ca764f5a7e5f4f42cf6818fa8b007ee49b5f989b28

Observation 1e3cd6c3-9c29-4b23-8a4a-9af3c335d1af · inbound

Sample-efficient LLM Optimization with Reset Replay cites this paper.

Sample-efficient LLM Optimization with Reset Replay Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T23:38:08.044450Z digest=sha256:710613169ae794a458355e61ba2bfb63b3598e20b11854748b47370046f6f3a4

Observation 564cad11-567b-4dad-a399-c7f3cfefca93 · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:bead2d34c5113068c630f47414594ca50a48712bb5fb34d38bf86cc04972dc71

Observation 808fbd4d-c3e0-4f1f-88bc-b8bff91c6cc2 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:25554ef38efcc6bc1cc1373b1a0736fb4a2c5c130d7983ce97e1615a82bd232b

Observation 49446334-2c1d-4385-a9e1-a519bf67dd97 · inbound

Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning cites this paper.

Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T14:41:24.891052Z digest=sha256:14f8ec08ae55b371df60ef958748999b0094e9cf91575823ff70d8b200e052ee

Observation eec23df6-7e9f-4b05-95c8-180768cf78a1 · inbound

SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T09:33:57.037949Z digest=sha256:ad5012f81426cc095c6f5e40830d264118f5ade771733f975f614d9c9bbd8b42

Observation c015ebd5-0bb5-4d99-86fc-21eb80bd959a · inbound

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.480368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.480368Z digest=sha256:f053de0ff608a4fe011cbc31ff497f58bfac4b54a57121866d1d1ebdf45f2d92

Observation 1c37e923-e5d4-4838-b5d0-fba9af9e4bdc · inbound

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling cites this paper.

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T03:15:57.744706Z digest=sha256:af31502fe761a3c31e1736fe7b9ab656c9a09acb94ccdb99274337f23ef57130

Observation fb12e888-3d24-44dd-b835-dc5ce12a995f · inbound

Hard Negative Sample-Augmented DPO Post-Training for Small Language Models cites this paper.

Hard Negative Sample-Augmented DPO Post-Training for Small Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T21:57:16.279587Z digest=sha256:e0ed099b0a0096135e4546278c215ab544872d5c18c186ff1f98c1ee0bf4eee1

Observation 3aee8528-579a-4fed-8205-11dc1dea540c · inbound

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation cites this paper.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:29.025781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:29.025781Z digest=sha256:a5f55e24dcb65278d70c3170189f27564b2672aec151ab353d5efa51ee97ef2f

Observation a7b9b8bd-b4c5-4748-82ce-a4918d724865 · inbound

Curr-RLCER:Curriculum Reinforcement Learning For Coherence Explainable Recommendation cites this paper.

Curr-RLCER:Curriculum Reinforcement Learning For Coherence Explainable Recommendation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T19:43:11.384787Z digest=sha256:3b5b01408fb1028dc5ddf9b48812af0397eec368eb8e621857ff5c190b43c256

Observation cddd1db6-ec81-4b3c-a7e5-c64dcf347b56 · inbound

Decomposing the Delta: What Do Models Actually Learn from Preference Pairs? cites this paper.

Decomposing the Delta: What Do Models Actually Learn from Preference Pairs? Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T17:22:00.563479Z digest=sha256:15c59fabba89a373b51021375b0bc27cdea5a3c0877ca84bc0d71c7cf703656b

Observation f443fea7-b8cd-4791-b487-5ca6d634b5ac · inbound

On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization cites this paper.

On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T19:35:59.832868Z digest=sha256:5fd694596939a3f4fa2c716c026f5d9a4f01d9deaf8be0a935ab2a99756c11f9

Observation 6c56e097-7310-4955-acda-b5063cb3d4bd · inbound

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding cites this paper.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-09T16:29:05.186607Z digest=sha256:3c6588ecf06d443e04bc2928e590472bc7f1bfda0c6e0add23dfb01f18fa8028

Observation 4cdde1c5-a6fa-4195-bf6e-12b1b00d862e · inbound

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction cites this paper.

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T01:22:15.608348Z digest=sha256:216e4c1d77c21c29504e454247006f5b6d6b400c9850be7376807a9a593c2cfe

Observation e145583c-698b-40c5-8677-aaac3491cc14 · inbound

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models cites this paper.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:455fab4b690a5b36bad2a7139bda5fbb5868f8cba606bc670a2d0df48f029589

Observation 8970bf92-f524-4ea3-b4d6-39a3f68b7ccf · inbound

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge cites this paper.

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T04:19:00.238602Z digest=sha256:b18af74a67d1251cbfc8618699a8aa82a65e9165b87660c3f2ccce59f43c3dd5

Observation f459a32c-788c-4e36-b98d-4479abf8e8f0 · inbound

YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning cites this paper.

YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T05:29:42.381280Z digest=sha256:f78d3380ebb92aadb69e3f368a20768e91442cc9ac8a48d16e9418e3cf8251d2

Observation d0349acd-d333-4845-8385-e9e4ec6bbcf1 · inbound

Enhancing LLM Metacognition via Cognitive Pairwise Training cites this paper.

Enhancing LLM Metacognition via Cognitive Pairwise Training Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:02:33.922491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T19:01:18.153145Z digest=sha256:84f18a939446ee8f8b3a05982e3a0871b42fa7f4480a9e66512cbfd6cccb4c2e

Observation b7776c39-e0cf-4b81-9ecd-8433735cd0b0 · inbound

DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling cites this paper.

DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T16:57:10.030422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T22:16:01.457276Z digest=sha256:25ff372cb6a422df464d5a4a0e2a62dc59b3c16f2bb9ca84c80e7d3e89afc4ef

Observation d6322ad5-456d-46c5-b859-1e7907ebcf31 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 178

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:25.924065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:39:44.696961Z digest=sha256:8af191b6df5fda43e2e8654f1d32c2cad20b4b4afc8e217456d590e472df30a4

Observation 2ff953ea-7321-44d2-9486-31e62b458d5e · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-02T12:05:18.089147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:05:18.089147Z digest=sha256:d2575fec3d0bf8c4cfffaf219d24350ac76084ad68c9248dfd492f88a35efab5

Observation 765c1947-6e7a-40f9-a237-21b4d3aeec03 · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:27:36.668377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:c212b137b3d750e3b01c1f86bd30e6b61fde70949d50afa6f49018b74116cb6e

Observation 0138b81b-b3bd-49ee-90ad-45c393ff9032 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:37:49.416543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:ef9eed077e8b5ddfc42eb8b03b4fc9af0391ced39c1854c6b9ca4af5b76ebd09

Observation d88db0e8-6e3c-4970-8b32-4f9170e4df84 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:27.879246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:27.879246Z digest=sha256:256f3af4502b207ac2d29b325a111db0e29578883fdf63bfc9a35cb60913c30f

Observation 165fabcd-4440-4baf-a1cd-42f140a0680d · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:08:33.402998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T06:39:34.199607Z digest=sha256:47059cc8a474c6083152af5e6f472fb971fbc8b2dced9e8b5cb562bdd383ef83

Observation 1a30c586-887c-4550-8011-78f8f7c1b6de · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:19.127147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:19.127147Z digest=sha256:6319631e8218f7a5f259c2ff3fb8a60b678cbdd7d6c864b0e191f873a93ad5d0

Observation fa580899-0bc5-4646-a750-1e0487fdcf0a · inbound

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence cites this paper.

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:38:44.174836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T03:57:19.028507Z digest=sha256:79d1b90f19fb2855c62dc58d43af27bc7ace0d7cf7693d1e44b7c965e125a9dc

Observation 6d9ae9eb-77f8-48ae-80d4-99f0a7491913 · inbound

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents cites this paper.

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:40:07.843363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T19:55:51.114244Z digest=sha256:8dcecca7009be76b4516ecbba2faa50d611f5d16bed19de88a429446491c79eb

Observation d4408eb1-6fd1-4e65-a041-39d1f3336854 · inbound

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM cites this paper.

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:33:28.303470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T13:27:45.796650Z digest=sha256:fbaad4e72da6a33b4f3c6a0efcca800e3867b87e4e74f71ba65cb858ec83916c

Observation dbbdc301-8aa1-432d-8792-d5fdf14ce555 · inbound

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models cites this paper.

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:09:50.060500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T05:32:59.640335Z digest=sha256:7c6eef71034226651d75e5d00e637a1e8b9cb5a1f20dcd0ddada3d54a7d751a2

Observation 1187e579-c9d9-4db7-bdca-02c9bfc88a57 · inbound

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models cites this paper.

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T17:23:45.531776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-29T05:19:07.877346Z digest=sha256:85da7eec9f59f8e45f8d2c016ba5160cec0cdf9de823f8d171bd6f72b4cc63f8

Observation bcd9bee0-47a9-4340-8155-476652639379 · inbound

Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement cites this paper.

Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T08:04:28.377800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T07:55:28.254309Z digest=sha256:5aed9ee801bc3b2e6b67a4597451ff8736049969fd133afd35005834facaf7fb

Observation cfa2c755-5b1e-4f8a-9995-6e32ecd56d6b · inbound

StrucTab: A Structured Optimization Framework for Table Parsing cites this paper.

StrucTab: A Structured Optimization Framework for Table Parsing Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:24:18.922918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T06:22:21.694355Z digest=sha256:bc208c5da0321bf10fff990e5cac750397c32dc0fd8c0b2e7c07039259dfff67

Observation 08ae0c01-cea2-4452-afda-ac733a5a2c7d · inbound

Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning cites this paper.

Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T02:23:04.515705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:23:04.515705Z digest=sha256:d2c6ed802e5763c4dddd37a4187374800251d59058e24056ab60375d06d3084d

Observation 52443909-5f8f-4eee-b2b7-9b01f26e98cb · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:75547f133ccb0b30f24b1213153d5ff38dd7528107474b9e9284b91b0d09aa60

Observation 85614d0d-d12e-484e-976e-8ffddea0d4f4 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:42.479453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:42.479453Z digest=sha256:df0dce87abb8661794b12c1a7640f38500dc3f3e0e418995ded29d6a690517d7

Observation 87b2b2cf-224f-4ae4-9ad6-dde87fe12866 · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:0cc3f7f4e5f9438c23f9664f03fc04edaa8b86c68fca57b03c192997e571dc90

Observation 6ad3e2a1-1796-4da4-ae1e-a62d7563a082 · inbound

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding cites this paper.

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T13:42:24.987811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:42:24.987811Z digest=sha256:69260ceaf4528aeb4db7fb5e4e218cd1203f5b3f8f22f5ebb5ff76e16186d046

Observation e050620a-34d5-46d5-b4d6-852f9bc765c1 · inbound

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization cites this paper.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.114561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.114561Z digest=sha256:2d17446f40f20e56646811278fce01148af21eb0cfa71ba9bcc2499473f387df