Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T23:58:29.040819Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 75 inbound Pith citation observations for arXiv:2406.18629.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T23:58:29.040819Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:20:13.925085Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
38 of 38 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation 6cc6da4a-7f20-472c-93af-ce641a20af4b · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5f81d01a-12e0-41f1-9710-d4a4fd8518b2 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Llemma: An Open Language Model For Mathematics
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8911546b-1d35-42c6-abb5-1068da101f26 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Qwen Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 95df9dbc-b51e-4716-84d4-a1ff874671d7 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs AlphaMath Almost Zero: Process Supervision without Process
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 036242d0-3f62-4f0a-9799-50689a486566 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0d4146c7-3b79-4bed-a7f4-c373ae647493 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4fdd0d13-c785-45f0-b0d0-cd0d1c5cbc16 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Measuring Mathematical Problem Solving With the MATH Dataset
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e81133a8-39b6-4626-b85e-6c14d1e98abd · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs ORPO: Monolithic Preference Optimization without Reference Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aa8c5b32-a312-4b29-ac0b-f19dd3af5479 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Common 7B Language Models Already Possess Strong Math Capabilities
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 489efe0a-d477-4c8b-94b2-47078a1d17a9 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c0eaa2a6-8da8-4363-8916-3cc0dc939799 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Let's Verify Step by Step
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4b075053-92a5-4a37-8f89-3f69aefd8a91 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Rho-1: Not All Tokens Are What You Need
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bb851368-450b-43be-a9a0-49a5e3f9781d · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fb61548a-ef33-449e-a772-18361de229fc · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Augmenting Math Word Problems via Iterative Question Composing
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9707572d-5283-4e31-93c5-8370a9918894 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Improving Large Language Model Fine-tuning for Solving Math Problems
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7ca1941d-7477-4e79-a4dd-3ff3a5abb369 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b50c459f-c26c-4ef3-86ac-73c93556487d · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bd279192-e885-45f5-9df8-41ea13196fb9 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Orca-Math: Unlocking the potential of SLMs in Grade School Math
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0d86d06d-2721-4ab5-af11-1b427c1e8f62 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2b98df74-23c5-4bb5-8021-7ac74090b29d · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Code Llama: Open Foundation Models for Code
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e8ed9e34-afb2-4b82-8dd0-7d5de9925bf6 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6de8ae2e-03b9-44ec-a1c8-60d64d7d8cbd · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MathScale: Scaling Instruction Tuning for Mathematical Reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 889dc090-5a34-4059-9d08-6d37ed4e0238 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d17fe6da-b6cc-4608-972f-a4cf8b53a4e7 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1a383ce5-5ad3-408e-8c76-85838ec2650c · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs LLaMA: Open and Efficient Foundation Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ec1fd29b-da13-44ce-961b-f320537be3cc · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Zephyr: Direct Distillation of LM Alignment
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d0abb843-beab-40b5-a7d1-4f466f8d5266 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b9903deb-c0fe-4bc6-b8b1-85f6d9478e51 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b0635685-0424-427f-954b-7dae3584314e · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 26791200-3458-4d05-8142-c370e8dc9306 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9e3522f4-4a64-419a-8bb9-b4f87a7d88a2 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Answering Questions by Meta-Reasoning over Multiple Chains of Thought
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 859f9129-73d9-43d8-84d4-4772f4109909 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 48280907-40fd-4801-94bb-c33f2fdc1998 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 94c1f6e1-2a00-4a6c-9985-8f2832498306 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 923df388-c29d-4125-bfce-3ad63a44236e · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs MAmmoTH2: Scaling Instructions from the Web
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 17f98e61-12f4-4375-90f7-b28e009959fa · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 80e66f49-39ce-4d1d-82d9-3d72e932817a · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b76230c7-eec0-420b-a315-402e3895a086 · outbound
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fa4bf688-ed37-43a3-8817-050c10a3402a · inbound
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5b0a0db2-e7bf-4f99-b7c4-43cfa6aff57d · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bf05dcc1-26d7-43c1-8876-d01f3cb28f33 · inbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ee085e-1b7d-49ac-a12d-16e73d18f24b · inbound
Preference Optimization for Reasoning with Pseudo Feedback Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90bd06fd-699c-4981-a096-0ddeb21a5be2 · inbound
Mars-PO: Multi-Agent Reasoning System Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 232c31f4-a495-4526-85c3-39f190e2e66e · inbound
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ad7b07-5d02-43ae-a588-8ce714176bf7 · inbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6784637-1a2b-4c90-840f-60079620f06e · inbound
A Systematic Examination of Preference Learning through the Lens of Instruction-Following Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 076dad93-3150-45f2-97f7-fa92d32b0265 · inbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca2d97d6-65c3-449a-a8e3-ee73b60c5396 · inbound
Plug-and-Play Training Framework for Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93ce969c-ecec-4fb8-aeb2-2c5092d15dd2 · inbound
Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c88df6c2-b598-4e56-bd5f-faeed1545a19 · inbound
SDPO: Segment-Level Direct Preference Optimization for Social Agents Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae5de45-65db-4af0-9cce-29c3e93d268a · inbound
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8edc1b-2d84-466a-aeca-f54a9f5890cb · inbound
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7fa5bfa-b430-4b78-a825-3df5b72de4c3 · inbound
Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30af0243-2239-4962-b88d-f4ad59925137 · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 169
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91d6dd8e-0e46-4c22-ad3b-e4a5371bbc10 · inbound
LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bce7f0a-c59f-4953-aeb5-9971afb79efa · inbound
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4f4307c2-cea9-4983-a12e-5a00a1d79f35 · inbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 290e26b9-8ff6-43c0-ada4-d2130980bf9c · inbound
Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c1bd92-9ef3-44f5-81fb-ac6f08efd1f9 · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 181
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 794dfe46-4d9f-44d7-bc45-92f56829af41 · inbound
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c5476062-0026-4499-be03-a071d6d1d788 · inbound
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6883d3c8-a050-40a8-8301-f9d80190edc0 · inbound
Large Language Models for Planning: A Comprehensive and Systematic Survey Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 124
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93167904-f699-4575-8d5c-3297e29e1fcb · inbound
Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1122bdd9-2470-4da0-a66b-0cc656f55190 · inbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c01072d5-3cee-4766-bfb9-26bbf4b826a7 · inbound
Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da8f3bc5-b4f0-4027-9da9-3c0cee68f900 · inbound
APT: Improving Specialist LLM Performance with Weakness Case Acquisition and Iterative Preference Training Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a52f76-7061-4763-814e-a502ca788c4c · inbound
A Survey on Large Language Models for Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d19fd706-3403-47a2-ad12-cb865ce56a1b · inbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40b769e-4c06-48e5-9d9d-e10fc34fc93c · inbound
MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9653dec-3380-4c78-8316-8b21d5210f57 · inbound
Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a898ff4-7998-4ad4-9dd9-40d056379cdf · inbound
From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation edf792c8-05a7-415b-8318-369e1b8a0856 · inbound
Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00092c23-b383-49c4-b201-bfa48500b59f · inbound
Mitigating Object Hallucinations via Sentence-Level Early Intervention Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c3f08595-d6ec-47d1-b174-aa944314cf68 · inbound
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3cd6c3-9c29-4b23-8a4a-9af3c335d1af · inbound
Sample-efficient LLM Optimization with Reset Replay Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 564cad11-567b-4dad-a399-c7f3cfefca93 · inbound
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 808fbd4d-c3e0-4f1f-88bc-b8bff91c6cc2 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 49446334-2c1d-4385-a9e1-a519bf67dd97 · inbound
Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eec23df6-7e9f-4b05-95c8-180768cf78a1 · inbound
SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c015ebd5-0bb5-4d99-86fc-21eb80bd959a · inbound
TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c37e923-e5d4-4838-b5d0-fba9af9e4bdc · inbound
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fb12e888-3d24-44dd-b835-dc5ce12a995f · inbound
Hard Negative Sample-Augmented DPO Post-Training for Small Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3aee8528-579a-4fed-8205-11dc1dea540c · inbound
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7b9b8bd-b4c5-4748-82ce-a4918d724865 · inbound
Curr-RLCER:Curriculum Reinforcement Learning For Coherence Explainable Recommendation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cddd1db6-ec81-4b3c-a7e5-c64dcf347b56 · inbound
Decomposing the Delta: What Do Models Actually Learn from Preference Pairs? Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f443fea7-b8cd-4791-b487-5ca6d634b5ac · inbound
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6c56e097-7310-4955-acda-b5063cb3d4bd · inbound
Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4cdde1c5-a6fa-4195-bf6e-12b1b00d862e · inbound
MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e145583c-698b-40c5-8677-aaac3491cc14 · inbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8970bf92-f524-4ea3-b4d6-39a3f68b7ccf · inbound
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f459a32c-788c-4e36-b98d-4479abf8e8f0 · inbound
YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d0349acd-d333-4845-8385-e9e4ec6bbcf1 · inbound
Enhancing LLM Metacognition via Cognitive Pairwise Training Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b7776c39-e0cf-4b81-9ecd-8433735cd0b0 · inbound
DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d6322ad5-456d-46c5-b859-1e7907ebcf31 · inbound
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2ff953ea-7321-44d2-9486-31e62b458d5e · inbound
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 178
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765c1947-6e7a-40f9-a237-21b4d3aeec03 · inbound
TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0138b81b-b3bd-49ee-90ad-45c393ff9032 · inbound
APPO: Agentic Procedural Policy Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d88db0e8-6e3c-4970-8b32-4f9170e4df84 · inbound
APPO: Agentic Procedural Policy Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 165fabcd-4440-4baf-a1cd-42f140a0680d · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1a30c586-887c-4550-8011-78f8f7c1b6de · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa580899-0bc5-4646-a750-1e0487fdcf0a · inbound
Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6d9ae9eb-77f8-48ae-80d4-99f0a7491913 · inbound
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d4408eb1-6fd1-4e65-a041-39d1f3336854 · inbound
Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dbbdc301-8aa1-432d-8792-d5fdf14ce555 · inbound
DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1187e579-c9d9-4db7-bdca-02c9bfc88a57 · inbound
DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bcd9bee0-47a9-4340-8155-476652639379 · inbound
Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cfa2c755-5b1e-4f8a-9995-6e32ecd56d6b · inbound
StrucTab: A Structured Optimization Framework for Table Parsing Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 08ae0c01-cea2-4452-afda-ac733a5a2c7d · inbound
Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52443909-5f8f-4eee-b2b7-9b01f26e98cb · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85614d0d-d12e-484e-976e-8ffddea0d4f4 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b2b2cf-224f-4ae4-9ad6-dde87fe12866 · inbound
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad3e2a1-1796-4da4-ae1e-a62d7563a082 · inbound
DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e050620a-34d5-46d5-b4d6-852f9bc765c1 · inbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.