Pith. sign in

Paper Citation Record · LEDGER

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.15706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15706 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:29:35.826782Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c61dd3a6-0f87-4fba-bed9-78da6f6c6430 · outbound

This paper cites Qwen Technical Report.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.069984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.069984Z digest=sha256:02decc3dbfb89fa24cdfcf9f2e127f18072bdd20f0cfa2af543add58994fd33a

Observation a9ad4ebe-4bfa-4155-a22e-7471c5277dd7 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning ORPO: Monolithic Preference Optimization without Reference Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.631728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.631728Z digest=sha256:94857f1044c62434033451ba944794c2ed0a7b3fc355564f6b6020b378b06f7e

Observation 3d4f9a7d-1972-40c9-b23c-eae567766c12 · outbound

This paper cites MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.957668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.957668Z digest=sha256:c61a257ab7de941cfdde64a547bc094917c56b84ac8e0dfa8023f651b4cf8784

Observation 946a8099-ad50-43ce-ada0-929c11b49985 · outbound

This paper cites Orca-Math: Unlocking the potential of SLMs in Grade School Math.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Orca-Math: Unlocking the potential of SLMs in Grade School Math

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.252687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.252687Z digest=sha256:2557b6a56d1908bd918867fb83cfbe3c078229bc61c87a19263f1ec164f7b29c

Observation 2d94a3e4-526c-4e4e-a214-63c1098c235c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.566083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.566083Z digest=sha256:19e33f8981bd6245c7496633c858a1e474979c471dbe173de0836f7d84a1d350

Observation 3b4612a2-3425-4484-9dae-1f571c030c6a · outbound

This paper cites MathPile: A Billion-Token-Scale Pretraining Corpus for Math.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning MathPile: A Billion-Token-Scale Pretraining Corpus for Math

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.680196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.680196Z digest=sha256:1e2d764f971b1750bcdc41480967c517fc335739bd1ef444ecac68a34c593233

Observation f6560daf-019a-4a7a-85c5-15fd3c253c84 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Automatic Chain of Thought Prompting in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.826782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.826782Z digest=sha256:6a193cd7469e7c29847a5a8ce0e5a035368d9f0a2287acb8b412c008ff65d3d6

Observation 9a397e16-b92e-468e-8660-2dbfc3d1c1e6 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Training Verifiers to Solve Math Word Problems

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.220523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.220523Z digest=sha256:914cc26d4391e662942151165a23269bef22bdf5e9961dcff970e1307a8f80a1

Observation ada7517d-270b-424d-bb77-6823d4c674a6 · outbound

This paper cites MathScale: Scaling Instruction Tuning for Mathematical Reasoning.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning MathScale: Scaling Instruction Tuning for Mathematical Reasoning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.374285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.374285Z digest=sha256:c36192d12e83bd78a1950cb17ee74bf87e51d78bea386fc2076b25d9ca698143

Observation afa863ba-4dbf-4806-9fc1-8135714b49ef · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning KTO: Model Alignment as Prospect Theoretic Optimization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.363833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.363833Z digest=sha256:c951004520b34c6bd2b1737d3988bd0de716454f27cd8d84b4e1c472a920c49e

Observation b40b769e-4c06-48e5-9d9d-e10fc34fc93c · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.795071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.795071Z digest=sha256:02526bc2a1d2b1faa11821143aa7eaba2422d5a73062f34e960cec98eafea3ea

Observation 83c80812-fb4f-4f46-8b67-046aef42f866 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.481780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.481780Z digest=sha256:0ebf0ffe4bbe2a4900eabdd445283b5c05314fb75658b9eced3b7911f30bc725

Observation 930ec7a3-8608-4099-8aa6-f78e94c68117 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.084817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.084817Z digest=sha256:bc7576ae7f9baa6d59bf80e766ea97f312ddcf4850e73b3e6c967a0ca94310ae

Pith citing papers

No inbound Pith citation observations are available.