Pith. sign in

Paper Citation Record · LEDGER

Step-level Value Preference Optimization for Mathematical Reasoning

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.10858.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10858 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:40:39.974766Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T09:42:04.234036Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ded0495b-0a64-467b-9aa0-caf6eaca9810 · inbound

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents cites this paper.

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents Step-level Value Preference Optimization for Mathematical Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:42:04.237687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T09:41:59.979595Z digest=sha256:b88357eb707091e5b0e2046f5ecf7d869e909fe2ccaad598d464000f6731b43b

Observation 547c4f5f-6ffd-45f1-9c91-fac1ba3bd1f3 · inbound

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs cites this paper.

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Step-level Value Preference Optimization for Mathematical Reasoning

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:51:29.518054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T15:51:29.022336Z digest=sha256:0bc8e87df66a8fa545fa1a4daf04a6fd7f1d89714e6d0de48e563d8dfbafa7bc

Observation 6f125f57-580a-479a-98b2-a7d853c4db5d · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Step-level Value Preference Optimization for Mathematical Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.490615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:9a4dd3b11ef0548dfe1df634a3b48fccf97db7f0fb65fd47ebc37119f96bbdf5

Observation b40e4971-2e58-423d-a19f-11be2b6f416e · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization Step-level Value Preference Optimization for Mathematical Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.805429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:4f2e7b50777223d576d70376a251184a053fb66c5ed241c5c7330890d89077f2

Observation 183659f9-af00-406b-ad67-4fea6c4eb65c · inbound

SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba cites this paper.

SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba Step-level Value Preference Optimization for Mathematical Reasoning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:46:12.439895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T09:44:53.290259Z digest=sha256:e32c99459188f8527228f21f2d11559d1a87f63c53cfac5b2a9dd0c0ffe32aa1

Observation 54ce838a-b3af-4aa0-ae8c-c181090fcbff · inbound

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning cites this paper.

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning Step-level Value Preference Optimization for Mathematical Reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:21:04.436629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T07:20:01.505216Z digest=sha256:d8a83a6d1aac5da7b725009f00d7708473b54eaf6e3f24154e3562e38dcb70bf

Observation eec809d2-b0ab-4299-903e-b35cfb1d93e6 · inbound

APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation cites this paper.

APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation Step-level Value Preference Optimization for Mathematical Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:26:25.191736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T05:22:25.475956Z digest=sha256:791d6c7a1cb2e3afc08be322b9c59621b64b00bb36d29fef7da35ab9a5b8a5d8

Observation 3dd2b738-438b-4692-8fc8-673e0334ace8 · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Step-level Value Preference Optimization for Mathematical Reasoning

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.952562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:39a6a221184e19f590da9cde24d4d12dd76b4952aae96469cb211b415c834545

Observation 7080c4b6-4ef2-4eaf-87ae-ec831fef7d5b · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:569657b76718c02cc929bce7848613c59d82cf9b0198c4bebda531246501e34a

Observation a9acd4b2-4ebd-42bf-90a9-62837c84ed50 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.974766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.974766Z digest=sha256:3a14e8e1eb5a95e1b0c5e37f2f2528aa8a7b8f454d74ebd0a85bfcfcbaf924ca