Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.10858.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:40:39.974766Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T09:42:04.234036Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation ded0495b-0a64-467b-9aa0-caf6eaca9810 · inbound
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents Step-level Value Preference Optimization for Mathematical Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 547c4f5f-6ffd-45f1-9c91-fac1ba3bd1f3 · inbound
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Step-level Value Preference Optimization for Mathematical Reasoning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6f125f57-580a-479a-98b2-a7d853c4db5d · inbound
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Step-level Value Preference Optimization for Mathematical Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b40e4971-2e58-423d-a19f-11be2b6f416e · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization Step-level Value Preference Optimization for Mathematical Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 183659f9-af00-406b-ad67-4fea6c4eb65c · inbound
SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba Step-level Value Preference Optimization for Mathematical Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 54ce838a-b3af-4aa0-ae8c-c181090fcbff · inbound
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning Step-level Value Preference Optimization for Mathematical Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eec809d2-b0ab-4299-903e-b35cfb1d93e6 · inbound
APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation Step-level Value Preference Optimization for Mathematical Reasoning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3dd2b738-438b-4692-8fc8-673e0334ace8 · inbound
Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Step-level Value Preference Optimization for Mathematical Reasoning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7080c4b6-4ef2-4eaf-87ae-ec831fef7d5b · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9acd4b2-4ebd-42bf-90a9-62837c84ed50 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.