Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2404.19733.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:49.338434Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 169cde8f-8921-4e81-b7f9-13c9b775fde9 · inbound
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Iterative Reasoning Preference Optimization
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba3dc64d-5934-44c6-af64-4f35cbbaee7b · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Iterative Reasoning Preference Optimization
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e87f634-90c2-4fe7-84c1-96679874929b · inbound
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search Iterative Reasoning Preference Optimization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08114dfe-9127-43df-b5e1-9ec5f849d039 · inbound
LIMO: Less is More for Reasoning Iterative Reasoning Preference Optimization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d098173b-6f47-467a-becf-6decefae2354 · inbound
Online Knowledge Distillation with Reward Guidance Iterative Reasoning Preference Optimization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef665640-8de3-49d2-8e93-6562c9d2f7a0 · inbound
Frictional Agent Alignment Framework: Slow Down and Don't Break Things Iterative Reasoning Preference Optimization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b8c74b3-c7d2-4fd9-9fa0-e88b8aaf2d78 · inbound
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Iterative Reasoning Preference Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5605bbe-f396-4310-8d6b-4e4fd10e59b7 · inbound
Control-R: Towards controllable test-time scaling Iterative Reasoning Preference Optimization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e3e4396-081c-46b7-96e1-774db6aa8875 · inbound
PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization Iterative Reasoning Preference Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8bc8ab-dfff-4b38-b740-7d91927a1c8d · inbound
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Iterative Reasoning Preference Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3fc33cf-e7f8-400a-914a-28d445eab0cf · inbound
Optimising Language Models for Downstream Tasks: A Post-Training Perspective Iterative Reasoning Preference Optimization
Reference 166
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2bfe77f-6232-485e-94f8-8b69dbbaa289 · inbound
From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought Iterative Reasoning Preference Optimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e72cd4f-3d77-4dc0-bd7a-2c2772aac227 · inbound
Technical Report of TeleChat2, TeleChat2.5 and T1 Iterative Reasoning Preference Optimization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32f8d4a6-ad75-4157-aa33-a7f4d287f706 · inbound
Learning to Configure Agentic AI Systems Iterative Reasoning Preference Optimization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a03a2dc-8c2b-4957-9d60-27c9f5ecf119 · inbound
Learning to Configure Agentic AI Systems Iterative Reasoning Preference Optimization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 225b7731-7a20-46dc-917b-d9c40ee5da45 · inbound
S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models Iterative Reasoning Preference Optimization
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2d574a8-4dd5-4852-b6cb-b2afba6acddc · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Iterative Reasoning Preference Optimization
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 590e9133-a022-4fef-9426-7f7586cede03 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Iterative Reasoning Preference Optimization
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a094ac28-9420-49a4-a90c-fe37fd65c469 · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Iterative Reasoning Preference Optimization
Reference 186
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1a7437f-f596-499d-ab0c-7fc41ea3c2cf · inbound
Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training Iterative Reasoning Preference Optimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69c53d56-ac18-4f74-8d92-7bb60224c84a · inbound
Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation Iterative Reasoning Preference Optimization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c380513d-5864-4ce4-8214-e693e6d539a8 · inbound
Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement Iterative Reasoning Preference Optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca1d4301-fb79-49f9-a511-757e94e3061b · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Iterative Reasoning Preference Optimization
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 511bdc20-4e3c-465d-951f-fbe531153738 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Iterative Reasoning Preference Optimization
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a635ae33-147a-4384-a69a-afc62e4a058e · inbound
Test-Time Scaling via Error Localization Iterative Reasoning Preference Optimization
Reference 178
Source-reported events for the cited work
Unavailable: canonical work link unavailable.