Pith. sign in

Paper Citation Record · LEDGER

Iterative Reasoning Preference Optimization

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2404.19733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.19733 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:49.338434Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 169cde8f-8921-4e81-b7f9-13c9b775fde9 · inbound

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive cites this paper.

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Iterative Reasoning Preference Optimization

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T23:04:44.466816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T23:04:44.287660Z digest=sha256:2bb2e4971324a96c77e6146a6130e9b25a84debef08264cf534d53156408bdf1

Observation ba3dc64d-5934-44c6-af64-4f35cbbaee7b · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Iterative Reasoning Preference Optimization

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:16:17.391561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:5d5d883c2ea74a0cdf32e92328a2148649d102727e8cfdfc2e5b57cdcba9635f

Observation 8e87f634-90c2-4fe7-84c1-96679874929b · inbound

Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search cites this paper.

Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search Iterative Reasoning Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:07:30.594150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T04:06:23.521344Z digest=sha256:b69952557fd8cfc67c482d6b47ead090e0b6bf871451ec5a1b82fa2ddc04be69

Observation 08114dfe-9127-43df-b5e1-9ec5f849d039 · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning Iterative Reasoning Preference Optimization

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:37.170645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:d8cc81e7132ebc6bcb2775dd0cb356b6cd0621224c7052b2f2a080184b87a80e

Observation d098173b-6f47-467a-becf-6decefae2354 · inbound

Online Knowledge Distillation with Reward Guidance cites this paper.

Online Knowledge Distillation with Reward Guidance Iterative Reasoning Preference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.338434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.338434Z digest=sha256:d232cf50c8ddaa406f0274472649b073185554a8e2396f5e974beb80ec1ca3a2

Observation ef665640-8de3-49d2-8e93-6562c9d2f7a0 · inbound

Frictional Agent Alignment Framework: Slow Down and Don't Break Things cites this paper.

Frictional Agent Alignment Framework: Slow Down and Don't Break Things Iterative Reasoning Preference Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:01.534554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:01.534554Z digest=sha256:2a4267d8f7aff3cff76e8e0cf1049df99c5ef6175720971e21fd334593ffab20

Observation 6b8c74b3-c7d2-4fd9-9fa0-e88b8aaf2d78 · inbound

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning cites this paper.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Iterative Reasoning Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.881145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.881145Z digest=sha256:4d504788a546fd80e885eba1b68b58f5c24ca0414227c3c8177cff34313ff1ce

Observation f5605bbe-f396-4310-8d6b-4e4fd10e59b7 · inbound

Control-R: Towards controllable test-time scaling cites this paper.

Control-R: Towards controllable test-time scaling Iterative Reasoning Preference Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:20.784306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:14:20.784306Z digest=sha256:facb3986ea6012e32ddef324007308bfcc48ba2ac89e35cbaa4da1c062ed8eae

Observation 2e3e4396-081c-46b7-96e1-774db6aa8875 · inbound

PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization cites this paper.

PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization Iterative Reasoning Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:47:50.045904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:47:50.045904Z digest=sha256:dabaec672a4af6b1c7788b26539daf6f4ec56ddce822f199276497e7e0547298

Observation 1c8bc8ab-dfff-4b38-b740-7d91927a1c8d · inbound

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment cites this paper.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Iterative Reasoning Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.682978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.682978Z digest=sha256:027e9bd155c693ff99a9bf72219a8e42f6a9a9a46bcc321ed14bbd30ccb1d480

Observation c3fc33cf-e7f8-400a-914a-28d445eab0cf · inbound

Optimising Language Models for Downstream Tasks: A Post-Training Perspective cites this paper.

Optimising Language Models for Downstream Tasks: A Post-Training Perspective Iterative Reasoning Preference Optimization

Reference 166

Resolution
unresolved
no resolver link, observed 2026-08-06T22:44:44.374791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:44:44.374791Z digest=sha256:cb20c8b5822d3cd1e6f71af526b9b4237648d93c4245e4bbb10fa99ee4b36524

Observation d2bfe77f-6232-485e-94f8-8b69dbbaa289 · inbound

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought cites this paper.

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought Iterative Reasoning Preference Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:37.514013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:37.514013Z digest=sha256:a844ff16533bf521bd90e3ca498009e08d3010becb052ec9d2b3f1cc15c91fa1

Observation 2e72cd4f-3d77-4dc0-bd7a-2c2772aac227 · inbound

Technical Report of TeleChat2, TeleChat2.5 and T1 cites this paper.

Technical Report of TeleChat2, TeleChat2.5 and T1 Iterative Reasoning Preference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:22.729159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:43:22.729159Z digest=sha256:f4ae0c2488ee993653338706ded35c23181bae5c4a68fcba521ed4be68640a53

Observation 32f8d4a6-ad75-4157-aa33-a7f4d287f706 · inbound

Learning to Configure Agentic AI Systems cites this paper.

Learning to Configure Agentic AI Systems Iterative Reasoning Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.393255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T13:06:56.207692Z digest=sha256:a91056542790046f642e3992ab2572bdddd4f1161fc77d960d68f8281aa8b2e1

Observation 0a03a2dc-8c2b-4957-9d60-27c9f5ecf119 · inbound

Learning to Configure Agentic AI Systems cites this paper.

Learning to Configure Agentic AI Systems Iterative Reasoning Preference Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:31:29.563857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T11:29:45.134060Z digest=sha256:cded3cc1f3592dcf7016d6e016b551724b89a22aed41af2116f5bb41c2fb1b76

Observation 225b7731-7a20-46dc-917b-d9c40ee5da45 · inbound

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models cites this paper.

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models Iterative Reasoning Preference Optimization

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:41:02.247061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:38:01.208136Z digest=sha256:cfb93e0f1b95992e82aa361c15e6377e5bd72df228a1f8e8f770096ccf71cf51

Observation e2d574a8-4dd5-4852-b6cb-b2afba6acddc · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Iterative Reasoning Preference Optimization

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.375076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:a49c1febb7cc8934b2e1efa5dc596eb3cde05cc70fe63c2a8c0bfb4f899fd259

Observation 590e9133-a022-4fef-9426-7f7586cede03 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Iterative Reasoning Preference Optimization

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.577200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:8e5432d3b026eacb7715c1b0704ae33840ea536dbc1605c2a063474af6b9b1d1

Observation a094ac28-9420-49a4-a90c-fe37fd65c469 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Iterative Reasoning Preference Optimization

Reference 186

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.562104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:3eb8c0567191aec08a0fde375a7ddc88b511ad5c25d87be2c9537b3e9838965f

Observation b1a7437f-f596-499d-ab0c-7fc41ea3c2cf · inbound

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training cites this paper.

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training Iterative Reasoning Preference Optimization

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:29:16.717946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T21:07:28.660119Z digest=sha256:8b29dad5d73ddc8ecf756fe7b5afa59969a9c469018e9a86c0dff539931c2413

Observation 69c53d56-ac18-4f74-8d92-7bb60224c84a · inbound

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation cites this paper.

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation Iterative Reasoning Preference Optimization

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:34:26.755731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T08:30:37.016856Z digest=sha256:35da61a7a09cbeff94020b977297f0a061eb5537277f91ce4053cd2d1eb4b392

Observation c380513d-5864-4ce4-8214-e693e6d539a8 · inbound

Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement cites this paper.

Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement Iterative Reasoning Preference Optimization

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:04:28.460655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:55:28.254309Z digest=sha256:cc7c2043a1513b8f794344f8930dad84bd427740dfa7da7486bab40293d49cbf

Observation ca1d4301-fb79-49f9-a511-757e94e3061b · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Iterative Reasoning Preference Optimization

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:72cf69e969d733427aa6b1e3ff95b640be0d5992740261a9aed5495c0df3ffb4

Observation 511bdc20-4e3c-465d-951f-fbe531153738 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Iterative Reasoning Preference Optimization

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.278322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.278322Z digest=sha256:af3c554e933429bfc6b7e38f520fa580ad8479a58c5af03699bab1cc0f4a35c8

Observation a635ae33-147a-4384-a69a-afc62e4a058e · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Iterative Reasoning Preference Optimization

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:36.276989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:36.276989Z digest=sha256:9d900a4fa3c0d625bb680cab3ef0838da893b0dfeacac13ebdad34705d0f4a2a