Pith. sign in

Paper Citation Record · LEDGER

Provably Robust DPO: Aligning Language Models with Noisy Feedback

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2403.00409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.00409 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:03:47.464904Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:46.406911Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2a42b6b5-f3ad-4eb1-96b3-e897de916829 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.252104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:10d81148c88e6d18fd69a954021d31f361538aa466bfce6e50d2293335042884

Observation 0c4164c7-dd8d-4467-a398-af8bbeda1ebf · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.679899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:5a82094ea555451b01e5a372fe1e0c0ae42fe01db486da82f65faf25960a27b9

Observation 6e9ef6e6-ad4d-4ddd-ac35-445bbf735d62 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.795434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.795434Z digest=sha256:308bddf30d7e4eaa59414690ff03748898fccfd780e224f5f2fc41998336f821

Observation 6f6217b9-2b21-4366-b350-b0d52c2f1c40 · inbound

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO cites this paper.

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:08.819245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:08.819245Z digest=sha256:ab0e09409b40c3ee02aa2ab6b88a614f45921773ed384f39482a8a0e962c1f8f

Observation 55493028-8900-4233-ab07-3174d5ea4748 · inbound

Incentivizing High-Quality Human Annotations with Golden Questions cites this paper.

Incentivizing High-Quality Human Annotations with Golden Questions Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:42:19.320740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T13:41:26.730528Z digest=sha256:13141f57b8107c82129f2c5d8c8fb1924929286c0ffc4f9d4e05d01a43a1c583

Observation c07e7b14-0763-4b65-88e6-5d97b0cb27d7 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.673138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.673138Z digest=sha256:1730db6e1eebc87514dcfc52ae1c5d3086a354c2323429da79c32f0c7a637c8a

Observation 18383f47-ebb8-49be-82e0-db1deea0be98 · inbound

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences cites this paper.

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:53.423556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:53.423556Z digest=sha256:1d545c70842837ab2a889151dcd9a6d44623f3e58bd2136ef2f8edc2880a95c0

Observation e72950c1-3456-4bd4-8789-0c0fb597fb4a · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:31.844690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:31.844690Z digest=sha256:cd61e8d2c5d525a6169801ed428c1dadcf221460c38ca864b6548be9aa8f54cb

Observation 879ddf9d-ee39-4e6d-bd23-4b8d4c93c977 · inbound

Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates cites this paper.

Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:01:34.180564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T12:58:26.626172Z digest=sha256:b158cde2e0b4b40b791a66c97f61091cabf32c0f2798036bc29bfe3d792147d5

Observation a0f62a07-09a5-43fe-8a7f-b05f5f17f77f · inbound

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization cites this paper.

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:27:22.570045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:27:22.570045Z digest=sha256:0354b78ef2ce2730518177c1f841cf104ca892ddb7ae143272ed800a0ecafff8

Observation b1db7cef-8d31-405d-91c7-233f62db05e9 · inbound

Users as Annotators: LLM Preference Learning from Comparison Mode cites this paper.

Users as Annotators: LLM Preference Learning from Comparison Mode Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:06.992814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T08:19:58.093621Z digest=sha256:999d6c218c44a8b89458d5f509d8573441f2f36ce84dfa82a470a3e163c90d8b

Observation 1f275270-dd03-4b92-9031-d2d2159fd45e · inbound

Multilingual Safety Alignment via Self-Distillation cites this paper.

Multilingual Safety Alignment via Self-Distillation Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.410770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T19:35:56.059362Z digest=sha256:fbde637369e1041fdd3460dd2bf50e1abfeea6220a535e9f3aa2e39a8e63f47f

Observation 10a79da6-4279-4b68-9ab8-028dec443f34 · inbound

Multilingual Safety Alignment via Self-Distillation cites this paper.

Multilingual Safety Alignment via Self-Distillation Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:41:00.754016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:08:32.264867Z digest=sha256:97bc301fdfa1ff66e869bf2d0bfe8e58e4c56e2fb7bd4667140c73b3e268e08f

Observation 59f00d15-ee8a-4cfe-8867-7658461bc1a0 · inbound

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training cites this paper.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:32:24.273715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T06:30:51.812541Z digest=sha256:fbbc422382dc580551ad0b64513a97a09168e11551bbf2899f42fbcefe008d42

Observation a29d9f3b-da26-431f-b3c0-00d7387e32e4 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.586731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:89f6f45fc2d789da80b44a2646ce3486d53a166ff65f1b917448da5788539e1e

Observation 66a041cc-61ca-4ecb-90bd-fc72154bdac8 · inbound

Which Pairs to Compare for LLM Post-Training? cites this paper.

Which Pairs to Compare for LLM Post-Training? Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:09:19.254368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T20:37:19.221848Z digest=sha256:88ed5437641900bfd33d32ba4812ccc14557f9031a74db2c75092fec5dfe8efd

Observation 0d736ac7-41f0-4a83-9703-22e57c8659ea · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.408648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:21e35a920ceba01f1efee32ca49ae8b9f4c1d181009ad5c3ae4069ff0040b4fa

Observation 23266883-97f4-4d2f-8f72-348a70fd816d · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 190

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.410065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.410065Z digest=sha256:4526a5f90c1eef95fe482cbc40b1683a4ffc7e8748a4e683b0d0b29ea8654684

Observation f903dc6c-5203-4d51-8355-c19cb17ee593 · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T15:37:01.391649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:37:01.391649Z digest=sha256:d507a05f911a37afb0e5daa2d4aa474e326e50dd831b182783ec98ae64b6b131

Observation 1b17ee32-4086-4ec1-9545-b8151cf549fd · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:00:53.937134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:00:53.937134Z digest=sha256:e3fa8ed25c5b166e08b8c1ef351639bea55646e55679195f3331da6bc122b798

Observation bff2af43-1386-4c93-8032-bc0d540df63a · inbound

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation cites this paper.

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:03:47.464904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:03:47.464904Z digest=sha256:41c6ca15078852334b7be9788f584cffc0989f71e74db7528f3f97a4c9f495f0