Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2403.00409.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T15:03:47.464904Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T11:09:46.406911Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 2a42b6b5-f3ad-4eb1-96b3-e897de916829 · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c4164c7-dd8d-4467-a398-af8bbeda1ebf · inbound
How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6e9ef6e6-ad4d-4ddd-ac35-445bbf735d62 · inbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6217b9-2b21-4366-b350-b0d52c2f1c40 · inbound
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55493028-8900-4233-ab07-3174d5ea4748 · inbound
Incentivizing High-Quality Human Annotations with Golden Questions Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c07e7b14-0763-4b65-88e6-5d97b0cb27d7 · inbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18383f47-ebb8-49be-82e0-db1deea0be98 · inbound
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e72950c1-3456-4bd4-8789-0c0fb597fb4a · inbound
A Technical Survey of Reinforcement Learning Techniques for Large Language Models Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879ddf9d-ee39-4e6d-bd23-4b8d4c93c977 · inbound
Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a0f62a07-09a5-43fe-8a7f-b05f5f17f77f · inbound
Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1db7cef-8d31-405d-91c7-233f62db05e9 · inbound
Users as Annotators: LLM Preference Learning from Comparison Mode Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1f275270-dd03-4b92-9031-d2d2159fd45e · inbound
Multilingual Safety Alignment via Self-Distillation Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 10a79da6-4279-4b68-9ab8-028dec443f34 · inbound
Multilingual Safety Alignment via Self-Distillation Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 59f00d15-ee8a-4cfe-8867-7658461bc1a0 · inbound
Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a29d9f3b-da26-431f-b3c0-00d7387e32e4 · inbound
TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66a041cc-61ca-4ecb-90bd-fc72154bdac8 · inbound
Which Pairs to Compare for LLM Post-Training? Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0d736ac7-41f0-4a83-9703-22e57c8659ea · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 190
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 23266883-97f4-4d2f-8f72-348a70fd816d · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 190
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f903dc6c-5203-4d51-8355-c19cb17ee593 · inbound
Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b17ee32-4086-4ec1-9545-b8151cf549fd · inbound
Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bff2af43-1386-4c93-8032-bc0d540df63a · inbound
Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.