Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:18:40.431498Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2505.01706.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:18:40.431498Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation da9d4998-f419-4811-846d-12cbe65310da · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17a81456-0a56-4f8f-8587-92106ba508e9 · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b48d90a-eaa2-442c-beb9-bcc31db1c918 · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Rank analysis of incomplete block designs: I
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 213a1309-5cca-4770-b8a1-14dfdeba5ee6 · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9675d9a-123f-4c1c-b829-6b7d8892609d · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b214a57-c63a-44cc-a86f-fc50abdf03cc · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Deep reinforcement learning from human preferences
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99edc449-12fa-4e1d-9084-e2a3c0e2577d · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Diverse Preference Optimization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46ec5300-70a3-4c60-a8f9-a7f7f07df7b7 · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm 2D-DPO: Scaling direct preference optimization with 2-dimensional supervision
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b99ca145-d560-4b7b-9ab4-46d5e33770e3 · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm A Survey of Direct Preference Optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d16bf50-25f6-4d67-9b7f-da58c07bff7c · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1341ab5e-4e74-41a3-b22a-a327ab782103 · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd6d03a-72fb-401b-ac24-a6ebaf2a3947 · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Direct preference optimization: Your language model is secretly a reward model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57710674-310b-406e-a080-2ada650490c2 · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9953d08a-c9b7-499c-bc11-e479704e9215 · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Large Language Model Alignment: A Survey
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 857476ab-f5b4-4fba-aeb0-c03eae9761fe · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Things we like: human preferences among similar organisms and implications for conservation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7f60104e-3e48-426c-8b1a-31cc8ce16fdd · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Aligning Large Language Models with Human: A Survey
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793f0b0b-2aae-441e-a55c-c64bb7e2653a · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cd4cba5-630c-4b7c-9577-6982d9e9e87a · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Token-level Direct Preference Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af46cbf7-a3e4-4e51-8ea6-658b6ce5e590 · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Beyond one-preference-for-all: Multi-objective direct preference optimization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 05846fd4-7eb7-416d-b9ab-dea33d3e4f0f · outbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Fine-Tuning Language Models from Human Preferences
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.