Pith. sign in

Paper Citation Record · LEDGER

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm

As of 16 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2505.01706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01706 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:18:40.431498Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation da9d4998-f419-4811-846d-12cbe65310da · outbound

This paper cites Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.907673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.907673Z digest=sha256:7c5eee840d3ac088e0426c98146861e3e221d9e9979f15a3544c516c6eff2e45

Observation 17a81456-0a56-4f8f-8587-92106ba508e9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.947889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.947889Z digest=sha256:0884e8f9f22e583977c9ceba3ae9a0d683e36cd6f0a385de371edc439eeeca67

Observation 5b48d90a-eaa2-442c-beb9-bcc31db1c918 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Rank analysis of incomplete block designs: I

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.952368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.952368Z digest=sha256:1fc54823b2c8d457f1cdf4138f86109e484cbcdcf937a4689b5d4c2bd5e5aaf8

Observation 213a1309-5cca-4770-b8a1-14dfdeba5ee6 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.956959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.956959Z digest=sha256:b6c3fefd811186cba83ed29f4799faa7e7fb0d6c41097d19ff9e6d9c6dcd1d07

Observation b9675d9a-123f-4c1c-b829-6b7d8892609d · outbound

This paper cites Provably Robust DPO: Aligning Language Models with Noisy Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.034871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.034871Z digest=sha256:5dc1d778a934a8a0fc9bef49cdc8c764f20a74dabd43c04013550c3bec7cf0c8

Observation 4b214a57-c63a-44cc-a86f-fc50abdf03cc · outbound

This paper cites Deep reinforcement learning from human preferences.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.092021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.092021Z digest=sha256:a3157f8206bfc138dfb5fe290eee1893ee8542628a64352093aa5c0679885ecb

Observation 99edc449-12fa-4e1d-9084-e2a3c0e2577d · outbound

This paper cites Diverse Preference Optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Diverse Preference Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.097323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.097323Z digest=sha256:d0097da4d470e2a56a91c95193efaaab9905295f311d05523ca60af0d7889f62

Observation 46ec5300-70a3-4c60-a8f9-a7f7f07df7b7 · outbound

This paper cites 2D-DPO: Scaling direct preference optimization with 2-dimensional supervision.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm 2D-DPO: Scaling direct preference optimization with 2-dimensional supervision

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:18:40.795844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:18:40.132613Z digest=sha256:2dc7d69ec8f33e6ecaf701c20dce50352f3a7d073f4cb9f6525f6d164a6cd960

Observation b99ca145-d560-4b7b-9ab4-46d5e33770e3 · outbound

This paper cites A Survey of Direct Preference Optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm A Survey of Direct Preference Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.171743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.171743Z digest=sha256:457c65c265cfd073b7b06bf594a8579410ad2f90639d5fef44aec9b1f08ac187

Observation 5d16bf50-25f6-4d67-9b7f-da58c07bff7c · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.205387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.205387Z digest=sha256:8542ca01064ea7478d3aefba253f8c44fe9ad37ae5f869c6d2030e2ee53a490a

Observation 1341ab5e-4e74-41a3-b22a-a327ab782103 · outbound

This paper cites RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.210266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.210266Z digest=sha256:2668fe4607a9f3cb0be5765832004a59c39a2fd2e2a8fbf260b3dcd4324c574e

Observation 3fd6d03a-72fb-401b-ac24-a6ebaf2a3947 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Direct preference optimization: Your language model is secretly a reward model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.215067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.215067Z digest=sha256:699a2f6dba7c5ed1f4d8a68d823a45f404d954ca8167e880b6a01492782a53df

Observation 57710674-310b-406e-a080-2ada650490c2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.224347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.224347Z digest=sha256:f3f7233839b191ca43a5d377ae2efed50c880b241239a3ebd5e536fbb733e33f

Observation 9953d08a-c9b7-499c-bc11-e479704e9215 · outbound

This paper cites Large Language Model Alignment: A Survey.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Large Language Model Alignment: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.269907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.269907Z digest=sha256:e962c5f9625e4678c932a03a101bd258302ba30878c6e1d71b3674609eeb737f

Observation 857476ab-f5b4-4fba-aeb0-c03eae9761fe · outbound

This paper cites Things we like: human preferences among similar organisms and implications for conservation.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Things we like: human preferences among similar organisms and implications for conservation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:18:40.775002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:18:40.307558Z digest=sha256:b77fa811b794a11a3c76f71834b6345c13be3f3dd4a6c81910164111f327ee08

Observation 7f60104e-3e48-426c-8b1a-31cc8ce16fdd · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Aligning Large Language Models with Human: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.312439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.312439Z digest=sha256:205c6eb01b6503af61cb8a7ce227c59d74cf8e74751fde5a521bb6a4657fc61c

Observation 793f0b0b-2aae-441e-a55c-c64bb7e2653a · outbound

This paper cites Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.317605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.317605Z digest=sha256:d241fff15649e6ee524b66bd73c2b465264b36b06971b35346d3d0e3e49382d0

Observation 5cd4cba5-630c-4b7c-9577-6982d9e9e87a · outbound

This paper cites Token-level Direct Preference Optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Token-level Direct Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.424127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.424127Z digest=sha256:5f1083883e9d5e860f89a2d89dc005b9e0f455d367c4e56c9ed90c4aa73e9fed

Observation af46cbf7-a3e4-4e51-8ea6-658b6ce5e590 · outbound

This paper cites Beyond one-preference-for-all: Multi-objective direct preference optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Beyond one-preference-for-all: Multi-objective direct preference optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:18:40.678798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:18:40.427730Z digest=sha256:0ecf205e023e69742030279eb9055c91dae630a11ab51a80809912b1f39a7ee5

Observation 05846fd4-7eb7-416d-b9ab-dea33d3e4f0f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Fine-Tuning Language Models from Human Preferences

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.431498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.431498Z digest=sha256:ecc751a4e77c4990b9e2133d0a2a0091add87620d463ac637f504bada88d3643

Pith citing papers

No inbound Pith citation observations are available.