Pith. sign in

Paper Citation Record · LEDGER

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm

As of 20 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2505.01706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01706 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:18:40.431498Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation da9d4998-f419-4811-846d-12cbe65310da · outbound

This paper cites Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.907673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.907673Z digest=sha256:3225f9c1e3626c5e65bb783359e23a555fa220260b73f742bb275aedef911f55

Observation 17a81456-0a56-4f8f-8587-92106ba508e9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.947889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.947889Z digest=sha256:e923e796b1ff5e430d90fdff9e04f53587685f704b84ae586a8764cb8b9a1190

Observation 5b48d90a-eaa2-442c-beb9-bcc31db1c918 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Rank analysis of incomplete block designs: I

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.952368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.952368Z digest=sha256:6f1c268752670c6a9c759aa525985db318755cbd51b1c9de60908926c5fb69ed

Observation 213a1309-5cca-4770-b8a1-14dfdeba5ee6 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.956959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.956959Z digest=sha256:326ff6e08e86dcc6123ff1daa51ded68d2a6ba7b62df9cfcfa63acfe28097897

Observation b9675d9a-123f-4c1c-b829-6b7d8892609d · outbound

This paper cites Provably Robust DPO: Aligning Language Models with Noisy Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.034871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.034871Z digest=sha256:840d2667f6bfe63409c6e5acf993642dd6874994f63c45a61dcaf3f93374d200

Observation 4b214a57-c63a-44cc-a86f-fc50abdf03cc · outbound

This paper cites Deep reinforcement learning from human preferences.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.092021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.092021Z digest=sha256:d3702c5fbcc3ba953542a5371f2147397bf77f9f41a70e726339fb2c27e23642

Observation 99edc449-12fa-4e1d-9084-e2a3c0e2577d · outbound

This paper cites Diverse Preference Optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Diverse Preference Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.097323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.097323Z digest=sha256:8fdbe9bc9d31bd4c96286fc37ccd19b218a176e3c0ccc416d355fded6eea1afd

Observation 46ec5300-70a3-4c60-a8f9-a7f7f07df7b7 · outbound

This paper cites 2D-DPO: Scaling direct preference optimization with 2-dimensional supervision.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm 2D-DPO: Scaling direct preference optimization with 2-dimensional supervision

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:18:40.795844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:18:40.132613Z digest=sha256:e63b19eb559f295a17e76d132ba96a4a3ab2d6a2fa6730495bd9a3366f30f757

Observation b99ca145-d560-4b7b-9ab4-46d5e33770e3 · outbound

This paper cites A Survey of Direct Preference Optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm A Survey of Direct Preference Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.171743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.171743Z digest=sha256:2d9691b59642c1907275063fa57fe336ed10e39823794861f4316bcecb3533a7

Observation 5d16bf50-25f6-4d67-9b7f-da58c07bff7c · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.205387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.205387Z digest=sha256:29559b6361896de3bdcd368e155899caf4221df7dbd6b69a13f70779b9c52738

Observation 1341ab5e-4e74-41a3-b22a-a327ab782103 · outbound

This paper cites RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.210266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.210266Z digest=sha256:ec4fe09585fab6f48f259c45934a259e40892514bb0908d17ac1d45a1d00e9d0

Observation 3fd6d03a-72fb-401b-ac24-a6ebaf2a3947 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Direct preference optimization: Your language model is secretly a reward model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.215067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.215067Z digest=sha256:955bf456f885f9a7e8cb2250d8b96dcc7bc0841f71686a598650f8bb744c0ae2

Observation 57710674-310b-406e-a080-2ada650490c2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.224347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.224347Z digest=sha256:e233da37d22710a192a0d8c0dbb24601edf306dc74835ba100d18bab8e4aa979

Observation 9953d08a-c9b7-499c-bc11-e479704e9215 · outbound

This paper cites Large Language Model Alignment: A Survey.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Large Language Model Alignment: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.269907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.269907Z digest=sha256:4ff188cde3f83b32db76ded275da1c7ad888effc657b4d3e4d32c54400fd2dcb

Observation 857476ab-f5b4-4fba-aeb0-c03eae9761fe · outbound

This paper cites Things we like: human preferences among similar organisms and implications for conservation.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Things we like: human preferences among similar organisms and implications for conservation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:18:40.775002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:18:40.307558Z digest=sha256:48a50aa81c4eaa5ab58f1143eb1cc37223ac9f5b8655621b7ad9587094e8cef9

Observation 7f60104e-3e48-426c-8b1a-31cc8ce16fdd · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Aligning Large Language Models with Human: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.312439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.312439Z digest=sha256:5f504f89f1550d5b9786769da30792e8f124b91e88e90941e152e3f6ce87ae68

Observation 793f0b0b-2aae-441e-a55c-c64bb7e2653a · outbound

This paper cites Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.317605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.317605Z digest=sha256:ac52e045b046efaa6275698a32056ba4a692ddc5ea1485b4a7274e32d4bf9536

Observation 5cd4cba5-630c-4b7c-9577-6982d9e9e87a · outbound

This paper cites Token-level Direct Preference Optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Token-level Direct Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.424127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.424127Z digest=sha256:2dff91565024c477552c674ee5742735f101227e9bd2b63673e7dd3a445e2105

Observation af46cbf7-a3e4-4e51-8ea6-658b6ce5e590 · outbound

This paper cites Beyond one-preference-for-all: Multi-objective direct preference optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Beyond one-preference-for-all: Multi-objective direct preference optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:18:40.678798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T04:18:40.427730Z digest=sha256:914aaad11d96cde31f76c84e486392e573cee72b40ac345864df3eeff54ab7ac

Observation 05846fd4-7eb7-416d-b9ab-dea33d3e4f0f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Fine-Tuning Language Models from Human Preferences

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.431498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.431498Z digest=sha256:199afcd01eafca09a7373db309c01e1dc3bf51ab60fd0bf5dbdd228ebb650a0d

Pith citing papers

No inbound Pith citation observations are available.