Pith. sign in

Paper Citation Record · LEDGER

Disentangling Length from Quality in Direct Preference Optimization

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2403.19159.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.19159 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:57:29.682556Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:07.153004Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cd80fb21-6efd-4b9a-b2a4-3b7ed043da32 · inbound

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization cites this paper.

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:05:46.916143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:03:53.675107Z digest=sha256:69e08491a7181daf8e0adba621483454908cb7493296ec00bf18a02a5192c2d8

Observation fbf0c15e-e879-4f34-b07e-e854e4c04e5f · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:29.682556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:29.682556Z digest=sha256:fe26f12bf0d66ef48fdbc7a0980035681b1ec08d91c7e247ed882b89bdddcd87

Observation a76c2dc1-1aba-4b8a-88f1-67ae4cc22924 · inbound

MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Framework cites this paper.

MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Framework Disentangling Length from Quality in Direct Preference Optimization

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:01.345378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:01.345378Z digest=sha256:03fa485013651a0a9804dbbccde8048961fc193e445098eec4dfd970c53a3f50

Observation ef4479f0-cb24-4db8-a1dd-270a815a3ae6 · inbound

Aligning Large Language Models with Implicit Preferences from User-Generated Content cites this paper.

Aligning Large Language Models with Implicit Preferences from User-Generated Content Disentangling Length from Quality in Direct Preference Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.110646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.110646Z digest=sha256:05a1acb79c215b3cee147a0e70ad3c9c591cfe0bbb26e9a0039b7710db7abb27

Observation 7faebdb5-bce8-4568-9819-96e8c32d9373 · inbound

Unlocking Recursive Thinking of LLMs: Alignment via Refinement cites this paper.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Disentangling Length from Quality in Direct Preference Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.051648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.051648Z digest=sha256:1b469241b814007f7274b3358ff6b06dfdd04b77955cb5b4a23343594f579753

Observation 2444c89a-c34f-44c2-a429-af7a17a9eb79 · inbound

Explicit Preference Optimization: No Need for an Implicit Reward Model cites this paper.

Explicit Preference Optimization: No Need for an Implicit Reward Model Disentangling Length from Quality in Direct Preference Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.966403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.966403Z digest=sha256:eb7f02241ed724da17f7d50f27acce93c849a6823d7970b4cb0d0504125225d6

Observation 51e5e0f4-f6c6-48d5-9ce1-3c8eade8b949 · inbound

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization cites this paper.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.113845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.113845Z digest=sha256:8725c5fef58c34c52270f7a4086b174147ff07033da108d174ff7f19496d8b1b

Observation 561a7e33-73f4-4e69-9cfa-a42e91b4fffd · inbound

Bridging Offline and Online Reinforcement Learning for LLMs cites this paper.

Bridging Offline and Online Reinforcement Learning for LLMs Disentangling Length from Quality in Direct Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.964261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.964261Z digest=sha256:68bf63917aa975b211ea275ded3601390286e2405faf9232bef8dff62288e55a

Observation fc58b32c-1150-498e-97fb-8a1730d2472d · inbound

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap cites this paper.

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap Disentangling Length from Quality in Direct Preference Optimization

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.708523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T23:46:24.208438Z digest=sha256:a794b13df535b7fc5917d57494cc70d4461937c5f11508e4c3ee18f14966a058

Observation 3cc6d8ee-4de4-4718-ae7a-cdf3005b8489 · inbound

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints cites this paper.

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints Disentangling Length from Quality in Direct Preference Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:33:14.576084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:33:14.576084Z digest=sha256:9a9cc19bbffd435aee38ebb62f5fd1596566c849b6a6903ec16b7de606db781f

Observation 4a85e76b-341b-4052-832a-3d4680e1b659 · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.001156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:65c551fb5d8496cafc1909640637b41432c5ad3a7c8d892eff4fa4504d3347c5

Observation f71c7d4f-3a4f-42b5-899e-613e0bf4b900 · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF Disentangling Length from Quality in Direct Preference Optimization

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.569063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:1ed2ad66602751e75187a82119d618bd55a52ce5f77441efbadb814c56b5b953

Observation 89d6fd95-903d-42a4-a73f-dbddd0e393f0 · inbound

AlignCultura: Towards Culturally Aligned Large Language Models? cites this paper.

AlignCultura: Towards Culturally Aligned Large Language Models? Disentangling Length from Quality in Direct Preference Optimization

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:05.154961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T02:36:36.854805Z digest=sha256:6c44ee88bf66d0f4837164c06d67b09102c235d437188a56239540934f481ff5

Observation b2a2d5ec-c627-4f8c-985f-2ce357cd246b · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Disentangling Length from Quality in Direct Preference Optimization

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.700043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:bc42448ec1da4620bca16ba747b6e5004a65cad50e25ec31f9adb9fd8ccf0dc2

Observation 4ae7dc73-eaa8-434b-82b5-c44ced77521b · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:56:05.480733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T16:57:49.396570Z digest=sha256:1ce828af1828f026dd5c2c350a21463742e76987fa76b7d9c2a8b33c50bb7a56

Observation dcb49329-1564-43f4-98c1-f57db2b90c7f · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:21:26.487774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:26:54.426050Z digest=sha256:895ce0984da946545fa3eb67f473dca49f4e3404eb06b60f0d6b6550037cda30

Observation 2910cfc1-58e0-4fb9-9087-1b443aeb8bef · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:12:28.831414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:08:39.328446Z digest=sha256:bc73e9f3bae4015e5e45e4f70089ca6720429b0739172dfa48285695f890f71f

Observation 12a70180-3549-4dcf-8df7-01cb3c4ee7cb · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T14:53:24.400757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:53:24.400757Z digest=sha256:2e6c8d0126c84d6cd2ef49b1d4f958fe7361f5ddee03a01bf94143518a8bf595

Observation 7777d650-97ff-4cb3-bdd2-66774e0b157d · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Disentangling Length from Quality in Direct Preference Optimization

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:59.553809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:b31ec330ddf63692f5fbf57a255cbfd71152e082d3b94e60e9a31737272c07e7

Observation abd32b47-9cb9-46ea-8461-5683afdb32b6 · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Disentangling Length from Quality in Direct Preference Optimization

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:38.226241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:3f5a6c70dba6c4de4d778a3412762195f189c64aa54fdd92230fe703d6212d8e

Observation 4915950a-6e37-4019-b487-ba51da449163 · inbound

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training cites this paper.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training Disentangling Length from Quality in Direct Preference Optimization

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:32:24.288800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T06:30:51.812541Z digest=sha256:82b10f3c6538c8b7bd17c490589b219c4ae22a964273af59ebbbc24f676f3889

Observation 228e011e-8d83-4c58-a34d-0abd2e02fe4e · inbound

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates cites this paper.

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates Disentangling Length from Quality in Direct Preference Optimization

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:33:24.381122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:29:55.729913Z digest=sha256:713640810a901e25192cdfa1cbc375ec3846c440acbc5e8eaeec46c3a2dd8ed4

Observation e6b289a2-75ca-4e9f-b630-5d3ef85d82b5 · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation Disentangling Length from Quality in Direct Preference Optimization

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.154736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T21:05:36.836361Z digest=sha256:83c436623311407168aecb4ad2e27e8dad0f2ea9e34750a1cb67b06e98ee0893

Observation 8d9a76b1-c24c-4b06-8528-4ab91745e03a · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation Disentangling Length from Quality in Direct Preference Optimization

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.562457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T06:30:27.178950Z digest=sha256:b587a46ecde193464d099ca5930da80334cf7d2515df8aee2a37da81511e2b70

Observation 67241da9-311f-417b-a35d-b930d892e7e1 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Disentangling Length from Quality in Direct Preference Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:756fafda948fc212ba1773a98691a727e21f5f9978f181f4dc92e9aa377241ae

Observation e9477601-761a-464d-ba86-7f81ed5390dd · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Disentangling Length from Quality in Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:35.672115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:35.672115Z digest=sha256:552a7194c0166b5199c304acb48b466fe7654c3620abceaf12e999b71647bc01

Observation 0f75957c-fcb7-430b-ad0a-edbae361afb7 · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Disentangling Length from Quality in Direct Preference Optimization

Reference 183

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:36.689869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:36.689869Z digest=sha256:2e4914d74ce0bdee3d87f280755ef557b5084c1e8d4adcc1ead536865ac0d3c4