Pith. sign in

Paper Citation Record · LEDGER

Disentangling Length from Quality in Direct Preference Optimization

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2403.19159.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.19159 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:25.220470Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:07.153004Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cd80fb21-6efd-4b9a-b2a4-3b7ed043da32 · inbound

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization cites this paper.

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:05:46.916143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-23T19:03:53.675107Z digest=sha256:8526c27e2f9268cc06f2c9825bebb98239896e7c862975c19d9141fdec9584f0

Observation 79b340e3-5408-475c-92a1-d6836787694b · inbound

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution cites this paper.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Disentangling Length from Quality in Direct Preference Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.797495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.797495Z digest=sha256:0e6f24159ed09ad5704c1b968e91a976370efdde3d92571588d740d45d87fc41

Observation 3abaff11-2063-48f9-bbb6-bd0beb6dd9d0 · inbound

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework cites this paper.

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework Disentangling Length from Quality in Direct Preference Optimization

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-12T18:15:15.895717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:15:15.895717Z digest=sha256:d22c0a46fd5d27ad86f0a4a37ff7733f070aa332ca0e2088f8bb20949c1a135d

Observation 3e6431bc-985d-43d4-955b-60e0b6ec7616 · inbound

Weighted-Reward Preference Optimization for Implicit Model Fusion cites this paper.

Weighted-Reward Preference Optimization for Implicit Model Fusion Disentangling Length from Quality in Direct Preference Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.144929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.144929Z digest=sha256:5f1d4ec0e837d4db6cb6b1b4236fe3a5dda2e2b192154b360d0e8a57093aa2d7

Observation 8f291d2b-6762-4ace-80c8-727e3371753d · inbound

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model cites this paper.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Disentangling Length from Quality in Direct Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.989023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.989023Z digest=sha256:afeb5759cdd113aced1cefad3f57374c7beba35fe91e62d906ef8e1a5c2318f0

Observation dac39ffc-405e-4b38-ab02-71e26a67d191 · inbound

Hansel: Output Length Controlling Framework for Large Language Models cites this paper.

Hansel: Output Length Controlling Framework for Large Language Models Disentangling Length from Quality in Direct Preference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:37:34.290665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:37:34.290665Z digest=sha256:d87b2a4ef864492773fa20962c8deff639b3f04f5e416782cfc933e29f35fd04

Observation ae6992c1-4b14-49d4-b764-c702636a6938 · inbound

AlphaPO: Reward Shape Matters for LLM Alignment cites this paper.

AlphaPO: Reward Shape Matters for LLM Alignment Disentangling Length from Quality in Direct Preference Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:51:08.514751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:51:08.514751Z digest=sha256:4e41def4b50fb57b507ce01a7d5e16946624256e525e85dfbd0dae8a83b253fe

Observation 376dbddb-1af1-4a53-9e9a-bc153e20df36 · inbound

Controllable Protein Sequence Generation with LLM Preference Optimization cites this paper.

Controllable Protein Sequence Generation with LLM Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:51.832360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:51.832360Z digest=sha256:b4723a7af27480347bd00adf0fd1e5576cca06f7a7428f7d4218d6a101cfd75f

Observation 5919e0e6-9054-4882-822b-42af4f2dac01 · inbound

SimulPL: Aligning Human Preferences in Simultaneous Machine Translation cites this paper.

SimulPL: Aligning Human Preferences in Simultaneous Machine Translation Disentangling Length from Quality in Direct Preference Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T18:21:56.118350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:21:56.118350Z digest=sha256:635a83d5ec217f1d283625235948e83b5762250bec9e96e80d83158cd6d0075d

Observation 5bb56b2a-80fe-4d09-bf07-d307e47eea2f · inbound

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling cites this paper.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling Disentangling Length from Quality in Direct Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.235461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.235461Z digest=sha256:032e2ea56fd76913fe32fb9166ab4fc099e3651fa4371dc9dc742f7ebd972e9d

Observation ec8409b5-0173-49ee-9f53-106c5931d72c · inbound

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective cites this paper.

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective Disentangling Length from Quality in Direct Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T04:08:51.815624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:08:51.815624Z digest=sha256:30b43c274cbe3d68da81e180c87545601dd275fb9cf0ebfa17b7055c5980a59e

Observation c94c4030-9263-4763-91fc-8508caf05beb · inbound

Design Considerations in Offline Preference-based RL cites this paper.

Design Considerations in Offline Preference-based RL Disentangling Length from Quality in Direct Preference Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.029675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.029675Z digest=sha256:6291bc15e04e9876ccbdf63373a7a2935d291ea22ad9e36e6b6ed4a4f9c8f437

Observation 5cc0fbaa-b4f8-4303-a375-4b58990c6ebf · inbound

DPO-Shift: Shifting the Distribution of Direct Preference Optimization cites this paper.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.959380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.959380Z digest=sha256:97b9a978db29356a10b19a8c8e9076168c6fdb06e11f48d027fa84a28c54e460

Observation aac4bc6e-f36f-4a2a-9635-8d77b4071ae2 · inbound

ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition cites this paper.

ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition Disentangling Length from Quality in Direct Preference Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:25.220470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:33:25.220470Z digest=sha256:66356d36501152b42db6ae701be285117cd126f1291b796a20e93c5ccabb317e

Observation c8e471bc-555a-41ba-81f2-98b146f7e896 · inbound

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models cites this paper.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Disentangling Length from Quality in Direct Preference Optimization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.448976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.448976Z digest=sha256:1d8853e653edfa9dae80d17ba6de9905e502ef0b8ac031c1119c8ab0f538799f

Observation 5457311b-7bd6-4dcc-a634-03e7ef430835 · inbound

Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation cites this paper.

Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation Disentangling Length from Quality in Direct Preference Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:02.287393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:02.287393Z digest=sha256:7eafcc2d8ab5eca8d189e8db1019f85dc3d671ddee51d8f4e765e1403faf8f2d

Observation 54a87a3e-fdf6-45bb-b0be-d7f441ad7a7e · inbound

A Survey on Progress in LLM Alignment from the Perspective of Reward Design cites this paper.

A Survey on Progress in LLM Alignment from the Perspective of Reward Design Disentangling Length from Quality in Direct Preference Optimization

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:07.033656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:52:07.033656Z digest=sha256:fd7bdb07023d542bae7905efdc23d5b80e3eec1249aedf95b7c3cb4f82487410

Observation b9af7862-07bc-4ea4-9094-1a5eba98fd14 · inbound

InfoPO: On Mutual Information Maximization for Large Language Model Alignment cites this paper.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Disentangling Length from Quality in Direct Preference Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.055880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.055880Z digest=sha256:40b124431eaa6545bd630a58f5d8ed4b082eb1be0bb36fa269c31bd8c11d49ce

Observation e3515bde-d35c-4f4d-a002-2d564b83b3fd · inbound

Preference Optimization for Combinatorial Optimization Problems cites this paper.

Preference Optimization for Combinatorial Optimization Problems Disentangling Length from Quality in Direct Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:55:41.022972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:55:41.022972Z digest=sha256:4753034855220e609c9f345ad92cb40e8d1000a9e7ea8ca482f7c5b4b5ca1629

Observation fbf0c15e-e879-4f34-b07e-e854e4c04e5f · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:29.682556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:29.682556Z digest=sha256:fb3cc57ed192f9461768d4615d81c5cf918c89443ee98f97513795cfe870c36e

Observation a76c2dc1-1aba-4b8a-88f1-67ae4cc22924 · inbound

MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Framework cites this paper.

MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Framework Disentangling Length from Quality in Direct Preference Optimization

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:01.345378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:01.345378Z digest=sha256:a3c9fdb9a6d936227e9a1e062dbf88b6e551c24cc4a3007fad4d28ae8084d645

Observation ef4479f0-cb24-4db8-a1dd-270a815a3ae6 · inbound

Aligning Large Language Models with Implicit Preferences from User-Generated Content cites this paper.

Aligning Large Language Models with Implicit Preferences from User-Generated Content Disentangling Length from Quality in Direct Preference Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.110646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.110646Z digest=sha256:37369d074eea16aacb554e3a8149eeb021e87582c72b328450263c3c0d049973

Observation 7faebdb5-bce8-4568-9819-96e8c32d9373 · inbound

Unlocking Recursive Thinking of LLMs: Alignment via Refinement cites this paper.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Disentangling Length from Quality in Direct Preference Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.051648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.051648Z digest=sha256:67545d7029d0b84cc5c561c2389e23a5075e039a086d3df2b55eca96809138d5

Observation 2444c89a-c34f-44c2-a429-af7a17a9eb79 · inbound

Explicit Preference Optimization: No Need for an Implicit Reward Model cites this paper.

Explicit Preference Optimization: No Need for an Implicit Reward Model Disentangling Length from Quality in Direct Preference Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.966403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.966403Z digest=sha256:366a4d95f181556072385da9fe378cf7df2d19708626e9f43511bee0d12b49e5

Observation 51e5e0f4-f6c6-48d5-9ce1-3c8eade8b949 · inbound

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization cites this paper.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.113845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.113845Z digest=sha256:529138b4f61fd9ce9ab649029794daddf411e947f964e7fc3aee5ed31be0f0bc

Observation 561a7e33-73f4-4e69-9cfa-a42e91b4fffd · inbound

Bridging Offline and Online Reinforcement Learning for LLMs cites this paper.

Bridging Offline and Online Reinforcement Learning for LLMs Disentangling Length from Quality in Direct Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.964261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.964261Z digest=sha256:41c95703668875776dc6a939fdfb698a3ceb72b1a2516fee42db7313376b3b90

Observation 57b33c6a-93d3-4bf4-ba41-21bb18489a82 · inbound

DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition cites this paper.

DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition Disentangling Length from Quality in Direct Preference Optimization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:47.517073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:47.517073Z digest=sha256:03635c91251e9b8ce3d9a93d782a21859cb49244432d838996b60dade6c6b8de

Observation ade77db5-f931-4a63-8a69-74c51116f0bb · inbound

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering cites this paper.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Disentangling Length from Quality in Direct Preference Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.513629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.513629Z digest=sha256:aba568c254e3242fb1597cfc489f891d68e3b37dec68fcd136a547fecad99377

Observation fc58b32c-1150-498e-97fb-8a1730d2472d · inbound

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap cites this paper.

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap Disentangling Length from Quality in Direct Preference Optimization

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.708523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T23:46:24.208438Z digest=sha256:1d9f68bf15a6105249121e0e14cfaf1c8d73317f81b59b4a3cb724e81fd78ac2

Observation 3cc6d8ee-4de4-4718-ae7a-cdf3005b8489 · inbound

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints cites this paper.

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints Disentangling Length from Quality in Direct Preference Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:33:14.576084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:33:14.576084Z digest=sha256:c5aded6ef26fb439edf5411e35e8d090ee3ebc51bbcad63d8540cd555ddc8516

Observation 4a85e76b-341b-4052-832a-3d4680e1b659 · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.001156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:fb25a1fba4dfe7bedb4dbbdeedfbba34d8286564bf1d1091a68a5c353a179cef

Observation f71c7d4f-3a4f-42b5-899e-613e0bf4b900 · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF Disentangling Length from Quality in Direct Preference Optimization

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.569063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:2bcd07a566c50252a615b88e77c034f8526ccd242565efc91fd270ad916bf0c3

Observation 89d6fd95-903d-42a4-a73f-dbddd0e393f0 · inbound

AlignCultura: Towards Culturally Aligned Large Language Models? cites this paper.

AlignCultura: Towards Culturally Aligned Large Language Models? Disentangling Length from Quality in Direct Preference Optimization

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:05.154961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T02:36:36.854805Z digest=sha256:335bf4ceea524dc76c89f13b7088240f229bbd7406f509ce09eaac91ffd35984

Observation b2a2d5ec-c627-4f8c-985f-2ce357cd246b · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Disentangling Length from Quality in Direct Preference Optimization

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.700043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:f9fff1c746be47c278e15af0dc1db4421af595d3ef7e56e79f223f9130b6f1b5

Observation 4ae7dc73-eaa8-434b-82b5-c44ced77521b · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:56:05.480733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T16:57:49.396570Z digest=sha256:ae1ed1ccbac647fa2fd3ffedaeb7cd696adc6b03d64a44998b29eaad0b3d2979

Observation dcb49329-1564-43f4-98c1-f57db2b90c7f · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:21:26.487774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T03:26:54.426050Z digest=sha256:e6ccbf03fe98d50879f541ac4c45227fb7981b3f4689f7dc8c5eb21da77d8e4f

Observation 2910cfc1-58e0-4fb9-9087-1b443aeb8bef · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:12:28.831414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T07:08:39.328446Z digest=sha256:35d057f6bf0a1eb29dc07ddaafbf6419f8f9cebe231fa9da7af223e283de2abc

Observation 12a70180-3549-4dcf-8df7-01cb3c4ee7cb · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T14:53:24.400757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:53:24.400757Z digest=sha256:6aac69063581d7250302b7a211c7294ff8a3b555bb792cfacbdddfd752d198ef

Observation 7777d650-97ff-4cb3-bdd2-66774e0b157d · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Disentangling Length from Quality in Direct Preference Optimization

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:59.553809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:8bb2b188e062034a2c5cdd278601429b057ab07d8384d90ac09ba62a32bbf868

Observation abd32b47-9cb9-46ea-8461-5683afdb32b6 · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Disentangling Length from Quality in Direct Preference Optimization

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:38.226241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:2f506f6a9535bdfa2dce2ce4e7a25f8d41a0520a76b9cef16de15d2612354377

Observation 4915950a-6e37-4019-b487-ba51da449163 · inbound

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training cites this paper.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training Disentangling Length from Quality in Direct Preference Optimization

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:32:24.288800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T06:30:51.812541Z digest=sha256:722ecf082bad1c4d471506f5be0908c0d15e7afef580cc102895ab28c829badc

Observation 228e011e-8d83-4c58-a34d-0abd2e02fe4e · inbound

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates cites this paper.

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates Disentangling Length from Quality in Direct Preference Optimization

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:33:24.381122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T12:29:55.729913Z digest=sha256:2eda16ff0f181bf319564d2a5c8682d37c5a702221b0d513fe6c5947e9a82920

Observation e6b289a2-75ca-4e9f-b630-5d3ef85d82b5 · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation Disentangling Length from Quality in Direct Preference Optimization

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.154736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-25T21:05:36.836361Z digest=sha256:0775a640659e0bd69a9ee9c26b7ab9b2110977122028dab98c806fb886adbc60

Observation 8d9a76b1-c24c-4b06-8528-4ab91745e03a · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation Disentangling Length from Quality in Direct Preference Optimization

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.562457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-01T06:30:27.178950Z digest=sha256:f3dfd70646f3c2f084183ccf58a8977c74ddca25ee7c3cb9f29311d044d62047

Observation 67241da9-311f-417b-a35d-b930d892e7e1 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Disentangling Length from Quality in Direct Preference Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:48b0baddd66c90044f57023a174e8f617d1c8b7318afe43365c891d2948dec10

Observation e9477601-761a-464d-ba86-7f81ed5390dd · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Disentangling Length from Quality in Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:35.672115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:35.672115Z digest=sha256:93125a344321a8f80ad5ffd01455e0df0a7cf4d8d6cb495deac1ff8c68101049

Observation 0f75957c-fcb7-430b-ad0a-edbae361afb7 · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Disentangling Length from Quality in Direct Preference Optimization

Reference 183

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:36.689869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:36.689869Z digest=sha256:18c26e3c07f880bd924004d334b5c6dabd64fe26797e1e22ad6a92cf87724cd1

Observation 8113fb18-92fb-4643-9e23-7645c11ed9b2 · inbound

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration cites this paper.

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration Disentangling Length from Quality in Direct Preference Optimization

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-15T14:33:57.591644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:33:57.591644Z digest=sha256:0237967659bb7e6836f82e491da9076553a14c27761d7464fbbd60a8a5d1059c