Pith. sign in

Paper Citation Record · LEDGER

Length Desensitization in Direct Preference Optimization

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2409.06411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.06411 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:37:34.210531Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.691954Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3aa25225-c77a-477a-9353-1cb33690dc55 · inbound

Hansel: Output Length Controlling Framework for Large Language Models cites this paper.

Hansel: Output Length Controlling Framework for Large Language Models Length Desensitization in Direct Preference Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:37:34.210531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:37:34.210531Z digest=sha256:5d765fb95af5d53eb5d2cd9e455a741ae6d426637838bdfe50263c6d00428347

Observation 5f33d3a0-bab8-4df3-8ffb-e0ad92964bbe · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Length Desensitization in Direct Preference Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:31.732769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:31.732769Z digest=sha256:13212e74d6a899e25f6c7c49bff17d0cbd3f15484d02167addd51c2c6a8eca9e

Observation abf57df3-9381-41e3-8389-0b164cbcf2f0 · inbound

TiCo: Time-Controllable Spoken Dialogue Model cites this paper.

TiCo: Time-Controllable Spoken Dialogue Model Length Desensitization in Direct Preference Optimization

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.752476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T00:38:52.182973Z digest=sha256:c5b15bb7a364f82b5d1cbb249ab75b6a100a0b111cdf0a86a607c9076e6158ca

Observation 340fa0a7-083c-4f74-9c30-7a9ce73f343e · inbound

GroupDPO: Memory efficient Group-wise Direct Preference Optimization cites this paper.

GroupDPO: Memory efficient Group-wise Direct Preference Optimization Length Desensitization in Direct Preference Optimization

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:48.709683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T09:43:18.432084Z digest=sha256:479d3d3cf94e7743ebcbe8c1245699f02bf889dadb701be6235e73b930fa91d7

Observation 0b5c5264-4c30-4508-b3c8-3df38d12a197 · inbound

Learning to Control Summaries with Score Ranking cites this paper.

Learning to Control Summaries with Score Ranking Length Desensitization in Direct Preference Optimization

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.366133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T06:41:27.954392Z digest=sha256:973e0a4a22faab9314806b6aa4ea18a88f80d6ab0d19d2d1df82cb819ee613f4

Observation df62599c-a6da-476b-8033-298497b91682 · inbound

From Coarse to Fine: Self-Adaptive Hierarchical Planning for LLM Agents cites this paper.

From Coarse to Fine: Self-Adaptive Hierarchical Planning for LLM Agents Length Desensitization in Direct Preference Optimization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:10.387921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-08T08:10:36.579810Z digest=sha256:b45807739d555e49e1ceb980374b2b227840755d9145e4933b7bf709d3268872

Observation bdb75aa9-7689-4f60-aefa-d4d84e10b311 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Length Desensitization in Direct Preference Optimization

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.693757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:710396ef0035327f54bd4ee3ed1beb76a323a7da45c1e3a7eb85a2f9220e0daa