Pith. sign in

Paper Citation Record · LEDGER

Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2309.16240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.16240 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:58:38.850516Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0c765219-aed3-4e29-b9f8-1d133a61dde1 · inbound

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization cites this paper.

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:05:46.880609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-23T19:03:53.675107Z digest=sha256:fbd15f9c245daa1a4b47adc51200f55ff5f7b1b12fb3869b274049c9cf47adc5

Observation 4f4fdb30-2753-4bae-b614-7210da193576 · inbound

Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment cites this paper.

Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:39:52.620622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:39:52.620622Z digest=sha256:8726b1ade21dfd647062953bd72b6d3ad54ba2abb4af5da57e151979e8c5d442

Observation 250473cb-99de-4577-8adf-89c68256c17d · inbound

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd cites this paper.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.925208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.925208Z digest=sha256:fae1b8bb2ecf2fe6bdffa970404c29977d24eabb366627b69f07247df56cbd31

Observation 4e4ee5fd-a9b9-4f08-95c6-cad6428fcaa1 · inbound

Diversity By Design: Leveraging Distribution Matching for Offline Model-Based Optimization cites this paper.

Diversity By Design: Leveraging Distribution Matching for Offline Model-Based Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T22:39:45.735113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:39:45.735113Z digest=sha256:35f3c7c67a1ec923b5f1d5f687212b723fd435fd2e18d959430fb394365469bd

Observation ce3367d6-9549-4bf5-bd50-2407d0d19eb9 · inbound

Policy-labeled Preference Learning: Is Preference Enough for RLHF? cites this paper.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.850516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.850516Z digest=sha256:89566664479750cbfd1079afba96d02e27b472b2b0786cbb11972ad4d2f07e74

Observation f549c0ef-2db6-4c81-8ddf-5fb06958979c · inbound

Bounding Neyman-Pearson Region with $f$-Divergences cites this paper.

Bounding Neyman-Pearson Region with $f$-Divergences Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:06:02.008752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:06:02.008752Z digest=sha256:20144b6ed057ca49032fcb343ded7465c291b6ee40fe0727abcb095f47612de3

Observation fe761636-83cd-4a26-a6ab-06a340bb0116 · inbound

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models cites this paper.

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:31.734703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:31.734703Z digest=sha256:41e328e47511ba7913b7272778b89e34b819e453d668349437453bb2ffda522d

Observation 19ee0da3-f26e-4d8f-bbe5-0387e8bcc2fb · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.062867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.062867Z digest=sha256:a3c7ad43c2b8e2c6c1c0e975509358c30a704c2d4d4192bb773ca2cd24adb584

Observation 76980b50-eb90-47f5-9a42-dc7bc440db34 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.083529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.083529Z digest=sha256:2c78f39e254e68e71ca2d295b0812ccedf7d86310f5c6c8994537ac3d5cf9024

Observation 8934c5be-c620-4ef0-b3b4-b4cf842d666f · inbound

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering cites this paper.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.537395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.537395Z digest=sha256:f947fb758916bcdd1049926e61e3e21dd0fb42ba66123510ff8ae19208406bf8

Observation 16949879-736a-40a2-b1a9-74e0f343206d · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.039547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:a9d6cc1f38d070cc98c9329a6fe777f03d10f1489e4a7b5f63d005ff930d87ca

Observation 951663bd-9add-4f69-9910-0ce7a4bcc7e2 · inbound

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning cites this paper.

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T11:27:22.818786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:27:22.818786Z digest=sha256:ac5e5d63f40cb9377076d3615e0dbb5f8dd7bb13b72b09951b7a41bdd6d5d740

Observation dfb966f1-9139-41de-8afa-bb2b116a9663 · inbound

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model cites this paper.

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T14:03:36.093027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:03:36.093027Z digest=sha256:189490bfae8718ae5cc50bbc6f75cd794ac632a08396b06851cc9826136f2a11

Observation 1fddb961-0819-47c9-b429-8a4ed6fc1a58 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-09T20:37:32.126851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:ba41dcc9c764207b28ce22716c82bed6a84dc3672ae823c4d53b727081c263e6

Observation 19973347-87d4-450a-83c0-7911744b8fe1 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.110302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:3b824686d1cce50482333776fad398e245992dbab352ddcb1c521c97eda13b40

Observation ada1bef8-3067-4a22-9776-478fcd096770 · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:46:10.397358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T13:55:22.422923Z digest=sha256:aa43a166fec6a3ec7a7d10e75966f9aae53555c5fc0e3bca9a49f5d184544fbf

Observation 326e914d-0920-427f-9129-5c084eb5195f · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T09:04:04.467772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-21T09:02:28.407064Z digest=sha256:4f1950d927e51cff014dce2d4ebed4fbdd54c841de021fd1419ce525c85ed34e

Observation c695d87b-dea2-4534-a79a-71f609a6356b · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:35:58.803821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:a8fd32c60b32f2e43c91ac5aa6f310d0fb6203fbc978d4458245435d13e8748b

Observation 220a304e-7a5e-4c27-91f4-1e32b9584b29 · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.019871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:0ae24a99045d1760fa11a4b867752db9af8a36cf1714e5b0d775ac824250713e

Observation e0144991-53b8-4943-8b53-df773b4bd54a · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:07:27.955425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T07:03:00.503644Z digest=sha256:c98305efc25221d4f8c8da1630bf9f7d6142463c21948ac109f78064a0b5bad9

Observation 0ba99aa1-1700-4fda-b7cc-82e8faffc46a · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:28.247901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T21:06:01.667173Z digest=sha256:3c19d1c7c818181aaaa2a60892ed6be80ca1d0202cfb8da55b495a289470d670

Observation 243a7812-7707-4f01-a0f6-c10371f33d4f · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.245928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:58556081d9d1a1abbb4f108f0c24891ee7d6bde03aee0917f1ce947a0aaec58b

Observation c9f3e4ec-e3a0-4560-9122-0cf4de55d24c · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.415473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:e3179f6d5c3ca00043ff79f96bc0fdefc57c7ce0d4415efb1778268795e8e5c3

Observation 6c795686-ad62-4af4-85df-7b8a941c1ae7 · inbound

Rethinking the Role of Temperature in Large Language Model Distillation cites this paper.

Rethinking the Role of Temperature in Large Language Model Distillation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:42:50.064284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T23:25:55.849389Z digest=sha256:98aa7c3df0e913b8c79d29035778aedd1db3c7e0612602cab3439eac918bfa54

Observation 3c744dee-790b-4f01-b78a-8ec48199d654 · inbound

SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter cites this paper.

SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:56.186095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T02:53:22.159132Z digest=sha256:e00a041957199b283c164c0f943d7e8dcc5ee15a6e04bbf55fda205fdca0c6f5

Observation d976553e-09a8-4127-a434-1fbf0a4f9173 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 189

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.380972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:7ec601d019deb9653092fe269a15f7357510fb774a5d2a223f4f06101a540a8b

Observation c6053d29-2b05-4f20-9194-e5acdafbac21 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.407194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.407194Z digest=sha256:26acfda113af5e561447002798faaebecd40b0ad6be327dded7cdda775ceb0c3

Observation ff8e3dd2-af1b-4ce2-b914-fb03b68635aa · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 158

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:ebafcb013d90e92fda4a6da96dae2622c27f2216866d7503e7d2cb2bff04eea9

Observation 1d1df6fb-b34e-4886-bc4b-e260e9ca3516 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 159

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:50.233277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:50.233277Z digest=sha256:8bd56cd3ebeefcbecd024eb74c4923910f25509fbaafafc91e87d5a1f4303741

Observation c8fc713c-d568-4789-a4ca-d889d967af12 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:58.977272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:58.977272Z digest=sha256:5d6cc917755af677f3cb0bea274fc30365fbe30d07e987a51a5bb37c15246364