Pith. sign in

Paper Citation Record · LEDGER

Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2309.16240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.16240 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:11:31.734703Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0c765219-aed3-4e29-b9f8-1d133a61dde1 · inbound

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization cites this paper.

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:05:46.880609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:03:53.675107Z digest=sha256:0e2e87df8cae35a660e2d99f0953bce6e04144cf391cb39713737b0ec79e71be

Observation fe761636-83cd-4a26-a6ab-06a340bb0116 · inbound

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models cites this paper.

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:31.734703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:31.734703Z digest=sha256:ef5b856096c6e6155ad627b66f99aec49f8ebf8997d97d1c4187a6a42d5d932e

Observation 19ee0da3-f26e-4d8f-bbe5-0387e8bcc2fb · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.062867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.062867Z digest=sha256:73321c71214fd75693d45a2bbc78afbee3b5508e1e7dd8d38f4ba28b0126dfdb

Observation 76980b50-eb90-47f5-9a42-dc7bc440db34 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.083529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.083529Z digest=sha256:fa0c15ea9b17a034ecbcbd85bb87353bc34b52a25f602462e46a8f448cbcf8e2

Observation 16949879-736a-40a2-b1a9-74e0f343206d · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.039547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:bbbeccc03fa584740f61e5041595a2b9f0f8bf876394ecb29c653bc573b27d7e

Observation 951663bd-9add-4f69-9910-0ce7a4bcc7e2 · inbound

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning cites this paper.

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T11:27:22.818786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:27:22.818786Z digest=sha256:14c317c8e2bbd09ee4650b4a8b27c707ce9307e775d73da674710bed72d898d9

Observation dfb966f1-9139-41de-8afa-bb2b116a9663 · inbound

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model cites this paper.

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T14:03:36.093027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:03:36.093027Z digest=sha256:f87ef12baaf4b5f6f7661e7bcc6effc0632dde162883e103c8bd429e670e8a2e

Observation 1fddb961-0819-47c9-b429-8a4ed6fc1a58 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-09T20:37:32.126851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:c74ed6d22a596cb7aad44dc4cfe85c5eec7d469baa2e98ae5b298f289133e7d1

Observation 19973347-87d4-450a-83c0-7911744b8fe1 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.110302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:ca859b342710fdee422095e6d556414bd59bebd7afa7780cc09cc98996536515

Observation ada1bef8-3067-4a22-9776-478fcd096770 · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:46:10.397358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T13:55:22.422923Z digest=sha256:b56e0e3a03d8c559b032be466981d123b2ba2999f5b0d37d24747e2e348a4d7a

Observation 326e914d-0920-427f-9129-5c084eb5195f · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T09:04:04.467772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T09:02:28.407064Z digest=sha256:819d0db09c7677de5e9f2edaadfbee57638a57596198ab46882b3c664df5c99a

Observation c695d87b-dea2-4534-a79a-71f609a6356b · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:35:58.803821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:2fc8dc743f7c9dd6fc1640f40e425257f686df8d384a8b6036b9cbdb8eebe896

Observation 220a304e-7a5e-4c27-91f4-1e32b9584b29 · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.019871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:545e6bfcfa28f0dbaf9600a56f8c3c9120075fd732ab4e8eb0165295ec2bf0b3

Observation e0144991-53b8-4943-8b53-df773b4bd54a · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:07:27.955425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:03:00.503644Z digest=sha256:20b9b562c1d25f350760fcb0c9121a435eb7b50da260df375e55cedd021354cf

Observation 0ba99aa1-1700-4fda-b7cc-82e8faffc46a · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:28.247901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:06:01.667173Z digest=sha256:28bd0756ced4a8d8d1f18e64f6c18ccb09fdcf6594fd463acbc7aaa6625b7577

Observation 243a7812-7707-4f01-a0f6-c10371f33d4f · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.245928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:64e29e15cfe9eb2d287c4034286df7d2968cbc4eebdfd7d234d47632c531cd11

Observation c9f3e4ec-e3a0-4560-9122-0cf4de55d24c · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.415473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:f8b08a785683db720c5a049a7a481d69d875d04646c2fff9b5848c40b6094cfb

Observation 6c795686-ad62-4af4-85df-7b8a941c1ae7 · inbound

Rethinking the Role of Temperature in Large Language Model Distillation cites this paper.

Rethinking the Role of Temperature in Large Language Model Distillation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:42:50.064284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T23:25:55.849389Z digest=sha256:b5feb54b6b383248e2165cdc6bfaafbdd27c1b0397a9266c1a8719ca958c642a

Observation 3c744dee-790b-4f01-b78a-8ec48199d654 · inbound

SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter cites this paper.

SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:56.186095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:53:22.159132Z digest=sha256:94ce358191338a40030f7a6cae43676d344cede88ed3688c462200d96a1a4df2

Observation d976553e-09a8-4127-a434-1fbf0a4f9173 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 189

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.380972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:74269eb0cf1171084bf006c740a174d631c8b95f21106a08655c11747dd9c617

Observation c6053d29-2b05-4f20-9194-e5acdafbac21 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.407194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.407194Z digest=sha256:a7939b07ad6f4c86dc1d61bc9401ccf1f91fbe89c90026e17c964d927c5dbe9c

Observation ff8e3dd2-af1b-4ce2-b914-fb03b68635aa · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 158

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:de1e6803182a451e00ee463646e0500e79dadbebca3faaf90710aac221f4dd1a

Observation 1d1df6fb-b34e-4886-bc4b-e260e9ca3516 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 159

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:50.233277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:50.233277Z digest=sha256:8ae3db4647c78e412f54a2fda9b3218c9505d16c12b607bb1ffa990909c69c62

Observation c8fc713c-d568-4789-a4ca-d889d967af12 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:58.977272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:58.977272Z digest=sha256:a66d76e0776618aaede7dfcc8180b06928bb18036ded16e4164aaab164c2faa0