Pith. sign in

Paper Citation Record · LEDGER

Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2407.13399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.13399 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:32:21.033460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T01:09:19.256384Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e3f076e6-6916-4baf-95d1-94ee9fcb0d54 · inbound

Preference learning made easy: Everything should be understood through win rate cites this paper.

Preference learning made easy: Everything should be understood through win rate Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.033460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.033460Z digest=sha256:72a467042999069da9c1f138352e359f6e6062f50d5fb42c85730891bc551956

Observation 3e81486b-232c-4b7b-b93b-ba130a1d6d1d · inbound

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO cites this paper.

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:09.183433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:09.183433Z digest=sha256:e1d4ac031bd301743e41b8237c892069a27626ac1729ee9983fbb734a33da10a

Observation 6c726fd7-51c0-4dcb-858c-88becab87c76 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.069235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.069235Z digest=sha256:c5dac967d76fb42fe89af2b4ec49627c3f935f9f08c0ba85694dfb62a9e77370

Observation cf75a51a-b4c8-4e9d-8880-63e46bd00e07 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.450115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.450115Z digest=sha256:e129947f8544539915a24c38004d71683b100c04904397574b58fcfd1bf3fa44

Observation b55ea79f-9f8e-422d-bb49-586c347b5e54 · inbound

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? cites this paper.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:08.217635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:08.217635Z digest=sha256:6f1d0aa978f40ecc0a27dd4c5e3fe472b9a707ea2480d7573a78be5be8d91130

Observation 261636fd-99f3-48c9-865a-8eac72226e6d · inbound

On a few pitfalls in KL divergence gradient estimation for RL cites this paper.

On a few pitfalls in KL divergence gradient estimation for RL Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:37.946422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:37.946422Z digest=sha256:942278f413e2ccda1df97faf53cbad0461a7f51ba41a158daf1215798cff76ab

Observation ba6da69a-17ba-4890-9235-cc4380afde66 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.869177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:ac3744f366d62952aa82d8f8b8e0bd510866be6142a261125d78c4b5e20d57e7

Observation 684ffb9a-8e32-4797-9e20-825ddc225477 · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:50.379903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:50.379903Z digest=sha256:7cf618c776cb6c9ec32e1fbac617b3283db1d8fc25b30e481e7cff64243ead36

Observation b0a0d9cd-8e8a-4aad-9c03-ab3df754fcbd · inbound

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model cites this paper.

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T14:03:35.433091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:03:35.433091Z digest=sha256:61a870a29fbc9fcc3fc69910349b319975187673bb55347de28063257c38f0fa

Observation dd096e3c-077c-45c6-87a2-994515c96177 · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:24.742624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:24.742624Z digest=sha256:88ac676a0a9fea5f39867d5c1e17f44cab82ceb75584de3257fa4f1ba42b033c

Observation 9e48175f-fc46-4c2b-961c-e58cc24da00e · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.570738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:ca0a933c02f92a4041e2f24b4f48386839a6e7b31454f04d9ac971f475adef61

Observation 793cd586-baf2-4a3c-b013-11bce3c0e45c · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.400015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:6d84d45de5abdf507a502b51976e496aba40480242b711a007772d5b1d605221

Observation f11d8a3a-a01a-40b3-872b-71693b7efc55 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:10:42.400543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:6099577ce6af3506d6a3b6ca642458926c4e43fcc9d4675544b75555db813b35

Observation e3f67899-95f2-467a-8963-69b9da1675cc · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:05:09.669377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:d168d8319435a949983bc9e42975e6cffc8860966df023d9ce38fa5cd99184e7

Observation 21ff4e1c-e768-4139-b5d8-986c02e676a9 · inbound

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients cites this paper.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:08.493642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:85fef2b4676a6d6c5e681de4342b7ea5bd1174e22b6f36315a2cd5967a889d3d

Observation 95c21f44-3d62-4d8f-a099-d65dd3e4bfe9 · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:58.682722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:82b348500323003bc02c5ff88a66ce80f4a5d727255b99f908d3948185116f38

Observation 5ea18bdd-b9d6-4be9-877d-4dd67df471c0 · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:30.550722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:99e398aff18f27240fab5fe71f09fa8689e8b00a9cf8ed105143c0d68b490a64

Observation 476175a0-9d98-4780-88c5-4ec70b1add6f · inbound

Which Pairs to Compare for LLM Post-Training? cites this paper.

Which Pairs to Compare for LLM Post-Training? Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:09:19.259738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T20:37:19.221848Z digest=sha256:c600f1f1a58f0001eeb5084a442310af7c2ffbd98f095a44e195879ac5003843

Observation 6af3265a-6d02-463b-8756-65fbcb111345 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:58.145317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:58.145317Z digest=sha256:e984c9efe4e0d3464731c0883906800ed5cc9e464bac1eaee7b10671147d4f96