Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2407.13399.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T18:32:21.033460Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T01:09:19.256384Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e3f076e6-6916-4baf-95d1-94ee9fcb0d54 · inbound
Preference learning made easy: Everything should be understood through win rate Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e81486b-232c-4b7b-b93b-ba130a1d6d1d · inbound
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c726fd7-51c0-4dcb-858c-88becab87c76 · inbound
Learning a Pessimistic Reward Model in RLHF Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf75a51a-b4c8-4e9d-8880-63e46bd00e07 · inbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b55ea79f-9f8e-422d-bb49-586c347b5e54 · inbound
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 261636fd-99f3-48c9-865a-8eac72226e6d · inbound
On a few pitfalls in KL divergence gradient estimation for RL Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba6da69a-17ba-4890-9235-cc4380afde66 · inbound
Failure Modes of Maximum Entropy RLHF Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 684ffb9a-8e32-4797-9e20-825ddc225477 · inbound
SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0a0d9cd-8e8a-4aad-9c03-ab3df754fcbd · inbound
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd096e3c-077c-45c6-87a2-994515c96177 · inbound
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e48175f-fc46-4c2b-961c-e58cc24da00e · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 793cd586-baf2-4a3c-b013-11bce3c0e45c · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f11d8a3a-a01a-40b3-872b-71693b7efc55 · inbound
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e3f67899-95f2-467a-8963-69b9da1675cc · inbound
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 21ff4e1c-e768-4139-b5d8-986c02e676a9 · inbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 95c21f44-3d62-4d8f-a099-d65dd3e4bfe9 · inbound
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5ea18bdd-b9d6-4be9-877d-4dd67df471c0 · inbound
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 476175a0-9d98-4780-88c5-4ec70b1add6f · inbound
Which Pairs to Compare for LLM Post-Training? Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6af3265a-6d02-463b-8756-65fbcb111345 · inbound
Normalized Rewards for Preference Optimization Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.