Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2404.03715.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.077519Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:49:37.721117Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation ff11211d-bc73-4838-908e-26f990d76380 · inbound
KTO: Model Alignment as Prospect Theoretic Optimization Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 35f6eb9c-ac9b-4b3f-858d-507259d3fdbb · inbound
Training Language Models to Self-Correct via Reinforcement Learning Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05481b0e-c6d2-4d15-9327-5f77272944cb · inbound
Process Reinforcement through Implicit Rewards Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3a9b81d3-fad4-400f-8a5b-9c988348fc5d · inbound
Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9384154-a853-48d5-8e9e-5e44e46c2217 · inbound
Reinforcement Learning from Human Feedback Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 205
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c8df08d7-fa7c-40af-90c7-3558c8025d87 · inbound
MPO: Multilingual Safety Alignment via Reward Gap Optimization Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf583308-af18-4836-b38e-fa24ff026969 · inbound
Learning a Pessimistic Reward Model in RLHF Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c612a59-241f-4f9b-a974-7f75e4c5f031 · inbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b6b2cad-3767-4e73-9981-da4ba7c71a92 · inbound
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f15eb2-1f01-4a41-b31c-5f4fe31cdcbc · inbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc2bec0-5495-40c1-a5e9-18cfb280fb3c · inbound
Debiasing Online Preference Learning via Preference Feature Preservation Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbf2b40-5962-4839-b7d1-094cdec1c94d · inbound
Multiplayer Nash Preference Optimization Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 375fda13-12d7-495a-ac33-4357f018d7f5 · inbound
Improved Bounds for Private and Robust Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f59d078f-612b-4b29-96c8-2b6e0079f15a · inbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13f26489-4421-4416-a739-9397259589c6 · inbound
CoAct: Co-Active LLM Preference Learning with Human-AI Synergy Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b200f40d-5da8-47c7-b62e-11bdda45921e · inbound
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5d79cd9f-3635-4db2-96a7-9fdd7e45ab3b · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 154
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a65fb4dd-ff9b-4306-8680-9a32351a9c0b · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 154
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 454872d7-679a-46f8-baa3-63fe12fbeb02 · inbound
Common-agency Games for Multi-Objective Test-Time Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 160f2efb-8f9e-48a8-9cb4-0241ff246eda · inbound
Learning from Language Feedback via Variational Policy Distillation Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eed979e2-bcfc-4804-a09e-539e2e4fa576 · inbound
Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39657248-29cf-426c-9954-0c4cfe4a1083 · inbound
S-SPPO: Semantic-Calibrated Self-Play Preference Optimization Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8a86bec4-31c5-4500-aca5-98e5479967fc · inbound
Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8af116e8-2e4a-4849-8af6-03779b44affc · inbound
AI Alignment From Social Choice Perspectives Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.