Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T13:09:54.433720Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2509.23102.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T13:09:54.433720Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T14:26:06.428076Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-20T14:28:21.295190Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 990ce93d-a989-401a-888f-8910532dc3fb · outbound
Multiplayer Nash Preference Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90faa925-f9ee-4c16-8373-4b82a5ff6928 · outbound
Multiplayer Nash Preference Optimization Human Alignment of Large Language Models through Online Preference Optimisation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87a54057-82c6-411a-95d6-f5c709ca948e · outbound
Multiplayer Nash Preference Optimization Evaluating Large Language Models Trained on Code
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f265f3d-330e-4ab6-b074-ad16d690f936 · outbound
Multiplayer Nash Preference Optimization Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a484db74-b16e-4104-8868-e16e157377c1 · outbound
Multiplayer Nash Preference Optimization Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 599039fa-d35e-4424-951b-8c5c78476c6a · outbound
Multiplayer Nash Preference Optimization UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f28c62cc-cde8-4b72-b882-e0e16d0f2cd3 · outbound
Multiplayer Nash Preference Optimization RLHF Workflow: From Reward Modeling to Online RLHF
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b6d2d59-44a1-493b-a62d-710e3aeb177d · outbound
Multiplayer Nash Preference Optimization Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c017d731-c26d-4c09-a187-5bde5969b007 · outbound
Multiplayer Nash Preference Optimization KTO: Model Alignment as Prospect Theoretic Optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76f6a24b-e98e-4788-ac89-2b54f12811da · outbound
Multiplayer Nash Preference Optimization Robust Preference Optimization through Reward Model Distillation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e4e2ea0-d071-4431-9fb4-0fad811c067d · outbound
Multiplayer Nash Preference Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e318a42-e6b4-4224-8dd0-dd1bc80105a0 · outbound
Multiplayer Nash Preference Optimization Measuring Massive Multitask Language Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f66dc85-7e8a-49d5-8fe7-94bcc5b8e78d · outbound
Multiplayer Nash Preference Optimization ORPO: Monolithic Preference Optimization without Reference Model
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a53c05b3-003c-453e-9298-06c4da0769c6 · outbound
Multiplayer Nash Preference Optimization From live data to high-quality benchmarks: The arena-hard pipeline.Blog post.[Accessed 07-02-2025]
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0315044d-b440-40a0-9fd4-df3c56a3b16a · outbound
Multiplayer Nash Preference Optimization TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8ab4fa3-0eda-43c9-b095-3099760038b8 · outbound
Multiplayer Nash Preference Optimization Decoupled Weight Decay Regularization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59e7694c-76f2-4396-a14c-70c57178ce24 · outbound
Multiplayer Nash Preference Optimization The hidden link between rlhf and contrastive learning.arXiv preprint arXiv:2506.22578
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60b45b80-0383-4fa7-8acb-c871fe844295 · outbound
Multiplayer Nash Preference Optimization Nash Learning from Human Feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f2724e5-37d9-45c3-8e8f-0be4bf8d592e · outbound
Multiplayer Nash Preference Optimization Pre-dpo: Improving data utilization in direct preference optimization using a guiding reference model.arXiv preprint arXiv:2504.15843
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a85e76b-341b-4052-832a-3d4680e1b659 · outbound
Multiplayer Nash Preference Optimization Disentangling Length from Quality in Direct Preference Optimization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bbf2b40-5962-4839-b7d1-094cdec1c94d · outbound
Multiplayer Nash Preference Optimization Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5cea3a4-c3cd-443b-96d8-cc17c2d6d8c4 · outbound
Multiplayer Nash Preference Optimization Proximal Policy Optimization Algorithms
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40c2dd33-6f91-4f15-9b98-a6e8e4c06910 · outbound
Multiplayer Nash Preference Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b233b10-28b9-47d9-b0cd-ab49eb3aba8d · outbound
Multiplayer Nash Preference Optimization A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a01f68e3-aa5b-46ee-b15d-5caafd728814 · outbound
Multiplayer Nash Preference Optimization Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e404c24-35ac-48b4-8b18-4832b6d28ae9 · outbound
Multiplayer Nash Preference Optimization Gemini: A Family of Highly Capable Multimodal Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5128dbfc-81aa-47b3-8f2c-44cc02513757 · outbound
Multiplayer Nash Preference Optimization Gemma 2: Improving Open Language Models at a Practical Size
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d30fb4d1-e33d-41f5-81fb-75d86b60e1bd · outbound
Multiplayer Nash Preference Optimization Causal Confusion and Reward Misidentification in Preference-Based Reward Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16949879-736a-40a2-b1a9-74e0f343206d · outbound
Multiplayer Nash Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e630a4a0-faaa-45a8-8123-a41233ba089d · outbound
Multiplayer Nash Preference Optimization Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c069f4b2-bd1d-4910-bd7c-1ac1976f5b5e · outbound
Multiplayer Nash Preference Optimization Self-Play Preference Optimization for Language Model Alignment
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc1d093a-2bae-44c8-a20e-caad02b45eec · outbound
Multiplayer Nash Preference Optimization Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a187e3c4-ed2a-4fd9-bc9c-25ad2c792f07 · outbound
Multiplayer Nash Preference Optimization Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea2b1167-7325-4b3d-a3e5-78e03367dc93 · outbound
Multiplayer Nash Preference Optimization Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2519bbae-2272-4422-9df6-d6745d1c3c19 · outbound
Multiplayer Nash Preference Optimization Qwen3 Technical Report
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d2ebbe0-42e6-41b1-a165-1904522704bc · outbound
Multiplayer Nash Preference Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b517741-e3b3-42c5-a97d-37e8906b71d3 · outbound
Multiplayer Nash Preference Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e24e9a2-3a0e-460d-bb4d-1c0251ef9412 · outbound
Multiplayer Nash Preference Optimization HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8df3fe8-36b7-4734-82a2-68f7d0244696 · outbound
Multiplayer Nash Preference Optimization Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da2b461c-efe9-45b4-95bb-25728d5e8ff4 · outbound
Multiplayer Nash Preference Optimization Instruction-Following Evaluation for Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95a1eeb1-752b-4012-a402-27bab3384241 · outbound
Multiplayer Nash Preference Optimization Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c95817d-50dd-43a6-b6c0-f3a4f254a0ed · outbound
Multiplayer Nash Preference Optimization WPO: Enhancing RLHF with Weighted Preference Optimization
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d4d6ceb-0672-4a88-89dc-d9c1badf1550 · outbound
Multiplayer Nash Preference Optimization Self-play methods like SPIN (Chen et al., 2024), SPPO (Wu et al., 2024), and INPO (Zhang et al., 2025b) use no- regret dynamics, while Pre-DPO (Pan et al
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a72b05d9-8af9-4b55-b7a2-d928e8501565 · outbound
Multiplayer Nash Preference Optimization More recent methods, such as ONPO (Zhang et al., 2025a) and EGPO (Zhou et al., 2025), introduce optimism and extragradient techniques for stable convergence under noisy preferences
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f46fac3-4fcb-466e-bc98-8f3d9eb3a37b · outbound
Multiplayer Nash Preference Optimization INPO (Zhang et al., 2025b) is reproduced according to the settings described in the paper
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c7501f8-5662-4941-ab8a-a1d0eefd0b6e · outbound
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e66d95a-469f-4e4b-a4b8-32ce090948d0 · outbound
Multiplayer Nash Preference Optimization logπθ y+t π′t y+t −logπθ y−t π′t y−t −η 2 #2 , whereπ′t= argminπ E(x,y+t ,y−t)∼Dt
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c4ff70f-d39e-4c71-b097-6b5b5a848949 · outbound
Multiplayer Nash Preference Optimization Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 069db180-de2f-4c4e-a638-63ad38bbb031 · inbound
Towards General Preference Alignment: Diffusion Models at Nash Equilibrium Multiplayer Nash Preference Optimization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1091f69-54be-48f8-bfb9-9b2a54ebea9a · inbound
Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment Multiplayer Nash Preference Optimization
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.