Pith. sign in

Paper Citation Record · LEDGER

Multiplayer Nash Preference Optimization

As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2509.23102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.23102 v3

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T13:09:54.433720Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T14:26:06.428076Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-20T14:28:21.295190Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact38
  • verified fuzzy6
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 990ce93d-a989-401a-888f-8910532dc3fb · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Multiplayer Nash Preference Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.010757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:e3899c113c755e9c2037421aaad5287817843eade938d06b23353af293bb8374

Observation 90faa925-f9ee-4c16-8373-4b82a5ff6928 · outbound

This paper cites Human Alignment of Large Language Models through Online Preference Optimisation.

Multiplayer Nash Preference Optimization Human Alignment of Large Language Models through Online Preference Optimisation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.031523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:493d61db17f95339b52c716ae48df525d1f38927961a4313384f23d7ac3a4c6b

Observation 87a54057-82c6-411a-95d6-f5c709ca948e · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Multiplayer Nash Preference Optimization Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:23.992454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:85615a1eb54e449a2f0c84044795d0fa50ba438370c003605c299f6cf6ff670d

Observation 4f265f3d-330e-4ab6-b074-ad16d690f936 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Multiplayer Nash Preference Optimization Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:23.996253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:436ad5dd47c77e3b2aca770302f7a83079e49dd608f3c94829f95708513a8d92

Observation a484db74-b16e-4104-8868-e16e157377c1 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Multiplayer Nash Preference Optimization Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:23.969722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:822b940bafadf944068505a8208bed47d7c49725c8598496391f455939a989ec

Observation 599039fa-d35e-4424-951b-8c5c78476c6a · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Multiplayer Nash Preference Optimization UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.049816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:f775b8fc3fc8f3a69678181b005e6b6ae8cce962aaabb653a533f763cc082ae5

Observation f28c62cc-cde8-4b72-b882-e0e16d0f2cd3 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Multiplayer Nash Preference Optimization RLHF Workflow: From Reward Modeling to Online RLHF

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.060389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:f811f62efac333cccf7f51de17887257c0562e88f13b72554c5df9b815ae8e61

Observation 4b6d2d59-44a1-493b-a62d-710e3aeb177d · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Multiplayer Nash Preference Optimization Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.046566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:701eb116ee869fac76ea22abd6e91fd5e11fee3c325bb2ef8720733ee2819d58

Observation c017d731-c26d-4c09-a187-5bde5969b007 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Multiplayer Nash Preference Optimization KTO: Model Alignment as Prospect Theoretic Optimization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.057089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:d29747b85cfc9c67ca44dc242e03dd537dfaacd1218a82519ad8cf2e99d38d92

Observation 76f6a24b-e98e-4788-ac89-2b54f12811da · outbound

This paper cites Robust Preference Optimization through Reward Model Distillation.

Multiplayer Nash Preference Optimization Robust Preference Optimization through Reward Model Distillation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.953881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:ccf3276a9c38455a8369fbc8e69a60febe735c6eaedde433cde7bc8c31e9cba1

Observation 1e4e2ea0-d071-4431-9fb4-0fad811c067d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Multiplayer Nash Preference Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.018070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:d8ec9f8f2bc3323016229b5bbc6f9781b9b9d6390f75fb47a09077350c1216df

Observation 8e318a42-e6b4-4224-8dd0-dd1bc80105a0 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Multiplayer Nash Preference Optimization Measuring Massive Multitask Language Understanding

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.035030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:2c969a3f8284e2accb847c79c1251af8ffb770994bb0cf5d34e45263fd89a2cc

Observation 9f66dc85-7e8a-49d5-8fe7-94bcc5b8e78d · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

Multiplayer Nash Preference Optimization ORPO: Monolithic Preference Optimization without Reference Model

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.071919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:260eda5119f4a18ca19af9f1d7a75e40b97e31b526102edf90f8c55eff13c5dc

Observation a53c05b3-003c-453e-9298-06c4da0769c6 · outbound

This paper cites From live data to high-quality benchmarks: The arena-hard pipeline.Blog post.[Accessed 07-02-2025].

Multiplayer Nash Preference Optimization From live data to high-quality benchmarks: The arena-hard pipeline.Blog post.[Accessed 07-02-2025]

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:11:25.001298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:e6a32023fa066e8ff5ce83cd82f4e2b41d6850a7d6953b8b2d91d0641d376c7d

Observation 0315044d-b440-40a0-9fd4-df3c56a3b16a · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Multiplayer Nash Preference Optimization TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.091395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:a998c54a674b2ff0ffa40a54e1c337643fb39e79c83e2d4a9b9d315fbca60ca6

Observation e8ab4fa3-0eda-43c9-b095-3099760038b8 · outbound

This paper cites Decoupled Weight Decay Regularization.

Multiplayer Nash Preference Optimization Decoupled Weight Decay Regularization

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:23.987973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:f9f9e1c7b7367a5e5fe0e979cc2c4aa61e002a58cc15d9cfaab70066545e0cb5

Observation 59e7694c-76f2-4396-a14c-70c57178ce24 · outbound

This paper cites The hidden link between rlhf and contrastive learning.arXiv preprint arXiv:2506.22578.

Multiplayer Nash Preference Optimization The hidden link between rlhf and contrastive learning.arXiv preprint arXiv:2506.22578

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.949194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:0f7e681bbe81ff622feabace6f9e2344c39059395021ab5df00d87453b0eaf79

Observation 60b45b80-0383-4fa7-8acb-c871fe844295 · outbound

This paper cites Nash Learning from Human Feedback.

Multiplayer Nash Preference Optimization Nash Learning from Human Feedback

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:11:23.958548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:ba82ba772ee9820e8f76449ccdf647ccc443f71794acf75247dca3a998ad9f5d

Observation 1f2724e5-37d9-45c3-8e8f-0be4bf8d592e · outbound

This paper cites Pre-dpo: Improving data utilization in direct preference optimization using a guiding reference model.arXiv preprint arXiv:2504.15843.

Multiplayer Nash Preference Optimization Pre-dpo: Improving data utilization in direct preference optimization using a guiding reference model.arXiv preprint arXiv:2504.15843

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.098915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:3f54e716ff22fd14e4d4957cccf348ba529d108febec5ce39f89a1834d4e9775

Observation 4a85e76b-341b-4052-832a-3d4680e1b659 · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Multiplayer Nash Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.001156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:65c551fb5d8496cafc1909640637b41432c5ad3a7c8d892eff4fa4504d3347c5

Observation 4bbf2b40-5962-4839-b7d1-094cdec1c94d · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Multiplayer Nash Preference Optimization Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.944650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:e8cadf9484699004a82ea5c2c13d53b0e339c7c8a51f8e213354d28b4f33de19

Observation d5cea3a4-c3cd-443b-96d8-cc17c2d6d8c4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Multiplayer Nash Preference Optimization Proximal Policy Optimization Algorithms

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:23.939607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:7b451baa27419a284c706af0baa5b98310d5981880031f865337ec16fe22baf1

Observation 40c2dd33-6f91-4f15-9b98-a6e8e4c06910 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Multiplayer Nash Preference Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:23.962835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:6702149538aee6ebd9c2a9aab234432a87bb43aa07c95435f998e271ca78a797

Observation 2b233b10-28b9-47d9-b0cd-ab49eb3aba8d · outbound

This paper cites A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games.

Multiplayer Nash Preference Optimization A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.978594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:b96d27ad612c035fed14152bc547cfb3fd8de258b14f764765086b05d7fda76b

Observation a01f68e3-aa5b-46ee-b15d-5caafd728814 · outbound

This paper cites Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment.

Multiplayer Nash Preference Optimization Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.076019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:6278936a247fe8b4d5aa8c5dbf723791dea50260c8494beb57213729f0599761

Observation 5e404c24-35ac-48b4-8b18-4832b6d28ae9 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Multiplayer Nash Preference Optimization Gemini: A Family of Highly Capable Multimodal Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.079371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:f77b15dc62bced654bcae56d3947900b5f2e362112564362d54199c03020b7e7

Observation 5128dbfc-81aa-47b3-8f2c-44cc02513757 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Multiplayer Nash Preference Optimization Gemma 2: Improving Open Language Models at a Practical Size

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.083212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:4a1107692d1339fab5b67b4a10dee5de7c742b7e6eb69b7625a888d166ea0b35

Observation d30fb4d1-e33d-41f5-81fb-75d86b60e1bd · outbound

This paper cites Causal Confusion and Reward Misidentification in Preference-Based Reward Learning.

Multiplayer Nash Preference Optimization Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:11:24.068249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:1df42457fadbc260304e47f7f18b177fe61c61bf1472965699f3930edb8647ef

Observation 16949879-736a-40a2-b1a9-74e0f343206d · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Multiplayer Nash Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.039547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:bbbeccc03fa584740f61e5041595a2b9f0f8bf876394ecb29c653bc573b27d7e

Observation e630a4a0-faaa-45a8-8123-a41233ba089d · outbound

This paper cites Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment.

Multiplayer Nash Preference Optimization Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.006252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:d43818d3efb572618a5b579abd9efa77c614e8a59f1b5e3c1af3335a4ede9af1

Observation c069f4b2-bd1d-4910-bd7c-1ac1976f5b5e · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Multiplayer Nash Preference Optimization Self-Play Preference Optimization for Language Model Alignment

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:11:24.087385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:72e93c32dfb79a735b4bf29e6b60f279641c3ed25846c726888268e8b2097bc1

Observation cc1d093a-2bae-44c8-a20e-caad02b45eec · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Multiplayer Nash Preference Optimization Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.053907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:9c092c12d1118601d763f676fc5a7be92a63fb65add973d20a1b1a851fb26430

Observation a187e3c4-ed2a-4fd9-bc9c-25ad2c792f07 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Multiplayer Nash Preference Optimization Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.027237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:2775e666f5cfda0070f267f1fa35ba8c7c086cce755de9826e39fa282a2fbacb

Observation ea2b1167-7325-4b3d-a3e5-78e03367dc93 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

Multiplayer Nash Preference Optimization Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.983816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:c4716d54289e921ab9ee89118f669a2335807030c9a31e5d3e79f85ec8a2abc8

Observation 2519bbae-2272-4422-9df6-d6745d1c3c19 · outbound

This paper cites Qwen3 Technical Report.

Multiplayer Nash Preference Optimization Qwen3 Technical Report

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.108463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:cff005e3cc4e46236f620e8dabc750078dfed4979973d7b02526d09c84f46e86

Observation 0d2ebbe0-42e6-41b1-a165-1904522704bc · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Multiplayer Nash Preference Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.043054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:87772d93a62509a541ec1c09477c6639486ec51878bbac26ac793f0aa7d749b1

Observation 7b517741-e3b3-42c5-a97d-37e8906b71d3 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Multiplayer Nash Preference Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.064514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:ca29e0b91eb3d9e7a9128a6aaf54c1a86a61b193ab8046f5754a8aa411e01855

Observation 6e24e9a2-3a0e-460d-bb4d-1c0251ef9412 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Multiplayer Nash Preference Optimization HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.014337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:4880bb4b5b9d63112b5ee1dadaeb3a4b90cb20af6198f25aa588c2416754288b

Observation d8df3fe8-36b7-4734-82a2-68f7d0244696 · outbound

This paper cites Improving LLM General Preference Alignment via Optimistic Online Mirror Descent.

Multiplayer Nash Preference Optimization Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.022728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:fc9f2e061a9f03bbfd4109740c7407e141aba69681aa35f599eee9348051cb09

Observation da2b461c-efe9-45b4-95bb-25728d5e8ff4 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Multiplayer Nash Preference Optimization Instruction-Following Evaluation for Large Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.094783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:0b8da894da0a33c1252a262fcd73367471c2a34ffb35082c7570b3586798c2c9

Observation 95a1eeb1-752b-4012-a402-27bab3384241 · outbound

This paper cites Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback.

Multiplayer Nash Preference Optimization Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.103703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:df179d750236921c7fb4e40c121888a642231ca6dd321400b3e2d79814594820

Observation 1c95817d-50dd-43a6-b6c0-f3a4f254a0ed · outbound

This paper cites WPO: Enhancing RLHF with Weighted Preference Optimization.

Multiplayer Nash Preference Optimization WPO: Enhancing RLHF with Weighted Preference Optimization

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.974150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:2b3b62e7b7c3b4710fa44219bf4b844c66f3f04e83eb3845ffd16b609e06dbca

Observation 3d4d6ceb-0672-4a88-89dc-d9c1badf1550 · outbound

This paper cites Self-play methods like SPIN (Chen et al., 2024), SPPO (Wu et al., 2024), and INPO (Zhang et al., 2025b) use no- regret dynamics, while Pre-DPO (Pan et al.

Multiplayer Nash Preference Optimization Self-play methods like SPIN (Chen et al., 2024), SPPO (Wu et al., 2024), and INPO (Zhang et al., 2025b) use no- regret dynamics, while Pre-DPO (Pan et al

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:11:24.993887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:d93025339bf0e007ca2aee788f5e276a55f44251136ea812e2d16116f0768b9d

Observation a72b05d9-8af9-4b55-b7a2-d928e8501565 · outbound

This paper cites More recent methods, such as ONPO (Zhang et al., 2025a) and EGPO (Zhou et al., 2025), introduce optimism and extragradient techniques for stable convergence under noisy preferences.

Multiplayer Nash Preference Optimization More recent methods, such as ONPO (Zhang et al., 2025a) and EGPO (Zhou et al., 2025), introduce optimism and extragradient techniques for stable convergence under noisy preferences

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:11:24.996355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:b4e187be61ecadf077ceb04316b159e8d59509597972a69082c0dbfa462c607e

Observation 0f46fac3-4fcb-466e-bc98-8f3d9eb3a37b · outbound

This paper cites INPO (Zhang et al., 2025b) is reproduced according to the settings described in the paper.

Multiplayer Nash Preference Optimization INPO (Zhang et al., 2025b) is reproduced according to the settings described in the paper

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:11:24.991274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:283fea676b3a5b0414df5ff097f24208ed23078d299dc9c3b96b3dab488deeb8

Observation 8c7501f8-5662-4941-ab8a-a1d0eefd0b6e · outbound

This paper cites minimal.

Multiplayer Nash Preference Optimization minimal

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:11:24.988894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:9ee3254ef1360fa5951e7c8553889fca3836067e130b52973d38078edc9506a3

Observation 8e66d95a-469f-4e4b-a4b8-32ce090948d0 · outbound

This paper cites logπθ y+t π′t y+t −logπθ y−t π′t y−t −η 2 #2 , whereπ′t= argminπ E(x,y+t ,y−t)∼Dt.

Multiplayer Nash Preference Optimization logπθ y+t π′t y+t −logπθ y−t π′t y−t −η 2 #2 , whereπ′t= argminπ E(x,y+t ,y−t)∼Dt

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:11:24.986614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:317de9f2f5cda6d0f2b17030619716c014d196bec9f1576701e984f9830bd95e

Observation 1c4ff70f-d39e-4c71-b097-6b5b5a848949 · outbound

This paper cites an unresolved cited work.

Multiplayer Nash Preference Optimization Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-18T13:11:24.998801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:2d717b9f195635d4e1fa331e7249572b5ca3144fd9e88467f3ad4a6943c67fea

Pith citing papers

Observation 069db180-de2f-4c4e-a638-63ad38bbb031 · inbound

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium cites this paper.

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium Multiplayer Nash Preference Optimization

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:56:06.889210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T16:54:58.732444Z digest=sha256:3f845655ed1ee87d3104fe0cf4b8fe9c6f3e6081cb24a27d0e3d41bfd8222537

Observation c1091f69-54be-48f8-bfb9-9b2a54ebea9a · inbound

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment cites this paper.

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment Multiplayer Nash Preference Optimization

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:28:21.296641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T14:26:06.428076Z digest=sha256:a3309672bbef55ab2d51177e59d4c38c6c52cc2ac1f5a7ff31471d6aa376b5a2