Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2407.00617.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:40:58.431783Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T06:15:06.516739Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 78f2aad9-cdc8-4194-835f-960b014972e0 · inbound
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bea63a5-2892-41ef-8df9-08596b5ac2cb · inbound
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3367dc2-ea39-4b96-bdba-d8f844872624 · inbound
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b22110-0364-4920-8647-55160f8ecb2a · inbound
Towards General Preference Alignment: Diffusion Models at Nash Equilibrium Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3de88b08-e481-495b-9ba6-eb8aae212352 · inbound
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 494ba96e-1b3c-4bb4-8a3a-7784ce071891 · inbound
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4b125e06-5f14-43d4-85bb-79ce9038f6cb · inbound
Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0b3fc829-0a8d-4059-8640-cab73bdfc385 · inbound
Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd179071-170e-4843-bc7f-41dad63fcca3 · inbound
Common-agency Games for Multi-Objective Test-Time Alignment Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.