Pith. sign in

Paper Citation Record · LEDGER

Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2407.00617.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.00617 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:40:58.431783Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T06:15:06.516739Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 78f2aad9-cdc8-4194-835f-960b014972e0 · inbound

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization cites this paper.

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:58.431783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:58.431783Z digest=sha256:0dc7eee233571a382c33f224494a90532e97c918f7fb0bf618ebbe775a124b17

Observation 0bea63a5-2892-41ef-8df9-08596b5ac2cb · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:25.524406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:25.524406Z digest=sha256:35b3cc26b218ace4cfd095eba924dd3089a94e1936cf4f8e9a4d1459c9d5349d

Observation e3367dc2-ea39-4b96-bdba-d8f844872624 · inbound

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective cites this paper.

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:24.161428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:24.161428Z digest=sha256:2dd5821843f176e7f0bf3e5e9ebd6e4cfa997a1be17e7e345399be446dbb80fd

Observation 12b22110-0364-4920-8647-55160f8ecb2a · inbound

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium cites this paper.

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:06.862801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:54:58.732444Z digest=sha256:5ef4a141ce8fde37ea8ca8ac283237002a369c7e3b4694ea0392c948eac7305d

Observation 3de88b08-e481-495b-9ba6-eb8aae212352 · inbound

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective cites this paper.

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:55.373741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T01:19:58.448344Z digest=sha256:b18bf8f5f72b17b68a199a4f1423bd54b856601f602e6fe82d7083fed01794be

Observation 494ba96e-1b3c-4bb4-8a3a-7784ce071891 · inbound

Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions cites this paper.

Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:31.596943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T03:46:35.908522Z digest=sha256:be1666ea3316bb56fca331d6049e7ea01d12f26544b8ebbc9db7d30ac3989c63

Observation 4b125e06-5f14-43d4-85bb-79ce9038f6cb · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:26.511794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T04:15:24.919355Z digest=sha256:0f087edf95c04728a314033d09ac89e230b72d41989a45f6daa17ac6c0dc0671

Observation 0b3fc829-0a8d-4059-8640-cab73bdfc385 · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:05.102397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T22:03:05.102274Z digest=sha256:fe65ead3a8d992e8eb608748ccdb618b1b0e0ae0b6b35abb079d10ec134acb30

Observation fd179071-170e-4843-bc7f-41dad63fcca3 · inbound

Common-agency Games for Multi-Objective Test-Time Alignment cites this paper.

Common-agency Games for Multi-Objective Test-Time Alignment Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:15:06.520722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T06:14:53.685486Z digest=sha256:2baa70a463987a4625a30734887790bb14f668d0f48b028e72ccd08bc400eea1