Pith. sign in

Paper Citation Record · LEDGER

Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2502.16852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.16852 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:19.895162Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T19:32:01.290815Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 46aa73ec-3aab-4942-ba24-86c20a2b6047 · inbound

SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence cites this paper.

SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T23:49:14.854778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:49:14.854778Z digest=sha256:374d8030b53233ed3bf7a8cb0455ab017bdd84b76d9366e9f9c6b9c66f708fa4

Observation 1bc552a7-eb6f-4ba2-be9c-c225b57e6375 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.293338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:6b9d6bd199da44cb5eb8732ab84e38b8a05ca7c2a46be2bfa99a7bdf0864fa1c

Observation b2ce1dd3-9f99-4ea9-896e-d1d721dafc0d · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:19.895162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:19.895162Z digest=sha256:0b8a27eef509e68ac75779983df14b3893019b357819a123c5721f0420392992

Observation d8df3fe8-36b7-4734-82a2-68f7d0244696 · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.022728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:4f1e96347a36b05647552aa2f662e28af2cb3a156683962bc8c5c543e57c8836

Observation 644bbec5-543a-4555-819a-4c4ca0d16576 · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:25.789502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:25.789502Z digest=sha256:d6598beccd26c4fc7d21b33d65a6298b595fb51f5fc3bf9f02e3a0b5990ce468

Observation 564865ed-5513-407d-87d0-01759a8b9402 · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:25.670135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:25.670135Z digest=sha256:afa9ef2273e1f3a6f2a938bd6019162f0ba8e12cd7361a8c77f565c053f495ef

Observation 7a2fc9aa-b845-4e20-82e8-52e78af90a38 · inbound

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium cites this paper.

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:06.812051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T16:54:58.732444Z digest=sha256:e6a5d436487fc4efff1323008d2af14fd59b75054e79d4e19160f881df4cbd3f

Observation b3a700e2-9e5a-403b-8dac-0574b3b2c91d · inbound

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective cites this paper.

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:55.154910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T01:19:58.448344Z digest=sha256:4663b8b940d0e14849a34f39306c06f755b5bbd8df83414c2b510c992f477450

Observation a50f61f9-77eb-403f-b166-1f1aa1d90a84 · inbound

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment cites this paper.

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.277911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T14:26:06.428076Z digest=sha256:7711bafeffc2c37ae10903f446f657f58675abaf41afbc9ad83ddf7943ce13b4

Observation c42f0dc0-1a24-4b14-b95b-906cda524044 · inbound

Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion cites this paper.

Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T02:24:42.953753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:24:42.953753Z digest=sha256:4653c144d70cef2afd7deeb79851cc8f541590824e72bcb1af894dc39520a3bb