Pith. sign in

Paper Citation Record · LEDGER

Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2502.16852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.16852 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:24:25.789502Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T19:32:01.290815Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1bc552a7-eb6f-4ba2-be9c-c225b57e6375 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.293338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:dcc7133a94a3da8e893049a51e7f8d18328eb7ecb7a8c9d7f5a28f8adc54d184

Observation d8df3fe8-36b7-4734-82a2-68f7d0244696 · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.022728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:8c4444dc55dccaece7eb4bd6b6c40dccb3a09db9e583da3aa9078062948f2f31

Observation 644bbec5-543a-4555-819a-4c4ca0d16576 · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:25.789502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:25.789502Z digest=sha256:3514f2b61867b5d0d35111c5cd282373b66bff5c98b99370091e5def2abffbe7

Observation 564865ed-5513-407d-87d0-01759a8b9402 · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:25.670135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:25.670135Z digest=sha256:76ddfa9e5c6a4611cd186b734b29277c70cc0c6cf9fd1ce3098839696e5b1e79

Observation 7a2fc9aa-b845-4e20-82e8-52e78af90a38 · inbound

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium cites this paper.

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:06.812051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:54:58.732444Z digest=sha256:f8e287d0839ee411f3e0b43dadb606f4073e643cacc006f8b640f4d478bf6413

Observation b3a700e2-9e5a-403b-8dac-0574b3b2c91d · inbound

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective cites this paper.

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:55.154910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T01:19:58.448344Z digest=sha256:11ae3d6e286a1b08ef499d1194274a8adf111a574f62352632059977e0b43f21

Observation a50f61f9-77eb-403f-b166-1f1aa1d90a84 · inbound

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment cites this paper.

Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.277911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T14:26:06.428076Z digest=sha256:bb7d3713c39f2aad9b84064148f2694e0ff2119722bc97545a50241d4f89ff2c

Observation c42f0dc0-1a24-4b14-b95b-906cda524044 · inbound

Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion cites this paper.

Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T02:24:42.953753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:24:42.953753Z digest=sha256:a48079b3a2720f23fac6cf79d450e6c254fd7ecea1a61194ce213016540dbf23