Pith. sign in

Paper Citation Record · LEDGER

RSPO: Regularized Self-Play Alignment of Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2503.00030.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.00030 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.173039Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:56:19.424988Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0fd650d0-0d55-4c60-b957-fb2e363d7bf7 · inbound

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers cites this paper.

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers RSPO: Regularized Self-Play Alignment of Large Language Models

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-07T22:50:33.173039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:50:33.173039Z digest=sha256:5bdebda5e60cc96a5ad7dcd602daaf33ce497890b26d25c7ac9f72e361e009a6

Observation 599bad6f-02f6-4358-84ad-e21cd6c4f3af · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models RSPO: Regularized Self-Play Alignment of Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:58.169779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:58.169779Z digest=sha256:32ea46e2a3452a71fd98c7d326f081170f91d4013cde32399f4dcc5c1991ac12

Observation b73ceb35-25ee-49ab-aab5-cfa2f48dba16 · inbound

GIFT: Games as Informal Training for Generalizable LLMs cites this paper.

GIFT: Games as Informal Training for Generalizable LLMs RSPO: Regularized Self-Play Alignment of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T11:39:36.921385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:39:36.921385Z digest=sha256:012e9d64df7beb4d2b1b090621d9cbf7ad3558e876a7fbe5f4060e06023d54c8

Observation 001f6bc8-e72c-42fb-9fd6-530167ee9e98 · inbound

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning cites this paper.

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning RSPO: Regularized Self-Play Alignment of Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:05.081278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T01:15:12.985803Z digest=sha256:c1db49839531629b953026b7152798018fe05fee52cef814b6a34b7139543e70

Observation 0103bc62-d9e5-4899-a9ed-243b87006504 · inbound

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models cites this paper.

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models RSPO: Regularized Self-Play Alignment of Large Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.350651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:44:53.969301Z digest=sha256:7422df46ed116ef449032efc890c30d8a0dde32cac09609646389615bbf4fd36

Observation fa118cab-4cca-4464-81e4-b8dfa3137d16 · inbound

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization cites this paper.

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization RSPO: Regularized Self-Play Alignment of Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.427569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T14:58:33.167891Z digest=sha256:911d34064e6986e52efefe10d8e97e0aace7dd368ba4a764ac9472a9afd3c613

Observation c6df849d-a624-4467-ab32-2e4767dd5460 · inbound

Meta-Learning Preferences for Multilingual LLM Alignment cites this paper.

Meta-Learning Preferences for Multilingual LLM Alignment RSPO: Regularized Self-Play Alignment of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:36:15.423449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:36:15.423449Z digest=sha256:4f44bf615b4554a917575f521b9e001fba4a8ace0c9ceb1b696a5f293f622d92

Observation 5c182e3e-92af-4e40-8414-c67a61633483 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization RSPO: Regularized Self-Play Alignment of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:58.832884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:58.832884Z digest=sha256:21b46d1da87b797756107077ca5ba8ed0df2b69a9f209651e0e59ae8250d0434