Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2503.00030.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.173039Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T22:56:19.424988Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 0fd650d0-0d55-4c60-b957-fb2e363d7bf7 · inbound
Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers RSPO: Regularized Self-Play Alignment of Large Language Models
Reference 139
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 599bad6f-02f6-4358-84ad-e21cd6c4f3af · inbound
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models RSPO: Regularized Self-Play Alignment of Large Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b73ceb35-25ee-49ab-aab5-cfa2f48dba16 · inbound
GIFT: Games as Informal Training for Generalizable LLMs RSPO: Regularized Self-Play Alignment of Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001f6bc8-e72c-42fb-9fd6-530167ee9e98 · inbound
IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning RSPO: Regularized Self-Play Alignment of Large Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0103bc62-d9e5-4899-a9ed-243b87006504 · inbound
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models RSPO: Regularized Self-Play Alignment of Large Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa118cab-4cca-4464-81e4-b8dfa3137d16 · inbound
S-SPPO: Semantic-Calibrated Self-Play Preference Optimization RSPO: Regularized Self-Play Alignment of Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c6df849d-a624-4467-ab32-2e4767dd5460 · inbound
Meta-Learning Preferences for Multilingual LLM Alignment RSPO: Regularized Self-Play Alignment of Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c182e3e-92af-4e40-8414-c67a61633483 · inbound
Normalized Rewards for Preference Optimization RSPO: Regularized Self-Play Alignment of Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.