Pith. sign in

Paper Citation Record · LEDGER

Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2411.00062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.00062 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:00:09.274466Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7bce9548-c0fb-4385-a5b5-0fbd53fc1461 · inbound

Lifelong Safety Alignment for Language Models cites this paper.

Lifelong Safety Alignment for Language Models Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.274466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.274466Z digest=sha256:b744dae2c77d5347998e53a41508b3f6a3eea970650dbe22631fe261cae05b33

Observation d49a9bb5-c74c-41b5-b597-caccd9fa8c15 · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:58.219409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:58.219409Z digest=sha256:13fcada353b366bd7226df949af880cbcb2d8357bffbf938cd0ab13baf3d2636

Observation 294a7bac-dff0-4969-b010-1ade01a7fecc · inbound

Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability cites this paper.

Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T07:59:10.131945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:59:10.131945Z digest=sha256:520ab2dbd9f276a4cd9a6a670ce191d8c9350bd25f85561f4d3ed94bdaccdea6

Observation 6ea7dfb6-a5bb-4ab1-b5fc-eb9b51d96dc9 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:03:28.955656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:d677076f0b746ae5234e3cc99fd91e6f17e3cb9e0a442c6c8c3a6a88fd793d32

Observation 00b19379-6e98-4624-ad30-e618639a9bb2 · inbound

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play cites this paper.

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:42:48.116666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T21:37:56.570173Z digest=sha256:7f483973c09ab85b050aaa4ca573677748d8f1dfbdda7291f8d64b162ee240be

Observation 7ff97028-1fd7-4a5f-a61d-4d5f6b572086 · inbound

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL cites this paper.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T19:23:10.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:23:10.328625Z digest=sha256:a555f11e53cc72374f6578c13c2d0ee0e29e31b448325febddcfab1be5e94a06