Pith. sign in

Paper Citation Record · LEDGER

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2506.07468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07468 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:07.051052Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:58:57.626705Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0122cf54-546d-493c-bbdb-139728a28dcf · inbound

Learning in Structured Stackelberg Games cites this paper.

Learning in Structured Stackelberg Games Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:47.218133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T21:19:07.958963Z digest=sha256:b6f346d2df3193d7ea899dad8bbf3278269c87e1fa893387f9b5328697f92099

Observation 87fb0fde-85a3-444f-bd59-22107d08e435 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:07.051052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:07.051052Z digest=sha256:14560df11a141208b7b311dff07c67aea463f9c7c4bd01e11b69dae9bca479df

Observation 3ce5c752-b121-4c84-817f-6ad9d19affd0 · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:22.502308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:22.502308Z digest=sha256:63170943f2e31b26dbc767718acd70214c3a1dc18569ff369f7521c096faa564

Observation 5cf49950-f877-4f14-98ac-021460384fb4 · inbound

ProbeLLM: Automating Principled Diagnosis of LLM Failures cites this paper.

ProbeLLM: Automating Principled Diagnosis of LLM Failures Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:07.037472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:07.037472Z digest=sha256:ebe012da5af9fa0c05a73a7e6a84dc5da2c88856d0709a60eb7e4f2709b62108

Observation 49079c01-45c4-4303-b942-dd408376f364 · inbound

Poster: ClawdGo: Endogenous Security Awareness Training for Autonomous AI Agents cites this paper.

Poster: ClawdGo: Endogenous Security Awareness Training for Autonomous AI Agents Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:47.218133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:31:16.838064Z digest=sha256:27f19afbc0bb20ff3632f2a9ea2b8edff0d7d92beaba109dff74e096f6bb1834

Observation 971b4905-a101-4658-8b8c-00ce6d42e8fd · inbound

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment cites this paper.

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:47.218133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T17:24:54.796037Z digest=sha256:5feab7e4f3d25d87c602b22ee0776950e4937b5cccca724805b5fa30f4eb0240

Observation 34ccd0d5-654f-4dd5-aff8-f4f4b8adbc6e · inbound

The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play cites this paper.

The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:47.218133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T00:51:44.093659Z digest=sha256:229a4103bb14f64c7fa2fc43b86b293f57ecaf8426d1f26741848de2ca384bfc

Observation b6e96ab9-4f73-44ad-865a-c8da413936cf · inbound

Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games cites this paper.

Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:47.218133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:03:49.101568Z digest=sha256:5d91e940187c20cdebeda700be5feffb3e91644c1d478966fe63988b7f44040b

Observation f2d594ef-34e0-4153-be32-593cc41954b0 · inbound

Addressing Over-Refusal in LLMs with Competing Rewards cites this paper.

Addressing Over-Refusal in LLMs with Competing Rewards Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:18:47.218133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T06:59:12.695984Z digest=sha256:0e9e9e3b338227b235e2f96f4149f44ecbe7c63b3abd201a3cfb635105e17013