Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T14:24:26.027983Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2512.20806.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T14:24:26.027983Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T06:59:12.695984Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-01T07:05:29.201563Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6bf5036e-9e1a-458d-b043-f7f18c0a1599 · outbound
Safety Alignment of LMs via Non-cooperative Games write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ff424ec-665f-44c7-a3a9-d2a0a85ba77a · outbound
Safety Alignment of LMs via Non-cooperative Games Detecting Language Model Attacks with Perplexity
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95f9ed95-80b5-4fba-af4e-df963133336f · outbound
Safety Alignment of LMs via Non-cooperative Games Bowman, Ethan Perez, Roger Baker Grosse, and David Duvenaud
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d058d0b3-5f42-4060-9235-c76e24e0b6db · outbound
Safety Alignment of LMs via Non-cooperative Games A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d70e4c-37b8-49f3-815d-20c75aaad14e · outbound
Safety Alignment of LMs via Non-cooperative Games fairseq2, 2023
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f33b9f68-23bc-41fa-b925-36bd194a164f · outbound
Safety Alignment of LMs via Non-cooperative Games Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1755611-ecbb-46bf-9c15-e265afd4f6ca · outbound
Safety Alignment of LMs via Non-cooperative Games Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55e80a70-57f3-4a7f-961b-905b65f616ad · outbound
Safety Alignment of LMs via Non-cooperative Games Human alignment of large language models through online preference optimisation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3ea4715-1646-4fc3-881d-10e6e7007410 · outbound
Safety Alignment of LMs via Non-cooperative Games Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb1361d-90dc-4864-bbf5-6d5116973cf1 · outbound
Safety Alignment of LMs via Non-cooperative Games Meta secalign: A secure foundation llm against prompt injection attacks, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e51d0a-c93e-4230-8877-7335e1f347d3 · outbound
Safety Alignment of LMs via Non-cooperative Games Self-Improving Robust Preference Optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d56223c-15ec-4bda-9d16-5631befa0c31 · outbound
Safety Alignment of LMs via Non-cooperative Games Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af809779-a57c-4792-a76c-53d8f178c1ca · outbound
Safety Alignment of LMs via Non-cooperative Games AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 332f2f8f-0d1f-4e1f-a2b5-edb9e47b25d5 · outbound
Safety Alignment of LMs via Non-cooperative Games Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54ba365-56ae-4497-8569-70873de80b7d · outbound
Safety Alignment of LMs via Non-cooperative Games WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58037a65-8d50-481a-8f54-57238aecf6dd · outbound
Safety Alignment of LMs via Non-cooperative Games Value-Free Policy Optimization via Reward Partitioning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08b1e2a9-25b6-4b9d-b802-b1e902060ee4 · outbound
Safety Alignment of LMs via Non-cooperative Games Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00abcb0-1e04-4554-b8e5-8bc5db70f088 · outbound
Safety Alignment of LMs via Non-cooperative Games Gradient-based adversarial attacks against text transformers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f7b995-cb23-4176-9e28-933b1a7e0e1e · outbound
Safety Alignment of LMs via Non-cooperative Games Direct Language Model Alignment from Online AI Feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92013bd9-5452-4c18-a068-03be9b71889c · outbound
Safety Alignment of LMs via Non-cooperative Games WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc5fa789-af12-4d5e-b77c-be604ffc1a5d · outbound
Safety Alignment of LMs via Non-cooperative Games Measuring Massive Multitask Language Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa270477-3fe6-43ad-9f0a-efc5bdcd085e · outbound
Safety Alignment of LMs via Non-cooperative Games WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202772d5-4b27-4529-a7b7-41fb2b972c45 · outbound
Safety Alignment of LMs via Non-cooperative Games Bridging Offline and Online Reinforcement Learning for LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 050f4c5c-0847-4f91-a0ba-2f68a0cbc320 · outbound
Safety Alignment of LMs via Non-cooperative Games JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86002d33-4fed-440d-a4e8-47d7e30f9c91 · outbound
Safety Alignment of LMs via Non-cooperative Games Gonzalez, and Ion Stoica
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b27e5fa-d2b2-4718-a468-56ce1e93c334 · outbound
Safety Alignment of LMs via Non-cooperative Games TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ce5c752-b121-4c84-817f-6ad9d19affd0 · outbound
Safety Alignment of LMs via Non-cooperative Games Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf4cb70e-cc03-4976-aef8-b663943f47a1 · outbound
Safety Alignment of LMs via Non-cooperative Games Understanding R1-Zero-Like Training: A Critical Perspective
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd839b38-c599-44f5-be17-1564713659d1 · outbound
Safety Alignment of LMs via Non-cooperative Games Evolving Diverse Red-team Language Models in Multi-round Multi-agent Games
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7034fb6b-576a-4fbe-9836-755f939fb788 · outbound
Safety Alignment of LMs via Non-cooperative Games Towards deep learning models resistant to adversarial attacks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c38e4bf-ce34-43f7-b9f4-567a9ff16f85 · outbound
Safety Alignment of LMs via Non-cooperative Games Harmbench github repository
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e4e4bc8-6001-49e3-bef0-5d101fb6705f · outbound
Safety Alignment of LMs via Non-cooperative Games Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024 b
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a60e2d0-57c5-4174-b728-1da098c41e4e · outbound
Safety Alignment of LMs via Non-cooperative Games Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e3c58f-85c6-4ec0-8879-505bb1053d42 · outbound
Safety Alignment of LMs via Non-cooperative Games The Llama 3 Herd of Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 279421bc-c660-4d90-8c4e-9ea37f1a6d0d · outbound
Safety Alignment of LMs via Non-cooperative Games Model Card - Prompt Guard
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1af94db-53b0-4dd8-9598-0ae13b266a4f · outbound
Safety Alignment of LMs via Non-cooperative Games Nash Learning from Human Feedback
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2effc27a-376d-4393-bc1d-98135915b902 · outbound
Safety Alignment of LMs via Non-cooperative Games Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83d3740f-98b2-4511-8301-36e340544399 · outbound
Safety Alignment of LMs via Non-cooperative Games Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fa5d27a-5894-4e1d-9bd3-3b94f81e319d · outbound
Safety Alignment of LMs via Non-cooperative Games A dv P rompter: Fast adaptive adversarial prompting for LLM s
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d0078a-55f8-4028-b665-4ec624028892 · outbound
Safety Alignment of LMs via Non-cooperative Games Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c0649af-e0e3-43a1-ba0e-b0f32b4b3ccf · outbound
Safety Alignment of LMs via Non-cooperative Games Generalizing verifiable instruction following, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d31b2b0-8037-4d2a-830c-320ce4d78d93 · outbound
Safety Alignment of LMs via Non-cooperative Games Qwen2.5 Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6555fb52-5a81-4d7c-a31d-47c9eba147db · outbound
Safety Alignment of LMs via Non-cooperative Games Strategic Deflection: Defending LLMs from Logit Manipulation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 39484ea9-6726-4041-a27d-55a23646f0d6 · outbound
Safety Alignment of LMs via Non-cooperative Games Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e63eba-5b2e-4105-a100-7d8781f04cdb · outbound
Safety Alignment of LMs via Non-cooperative Games XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d4a3a3-9da1-468d-b49b-36244ac89e6e · outbound
Safety Alignment of LMs via Non-cooperative Games DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4860f30f-d90c-46c3-8088-98ae623440ee · outbound
Safety Alignment of LMs via Non-cooperative Games ``Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e6e7293-f947-406f-81aa-c91cd60daeb6 · outbound
Safety Alignment of LMs via Non-cooperative Games Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517f4396-9557-4299-904e-761cabb57889 · outbound
Safety Alignment of LMs via Non-cooperative Games On general minimax theorems
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f983cde-fb04-4c61-8090-0c03e82e9597 · outbound
Safety Alignment of LMs via Non-cooperative Games Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8a4db2b-d7e7-4e60-9781-7209a11ba1b1 · outbound
Safety Alignment of LMs via Non-cooperative Games Rl is a hammer and llms are nails: A simple reinforcement learning recipe for strong prompt injection, 2025
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d17e2fce-c79d-4104-9613-1feaf1273f8d · outbound
Safety Alignment of LMs via Non-cooperative Games Self-Play Preference Optimization for Language Model Alignment
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d877aa-a2ce-4090-bc42-e5d16d31ad17 · outbound
Safety Alignment of LMs via Non-cooperative Games The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 644bbec5-543a-4555-819a-4c4ca0d16576 · outbound
Safety Alignment of LMs via Non-cooperative Games Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac1433e-488e-4844-b5b9-cc8d9c67cb84 · outbound
Safety Alignment of LMs via Non-cooperative Games Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c686451d-1ddf-4bdf-8e63-ff72c94314b6 · outbound
Safety Alignment of LMs via Non-cooperative Games Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54d6df7-02b4-41e3-8926-b52b60a0c097 · inbound
Min-Max Optimization Requires Exponentially Many Queries Safety Alignment of LMs via Non-cooperative Games
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4899a69f-4039-4ebf-ad0a-d966705a4b98 · inbound
Addressing Over-Refusal in LLMs with Competing Rewards Safety Alignment of LMs via Non-cooperative Games
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.