Pith. sign in

Paper Citation Record · LEDGER

Safety Alignment of LMs via Non-cooperative Games

As of 15 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2512.20806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.20806 v3

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:24:26.027983Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:59:12.695984Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T07:05:29.201563Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6bf5036e-9e1a-458d-b043-f7f18c0a1599 · outbound

This paper cites write newline.

Safety Alignment of LMs via Non-cooperative Games write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:19.119202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:19.119202Z digest=sha256:ce0da3d3929b9197e2b44bac75f00fec9455f82760adc3031ce782ea5a5aee68

Observation 7ff424ec-665f-44c7-a3a9-d2a0a85ba77a · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Safety Alignment of LMs via Non-cooperative Games Detecting Language Model Attacks with Perplexity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:19.243260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:19.243260Z digest=sha256:5c5cdf6101d8407e607598edbab87ea4e9e9d7eded369c4d3e477276e2c3c9ef

Observation 95f9ed95-80b5-4fba-af4e-df963133336f · outbound

This paper cites Bowman, Ethan Perez, Roger Baker Grosse, and David Duvenaud.

Safety Alignment of LMs via Non-cooperative Games Bowman, Ethan Perez, Roger Baker Grosse, and David Duvenaud

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:19.380861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:19.380861Z digest=sha256:1af3c1aa7e77168662896e6984b98800b48401f6fa2e24e3be179784ff67202b

Observation d058d0b3-5f42-4060-9235-c76e24e0b6db · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

Safety Alignment of LMs via Non-cooperative Games A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:19.502646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:19.502646Z digest=sha256:1463f7d7cbf2febe52e014e44b6eb7b23322fd5fff8cb21d248514185348d665

Observation b8d70e4c-37b8-49f3-815d-20c75aaad14e · outbound

This paper cites fairseq2, 2023.

Safety Alignment of LMs via Non-cooperative Games fairseq2, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:19.645830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:19.645830Z digest=sha256:081890b9832d0f957879b283b0f85266a4856bfdb42d5c81aabc4828a979a037

Observation f33b9f68-23bc-41fa-b925-36bd194a164f · outbound

This paper cites an unresolved cited work.

Safety Alignment of LMs via Non-cooperative Games Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:19.798316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:19.798316Z digest=sha256:69926de7b9696197e48f11374c49b6c1f6d361c5a46e3b71da53214961026fbc

Observation c1755611-ecbb-46bf-9c15-e265afd4f6ca · outbound

This paper cites an unresolved cited work.

Safety Alignment of LMs via Non-cooperative Games Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:19.900634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:19.900634Z digest=sha256:bb3b6324f130ae6d78234c6bc85d6e1a6d6467c1fb565a0706416b109d4feacd

Observation 55e80a70-57f3-4a7f-961b-905b65f616ad · outbound

This paper cites Human alignment of large language models through online preference optimisation.

Safety Alignment of LMs via Non-cooperative Games Human alignment of large language models through online preference optimisation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.116626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.116626Z digest=sha256:38a56b3663ce80567e829fe5ad1440a3d0cceea303b3ec69cd38ed85e2241330

Observation c3ea4715-1646-4fc3-881d-10e6e7007410 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Safety Alignment of LMs via Non-cooperative Games Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.197076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.197076Z digest=sha256:79cc5ec5fd7c5658d21464834b816d2a63ecbffdd85de941518c1b4ddccefc43

Observation 5bb1361d-90dc-4864-bbf5-6d5116973cf1 · outbound

This paper cites Meta secalign: A secure foundation llm against prompt injection attacks, 2025.

Safety Alignment of LMs via Non-cooperative Games Meta secalign: A secure foundation llm against prompt injection attacks, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.273057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.273057Z digest=sha256:6ff4cb3d48b18fdb724a1c801526aa6607a6bf4e2e5d12c8c04a62651aa82f68

Observation b8e51d0a-c93e-4230-8877-7335e1f347d3 · outbound

This paper cites Self-Improving Robust Preference Optimization.

Safety Alignment of LMs via Non-cooperative Games Self-Improving Robust Preference Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.383932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.383932Z digest=sha256:481af1f4f060d90ad3b8cfb9bdb22f5829ef3f662d29d23f5443571dec94a6d8

Observation 7d56223c-15ec-4bda-9d16-5631befa0c31 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Safety Alignment of LMs via Non-cooperative Games Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.525152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.525152Z digest=sha256:da3883d67b72f8e38b441e49f13cbdd3aea3c72863ddf7f00598323cf1d7576c

Observation af809779-a57c-4792-a76c-53d8f178c1ca · outbound

This paper cites AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.

Safety Alignment of LMs via Non-cooperative Games AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.592814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.592814Z digest=sha256:a96000b8b4a2879cf20a3dcf33fc5038f775c47bb9502a68f2ac6e1a73da0d44

Observation 332f2f8f-0d1f-4e1f-a2b5-edb9e47b25d5 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Safety Alignment of LMs via Non-cooperative Games Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.677213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.677213Z digest=sha256:e9dd300781afbcc88c95763365187302d03cec62c8cdc1ab0a623dae3362ff1e

Observation e54ba365-56ae-4497-8569-70873de80b7d · outbound

This paper cites WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks.

Safety Alignment of LMs via Non-cooperative Games WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.798535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.798535Z digest=sha256:8a94cc8dde8d902f8cd5a1f1ffb16e00972f3919a52f51e4ea1913971412c442

Observation 58037a65-8d50-481a-8f54-57238aecf6dd · outbound

This paper cites Value-Free Policy Optimization via Reward Partitioning.

Safety Alignment of LMs via Non-cooperative Games Value-Free Policy Optimization via Reward Partitioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.906356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.906356Z digest=sha256:9999883c604fa308f12da04546e632144a38c90511e4d87d3883dd4581d0e85d

Observation 08b1e2a9-25b6-4b9d-b802-b1e902060ee4 · outbound

This paper cites Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022.

Safety Alignment of LMs via Non-cooperative Games Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:20.980576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:20.980576Z digest=sha256:c1005e2efa27af5a9dbaa9fd0cc21e6458078cddf2e525da974299bac2704107

Observation f00abcb0-1e04-4554-b8e5-8bc5db70f088 · outbound

This paper cites Gradient-based adversarial attacks against text transformers.

Safety Alignment of LMs via Non-cooperative Games Gradient-based adversarial attacks against text transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.048650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.048650Z digest=sha256:62c18445de484ac0e643bb64cb25f6b494c9bd1a90229ba426079b232613fd45

Observation 50f7b995-cb23-4176-9e28-933b1a7e0e1e · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Safety Alignment of LMs via Non-cooperative Games Direct Language Model Alignment from Online AI Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.154235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.154235Z digest=sha256:8d93959856820ad79bfb932a017ce4426a67ac232545b446770aecbbfedb26ad

Observation 92013bd9-5452-4c18-a068-03be9b71889c · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

Safety Alignment of LMs via Non-cooperative Games WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.206299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.206299Z digest=sha256:deff03b48ea79b881bc6a93c56b8935adc9f424ddfca3d5877b7541880ff28ad

Observation fc5fa789-af12-4d5e-b77c-be604ffc1a5d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Safety Alignment of LMs via Non-cooperative Games Measuring Massive Multitask Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.308997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.308997Z digest=sha256:829fb9453171e30ccbd2eae3284f6b45e88e8f5758c2f2259ec2962fe2b643b0

Observation fa270477-3fe6-43ad-9f0a-efc5bdcd085e · outbound

This paper cites WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models.

Safety Alignment of LMs via Non-cooperative Games WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.473425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.473425Z digest=sha256:77a335c60844f4e7250906311299f3b5472dde5a80eb81b8c010b84df532c36c

Observation 202772d5-4b27-4529-a7b7-41fb2b972c45 · outbound

This paper cites Bridging Offline and Online Reinforcement Learning for LLMs.

Safety Alignment of LMs via Non-cooperative Games Bridging Offline and Online Reinforcement Learning for LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.679930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.679930Z digest=sha256:3d7be7a08e2580074ec6068e5a2c98990a39672a3d7442e389744eb5f9c022b9

Observation 050f4c5c-0847-4f91-a0ba-2f68a0cbc320 · outbound

This paper cites JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs.

Safety Alignment of LMs via Non-cooperative Games JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.893602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.893602Z digest=sha256:19081bd5466588548f5defdd5d8931101f845b1759c26f4a7d3900c7c9500f04

Observation 86002d33-4fed-440d-a4e8-47d7e30f9c91 · outbound

This paper cites Gonzalez, and Ion Stoica.

Safety Alignment of LMs via Non-cooperative Games Gonzalez, and Ion Stoica

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:22.126595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:22.126595Z digest=sha256:c933abd3326a039bd05801743357f94da901a04713791f673e715feb0067c37e

Observation 3b27e5fa-d2b2-4718-a468-56ce1e93c334 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Safety Alignment of LMs via Non-cooperative Games TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:22.304412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:22.304412Z digest=sha256:c4b2a8816167f708aa7301157ba3478514afcf920de567b890a7f6c11d146d7c

Observation 3ce5c752-b121-4c84-817f-6ad9d19affd0 · outbound

This paper cites Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models.

Safety Alignment of LMs via Non-cooperative Games Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:22.502308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:22.502308Z digest=sha256:2a22bac8e0d94c60735f3e86b99becfdb69e96383b1db26a099236131bcb7a98

Observation bf4cb70e-cc03-4976-aef8-b663943f47a1 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Safety Alignment of LMs via Non-cooperative Games Understanding R1-Zero-Like Training: A Critical Perspective

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:22.672881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:22.672881Z digest=sha256:d93e668ff75ab80634c556c9ba2943e0a846b1ad15694db748671fc79b3c1873

Observation bd839b38-c599-44f5-be17-1564713659d1 · outbound

This paper cites Evolving Diverse Red-team Language Models in Multi-round Multi-agent Games.

Safety Alignment of LMs via Non-cooperative Games Evolving Diverse Red-team Language Models in Multi-round Multi-agent Games

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:22.866024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:22.866024Z digest=sha256:8f828e2930b26a9662cc550ec9b77736c5f82b1b6775f6cddc15ea3b19de1630

Observation 7034fb6b-576a-4fbe-9836-755f939fb788 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

Safety Alignment of LMs via Non-cooperative Games Towards deep learning models resistant to adversarial attacks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.045786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.045786Z digest=sha256:313ed842cef7be7d581e92f26510287828d36317f3f6fc23c7ee459d06aac36a

Observation 0c38e4bf-ce34-43f7-b9f4-567a9ff16f85 · outbound

This paper cites Harmbench github repository.

Safety Alignment of LMs via Non-cooperative Games Harmbench github repository

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.137102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.137102Z digest=sha256:1e4a5ad7e9c03045d361a97d9eab5ce9e669e8ef0338567aa672ae6ee90678bf

Observation 6e4e4bc8-6001-49e3-bef0-5d101fb6705f · outbound

This paper cites Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024 b.

Safety Alignment of LMs via Non-cooperative Games Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024 b

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.231764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.231764Z digest=sha256:2d334d4ea813ff75d63f4c818ccccb693e85a90a5042c11f442963975b659747

Observation 2a60e2d0-57c5-4174-b728-1da098c41e4e · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

Safety Alignment of LMs via Non-cooperative Games Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.335344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.335344Z digest=sha256:ad92c960ff70a0af7f179bcb2178705f9cd4b2aeb8855bc7e0953979c4e07667

Observation 15e3c58f-85c6-4ec0-8879-505bb1053d42 · outbound

This paper cites The Llama 3 Herd of Models.

Safety Alignment of LMs via Non-cooperative Games The Llama 3 Herd of Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.412067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.412067Z digest=sha256:11e3874e643b88f2311bad8b0872aa23287261a224ca909784d4ce2086abb6ca

Observation 279421bc-c660-4d90-8c4e-9ea37f1a6d0d · outbound

This paper cites Model Card - Prompt Guard.

Safety Alignment of LMs via Non-cooperative Games Model Card - Prompt Guard

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.506201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.506201Z digest=sha256:dcbc269b16f4c407612dd17326747bb9a5bd9083de89f8d6ef2cef2d81683549

Observation e1af94db-53b0-4dd8-9598-0ae13b266a4f · outbound

This paper cites Nash Learning from Human Feedback.

Safety Alignment of LMs via Non-cooperative Games Nash Learning from Human Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.615297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.615297Z digest=sha256:afb147f86407a280280bd58a144424305372b81e53388c2353f689bbc35f26d7

Observation 2effc27a-376d-4393-bc1d-98135915b902 · outbound

This paper cites an unresolved cited work.

Safety Alignment of LMs via Non-cooperative Games Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.698021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.698021Z digest=sha256:1ca73d5a68f04f6500d4fb8d690a7a5b26f95d2dfbb34a87a5cb420af180b5a6

Observation 83d3740f-98b2-4511-8301-36e340544399 · outbound

This paper cites an unresolved cited work.

Safety Alignment of LMs via Non-cooperative Games Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.829322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.829322Z digest=sha256:aceb8f9ec9e2f04f18b6737b6e807068275e13ad02c50dfa820e8af99b30eb7b

Observation 6fa5d27a-5894-4e1d-9bd3-3b94f81e319d · outbound

This paper cites A dv P rompter: Fast adaptive adversarial prompting for LLM s.

Safety Alignment of LMs via Non-cooperative Games A dv P rompter: Fast adaptive adversarial prompting for LLM s

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.902182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.902182Z digest=sha256:360552cf06085f0b137aecd9fcd14ed75aaf4347c60fc719198e43610e62546d

Observation 15d0078a-55f8-4028-b665-4ec624028892 · outbound

This paper cites Automated Red Teaming with GOAT: the Generative Offensive Agent Tester.

Safety Alignment of LMs via Non-cooperative Games Automated Red Teaming with GOAT: the Generative Offensive Agent Tester

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:23.978444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:23.978444Z digest=sha256:aef23f48b2bd0b78d1f648f9a089e80c539d4b588c6a3ec1a282278a1eadf963

Observation 0c0649af-e0e3-43a1-ba0e-b0f32b4b3ccf · outbound

This paper cites Generalizing verifiable instruction following, 2025.

Safety Alignment of LMs via Non-cooperative Games Generalizing verifiable instruction following, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:24.065225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:24.065225Z digest=sha256:7d3d2e02f9c6d5bf58fe81fd5642e9cfabd6141f37e7eff17342872a5c9da04f

Observation 3d31b2b0-8037-4d2a-830c-320ce4d78d93 · outbound

This paper cites Qwen2.5 Technical Report.

Safety Alignment of LMs via Non-cooperative Games Qwen2.5 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:24.145631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:24.145631Z digest=sha256:22dee364611015822015a266526c96340107125025cc9a4c37b7db5f21decaeb

Observation 6555fb52-5a81-4d7c-a31d-47c9eba147db · outbound

This paper cites Strategic Deflection: Defending LLMs from Logit Manipulation.

Safety Alignment of LMs via Non-cooperative Games Strategic Deflection: Defending LLMs from Logit Manipulation

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-03T14:29:07.776313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-03T14:24:24.267269Z digest=sha256:efa7504166d0e66ee49059a9a4fe72cc9c776628d3d4395097e910e19632d9bc

Observation 39484ea9-6726-4041-a27d-55a23646f0d6 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Safety Alignment of LMs via Non-cooperative Games Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:24.358181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:24.358181Z digest=sha256:531cc2436b514370576e3a75b929d0bf999989676b2a28141d7398a5b3264fae

Observation 63e63eba-5b2e-4105-a100-7d8781f04cdb · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Safety Alignment of LMs via Non-cooperative Games XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:24.412524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:24.412524Z digest=sha256:3301004c4e24c5bdaa8c525ec1df62852916450a8234a5883d082ecbbce21654

Observation 09d4a3a3-9da1-468d-b49b-36244ac89e6e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Safety Alignment of LMs via Non-cooperative Games DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:24.506084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:24.506084Z digest=sha256:88bfa1e66d62c6e24f3beee59916d24a4d224d6edccfb5bed6f62098c621f6e0

Observation 4860f30f-d90c-46c3-8088-98ae623440ee · outbound

This paper cites ``Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

Safety Alignment of LMs via Non-cooperative Games ``Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:24.669188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:24.669188Z digest=sha256:4b9510cffe67f436a2fe55b0d78a442fd508656b26c73b8a59880f18d4a177fb

Observation 2e6e7293-f947-406f-81aa-c91cd60daeb6 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

Safety Alignment of LMs via Non-cooperative Games Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:24.762257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:24.762257Z digest=sha256:7c34cd843cb2d6c5ef269ce19bacdc82d4a6cad1e1c47ef018dfbaf646f093d4

Observation 517f4396-9557-4299-904e-761cabb57889 · outbound

This paper cites On general minimax theorems.

Safety Alignment of LMs via Non-cooperative Games On general minimax theorems

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:24.897317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:24.897317Z digest=sha256:1e31ae9eab65b69c57931e5f18586b6192603ec89a0fb5ead4bdfe97733bb86f

Observation 3f983cde-fb04-4c61-8090-0c03e82e9597 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Safety Alignment of LMs via Non-cooperative Games Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:25.064399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:25.064399Z digest=sha256:6ba3ca2d47900cc677d49e8cfadfdfc8b622286a92e37ba3032b5353066243f7

Observation e8a4db2b-d7e7-4e60-9781-7209a11ba1b1 · outbound

This paper cites Rl is a hammer and llms are nails: A simple reinforcement learning recipe for strong prompt injection, 2025.

Safety Alignment of LMs via Non-cooperative Games Rl is a hammer and llms are nails: A simple reinforcement learning recipe for strong prompt injection, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:25.253356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:25.253356Z digest=sha256:627d279ec7114041f0007e2e9693ce5ffc31d62b8bf4a1db18a057b6f6d8e0de

Observation d17e2fce-c79d-4104-9613-1feaf1273f8d · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Safety Alignment of LMs via Non-cooperative Games Self-Play Preference Optimization for Language Model Alignment

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:25.355700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:25.355700Z digest=sha256:ebc6543073ac5a2bfc650784fbf2424692d64195e7bafcc7b5cf4b8cba5b15db

Observation 88d877aa-a2ce-4090-bc42-e5d16d31ad17 · outbound

This paper cites The Alignment Waltz: Jointly Training Agents to Collaborate for Safety.

Safety Alignment of LMs via Non-cooperative Games The Alignment Waltz: Jointly Training Agents to Collaborate for Safety

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:25.516726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:25.516726Z digest=sha256:6b0225a44f15973a85cad4356095f601b89bf9039b24c1012d4778826ca78083

Observation 644bbec5-543a-4555-819a-4c4ca0d16576 · outbound

This paper cites Improving LLM General Preference Alignment via Optimistic Online Mirror Descent.

Safety Alignment of LMs via Non-cooperative Games Improving LLM General Preference Alignment via Optimistic Online Mirror Descent

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:25.789502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:25.789502Z digest=sha256:3b304e6da873ff45ec9d3a33a32608d9ed7d4e496435901ecc3772af3476331a

Observation 7ac1433e-488e-4844-b5b9-cc8d9c67cb84 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Safety Alignment of LMs via Non-cooperative Games Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:25.902197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:25.902197Z digest=sha256:c153c25504e48232c54e323c767a96ca12f6f4f6550103eab8d70b2986183317

Observation c686451d-1ddf-4bdf-8e63-ff72c94314b6 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Safety Alignment of LMs via Non-cooperative Games Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:26.027983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:26.027983Z digest=sha256:32caaaf1fb5751b3dac47294a55e27553bfa16bf97533dd510e0d82d921ba929

Pith citing papers

Observation e54d6df7-02b4-41e3-8926-b52b60a0c097 · inbound

Min-Max Optimization Requires Exponentially Many Queries cites this paper.

Min-Max Optimization Requires Exponentially Many Queries Safety Alignment of LMs via Non-cooperative Games

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-02T03:04:03.431172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T17:32:35.323096Z digest=sha256:5c21404782d8fd9e9172089eb74e162b9c4699a2e8104650d695ba25f9ddc064

Observation 4899a69f-4039-4ebf-ad0a-d966705a4b98 · inbound

Addressing Over-Refusal in LLMs with Competing Rewards cites this paper.

Addressing Over-Refusal in LLMs with Competing Rewards Safety Alignment of LMs via Non-cooperative Games

Reference 102

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:05:29.203499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-01T06:59:12.695984Z digest=sha256:70857d485be3edacbcfada7831482c27a8d1fdc2787df281083bb780f1dd94e9