Pith. sign in

Paper Citation Record · LEDGER

GPT-Red: Automated Red Teaming via Self-Play at Scale

As of 9 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2607.26115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26115 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T01:12:44.201592Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7be87c14-6d5f-4a38-83dc-6b1456f5942a · outbound

This paper cites AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.

GPT-Red: Automated Red Teaming via Self-Play at Scale AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.130108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.130108Z digest=sha256:e7da59e841e00f91eee74eb62b3822eeb998cde4c224cb5850cd7ca53d57b639

Observation e14ae957-196f-4a65-a724-239d6e7160d5 · outbound

This paper cites How vulnerable are ai agents to indirect prompt injections? insights from a large‐ scalepubliccompetition.

GPT-Red: Automated Red Teaming via Self-Play at Scale How vulnerable are ai agents to indirect prompt injections? insights from a large‐ scalepubliccompetition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.135153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.135153Z digest=sha256:f11756addf14c3c06d6447f0480aea3d6a3a8a3661fa32a2e3f1107e822a54f4

Observation 64677a4b-47eb-40e2-af4b-f333ce154e4c · outbound

This paper cites MART: Improving LLM Safety with Multi-round Automatic Red-Teaming.

GPT-Red: Automated Red Teaming via Self-Play at Scale MART: Improving LLM Safety with Multi-round Automatic Red-Teaming

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.140186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.140186Z digest=sha256:ee7e88ba3a918e0c33efb8b0af778b9d9b8a16c61082116fbb6d45f84688edd3

Observation 6ad47015-6eeb-423d-8168-b57de8d569cb · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

GPT-Red: Automated Red Teaming via Self-Play at Scale Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.145869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.145869Z digest=sha256:e0e156d3d96909bd5f83fc4866307bf2c1ff1d41120bec247b29e5f930a18803

Observation 04b3dc29-8c72-48d4-bcbf-e2af7d5700b8 · outbound

This paper cites Red Teaming Language Models with Language Models.

GPT-Red: Automated Red Teaming via Self-Play at Scale Red Teaming Language Models with Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.151768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.151768Z digest=sha256:99eac7a597c9d539e0f4507c25e2cd32a64431c7d9a6bd797dacb182c62ca045

Observation bd147938-f0fd-44a5-bc2b-366b870efa8e · outbound

This paper cites Toyer, S., Watkins, O., Mendes, E.

GPT-Red: Automated Red Teaming via Self-Play at Scale Toyer, S., Watkins, O., Mendes, E

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.164253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.164253Z digest=sha256:209dcf2f0448492877cdb25af0fdb7cb8f5ab182b3d64ef176a439c0154838f0

Observation c76f5e9b-e5d3-4311-b3ff-1e806e38fac5 · outbound

This paper cites AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents.

GPT-Red: Automated Red Teaming via Self-Play at Scale AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.169745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.169745Z digest=sha256:f9815d400c661f2b40d182a567860726aa6d3a2e3f52b3749975e0287919a775

Observation 717e47b6-1b6c-41c6-9307-4c7c5dc2ca33 · outbound

This paper cites Xu, H., Zhang, W., Wang, Z., Xiao, F., Zheng, R., Feng, Y., Ba, Z., and Ren, K.

GPT-Red: Automated Red Teaming via Self-Play at Scale Xu, H., Zhang, W., Wang, Z., Xiao, F., Zheng, R., Feng, Y., Ba, Z., and Ren, K

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.174390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.174390Z digest=sha256:8e6b2f3c2efa9c0666fcf1fc01a31105dab3d469520c012c7c780ae16ee08a59

Observation b3b09427-a1e8-4a06-a09e-bb0a28d19412 · outbound

This paper cites From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training.

GPT-Red: Automated Red Teaming via Self-Play at Scale From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.184834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.184834Z digest=sha256:909e0a455ecd560b52fe2be7cb5206e9d9189fa6210d999b6a469e4018ce5e61

Observation 2c4aff08-d08a-456a-9bd9-6ea58d236fa6 · outbound

This paper cites InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents.

GPT-Red: Automated Red Teaming via Self-Play at Scale InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.189784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.189784Z digest=sha256:b9b4b581a809a7a18b7d4eab3d4c64f41dc053de8351b55883fa6500c134d262

Observation 03156622-021f-43a8-884d-63311c262b19 · outbound

This paper cites Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents.

GPT-Red: Automated Red Teaming via Self-Play at Scale Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.195947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.195947Z digest=sha256:857d00faaeae3827f206265398fb297f77553dbc021d009ec121c2ed2f992f87

Observation c442ccb6-1671-4e00-a6cb-4a169654a8a0 · outbound

This paper cites Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition.

GPT-Red: Automated Red Teaming via Self-Play at Scale Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.201592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.201592Z digest=sha256:d8ad149e42bd96eb5908b39b21da3585c73fe87a08ecb8a80553f40b92d97a61

Observation 8c0e98fb-81ec-46f8-8f73-c93b7bb478cb · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

GPT-Red: Automated Red Teaming via Self-Play at Scale Constitutional AI: Harmlessness from AI Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.113665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.113665Z digest=sha256:4e8f1e9d31758b3bfaea5f1547ff26ea2d77d2d930922f982af13b09d246fbea

Observation ed6572bb-4fb9-4054-b5d4-bdafaec6e790 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

GPT-Red: Automated Red Teaming via Self-Play at Scale Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.124317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.124317Z digest=sha256:9a179a8bb9d4f91a89deeba3336582145c181b7b10ff17ae2de0a6d0d84681a0

Observation e9a38d2f-ccfa-47e0-bf9f-04bd2014c4c3 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

GPT-Red: Automated Red Teaming via Self-Play at Scale AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.107279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.107279Z digest=sha256:dbac26114562b6b0f5de6d5f3d36fa22e5a264f1a66e539422ed2b8dfdc6f056

Observation e8a800fe-adaf-4bc1-b6ff-5bf651c93279 · outbound

This paper cites Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning.

GPT-Red: Automated Red Teaming via Self-Play at Scale Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.118461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.118461Z digest=sha256:8b9a75dd680161af7d823b8d6b6148d14cd67c6814b365aaf323c567285d9b14

Observation 87c63fa4-9d67-41ee-948b-c36bd79adb0b · outbound

This paper cites Lessons from Defending Gemini Against Indirect Prompt Injections.

GPT-Red: Automated Red Teaming via Self-Play at Scale Lessons from Defending Gemini Against Indirect Prompt Injections

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.157635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.157635Z digest=sha256:3edbd740da4e61efd9d98e6a950fea829627cc8b15e74e8ce5bbaef419f1f468

Observation 1582961c-4a0f-43f7-ac87-96b5e90e24af · outbound

This paper cites Prompt Injection as Role Confusion.

GPT-Red: Automated Red Teaming via Self-Play at Scale Prompt Injection as Role Confusion

Reference 6667

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.179664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.179664Z digest=sha256:2d4a9188d5a2eb774f1759d13877a468719f95076753dea73fd0ad216ba5701d

Pith citing papers

No inbound Pith citation observations are available.