Pith. sign in

Paper Citation Record · LEDGER

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

As of 16 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2506.14682.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14682 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:53:33.486755Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:25:15.771731Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T23:51:45.015967Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved25
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5786a459-44e7-455b-ae16-95ad0238f8f1 · outbound

This paper cites IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.591918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.591918Z digest=sha256:ecd86b71900b77a01baf375b1ac9f268374bd3deb8105628a568b8df579e02cf

Observation d6d8958e-1ec8-4462-9ae4-d15fa45b3c4b · outbound

This paper cites LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.622526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.622526Z digest=sha256:5dd4bf2ef19540712cc443bc62edc93a7698a7add2fd0f111bf7ae5cdfce6f8c

Observation a2bdf34d-1c59-44e8-ba29-1d6dd7929034 · outbound

This paper cites An Empirical Evaluation of LLMs for Solving Offensive Security Challenges.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.627164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.627164Z digest=sha256:0bc4d240fb00be417661f2dbbfdf676efe76d05411a159c0548676606fef90a7

Observation 8cdf4899-13d1-4341-9b92-d3897291a8c3 · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLM Agents can Autonomously Hack Websites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.631740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.631740Z digest=sha256:d6d0d7c2ed73e5ea8ce1e2c523621f2d1d4c4feb2a350296306bc6ef2e533e32

Observation b8956961-49fd-4354-b263-91394b2e0055 · outbound

This paper cites Llm4decompile: Decompilingbinarycodewithlargelanguagemodels,.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Llm4decompile: Decompilingbinarycodewithlargelanguagemodels,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.298021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:32.637031Z digest=sha256:c193523c7f092eb87734fcc829b593fe3e4bbd42d01d29acd495e7f1802b762a

Observation f6cc0e37-7723-497a-a17c-3d861bd3a300 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.644504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.644504Z digest=sha256:c0d37c13b263dd6cc4a340f5b7c8ced5be43a96c93b6acff183c081f257f5fa8

Observation b17f629e-780e-4ca8-b4e4-0430f937be60 · outbound

This paper cites ATLAS - Adversarial Threat Landscape for Artificial-Intelligence Systems.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models ATLAS - Adversarial Threat Landscape for Artificial-Intelligence Systems

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.287871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:32.674375Z digest=sha256:aa3d817f392ed2415d76661daf24aeb9b82bc7eaa0d64c2a008ca44e2ef85a22

Observation a10781cf-654c-441a-ac0e-f37410e03917 · outbound

This paper cites OWASP Top Ten for Large Language Model Applications.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OWASP Top Ten for Large Language Model Applications

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.276513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:32.679927Z digest=sha256:1277ff180b097d565de169a8acbabb9db30c283ec05e3b1fb7299c1f30d0d481

Observation cd5fef5e-240c-472a-953b-0c62d239a656 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Measuring Massive Multitask Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.684251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.684251Z digest=sha256:f2ebc3d75cba9a1e46b2b931e0dba36291e01b97228f81a0cffcc979d4ed9ff5

Observation ee12daaa-f90a-435c-a0a1-70b9f18a4511 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.689085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.689085Z digest=sha256:336c84c3ce8e0cb7c7d9d936970eb5fd3063114ec06fc18f5390bab23ee9be90

Observation 786fffbf-ad22-4c29-9c6b-fc7f941845ed · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.693437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.693437Z digest=sha256:1e5f8ca9e8a3b6f4fcd0dd8a3607cf62c9d8b73fdc3522cf9154e1eb9b4c5e8c

Observation 1aec31e9-929e-4d0b-91d4-37a68ad00980 · outbound

This paper cites OpenAI o1 System Card.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OpenAI o1 System Card

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.262382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:32.697337Z digest=sha256:7bdd58e74d8f79cda842f8ac7c0ba7c53ccb6342f801597eccf33d111f704dae

Observation 648e0d2c-4834-4711-b100-5d48e8b50c1e · outbound

This paper cites Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.701223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.701223Z digest=sha256:a32507a5d2c1566fb2476ab2e9b1ef07ccb49e25010c439c973498398d0ed16d

Observation 37b18dbc-7422-4518-9dc3-0f954283512b · outbound

This paper cites Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.705945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.705945Z digest=sha256:f536ac61bd1de35f03311290b448f3c094e87e3d9784582e7f275b26a861b837

Observation 415e0e21-ec74-46e4-b7a1-af7ae258fcce · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.725506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.725506Z digest=sha256:ed281a8c52de3256810809366d39e616a270d783bc444cb4c091e07be7e9bde5

Observation 08056cae-2d1e-4ed9-962f-f7d42c40eb71 · outbound

This paper cites Jimenez, John Yang, Kai Liu, and Aleksander Madry.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Jimenez, John Yang, Kai Liu, and Aleksander Madry

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.249705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:32.756599Z digest=sha256:502181377a543a5a6528add8f033f758eb5aa423392e5e0600ed47a4ec282b5c

Observation 113acd9c-9a00-4443-9b82-41d4088c02d2 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.783357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.783357Z digest=sha256:f2e2b7039b048bbe0174d28229b87bea54d5308a5cf196e7d89b9631d8998c51

Observation a39166fe-2e49-45b5-96c4-593488160ad8 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models AgentBench: Evaluating LLMs as Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.791730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.791730Z digest=sha256:d4de3db68b67d76e90a6ea5240dfce623c0af3639702e48eb85e61fac2e45bd5

Observation 1ac4ff36-5be3-4839-8875-091815b144b3 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.795634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.795634Z digest=sha256:72dff3314f85ee58a3226e1aceafcbdf06772be695928632b29420dbcc1f542f

Observation b8383111-e426-45e0-b10a-14028c21132d · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Mind2Web: Towards a Generalist Agent for the Web

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.803200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.803200Z digest=sha256:8beb5778a6c6508f73dfe6ad7a3429f518f2d7e98f3b58da99a72fe1296f968a

Observation 1608e2a3-8a37-42b2-9663-684db5f7e669 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.807624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.807624Z digest=sha256:867797cbcef5c3acd375f7ab485f6b9626afb4345569876f12ecf2fbbd56cd61

Observation dd6e4679-7a17-4487-8954-34b71f13326a · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:53:34.149152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:32.814010Z digest=sha256:f72fe6031447eb32b71b64e186bd52f9faebf467d83bece57b3d95e259830db9

Observation 0a2b9505-3ecd-448a-985a-de94353d94c7 · outbound

This paper cites InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.840618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.840618Z digest=sha256:95265f18471a481a366e6192c2dfb2e3fe154ed63aeeeca7ad50de980404c50b

Observation ef0eb5b6-ae4b-416e-a836-0b3f445ab6ca · outbound

This paper cites NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.863240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.863240Z digest=sha256:44f55e3d2e4aead15226013e656f77e548cebe3cc9451f81777949bc2aeb16cb

Observation 8cbe8746-024a-480d-8048-9548adcda81f · outbound

This paper cites EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.885270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.885270Z digest=sha256:d32b2b4a82d4c058510de2cbba6e70c05c10db360f6d1f03715e543479fc30ca

Observation 24625837-5955-4a37-a26f-c90c28372f39 · outbound

This paper cites AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.995201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.995201Z digest=sha256:6f1467f30e4b919ed3fb30d0dfd1f097ed5353e42a917c93e24fd17a031ab8fd

Observation 38c7678a-0a84-41de-9196-0ece445f7681 · outbound

This paper cites Picoctf learning.https://www.picoctf.org/, 2024.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Picoctf learning.https://www.picoctf.org/, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.077377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.103438Z digest=sha256:655040bd6dfcf1418406abecaf498a838283a050441cc38f5507659f8938c854

Observation 54ec6960-1b49-4018-b144-1c0b1af26992 · outbound

This paper cites Ai village capture the flag @ defcon31.https://kaggle.com/ competitions/ai-village-capture-the-flag-defcon31, 2023.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Ai village capture the flag @ defcon31.https://kaggle.com/ competitions/ai-village-capture-the-flag-defcon31, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.026552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.122420Z digest=sha256:c69b36351bdd8619c74ccfed1c4e48dbb28816dc3fd2f1bd586ca62ab4db7c75

Observation d877e2e1-e64c-4c89-8521-0132a06329a0 · outbound

This paper cites Jupyter datascience notebook docker image, 2025.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Jupyter datascience notebook docker image, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.016063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.131197Z digest=sha256:3ae9f1082c774a6ce0375bd7e6efb7647d160995497d510b328427b5e25ce4bb

Observation 72825d09-5f60-42c1-8844-985a27434988 · outbound

This paper cites Optimizing Large Language Model Hyperparameters for Code Generation.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Optimizing Large Language Model Hyperparameters for Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:33.154793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:33.154793Z digest=sha256:620d9da536b22ca13e61d2342c6084a24e0c30a9ce785ec4a919d1e713c7129d

Observation dbdd6fc6-cfb3-4bf4-9e40-81635bd17218 · outbound

This paper cites The Automation Advantage in AI Red Teaming.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models The Automation Advantage in AI Red Teaming

Reference 31

Resolution
malformed identifier
no resolver link, observed 2026-08-15T19:53:33.168800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:33.168800Z digest=sha256:066e1d10dc30f48c88a3a345a8bcd6490412d27c2bde5da3ceeb2d768c3cac8c

Observation 4d833146-c9cd-4247-aa55-ccfec3df252c · outbound

This paper cites For example : ‘t = turtle.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models For example : ‘t = turtle

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.981683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.281100Z digest=sha256:51c3e959fcead87874e18d923e8d54f0b436ab6c201be1314a6bb52be3318f91

Observation 69d706f8-82b6-4039-bad0-b471a24bdfc3 · outbound

This paper cites S Y S T E M _ C O M M A N D _ E X E C U T E D _ V I A _ E X E C.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models S Y S T E M _ C O M M A N D _ E X E C U T E D _ V I A _ E X E C

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.970997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.295834Z digest=sha256:3807ff889f76eb4460d32be1a7b548755e5a5af3ddbc6bc200b6cb19a264a3a1

Observation ec659ead-6b3c-40ff-8c8f-0b0fbabd57e1 · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:53:33.959915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.306754Z digest=sha256:d468c9d9e5f7a7c8ac2a5d813cc074d642bf40142915a47de00e136ce6da3948

Observation d19cd96b-6aff-4075-9e4c-415cf88b77b7 · outbound

This paper cites For example : ‘t.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models For example : ‘t

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.949128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.310734Z digest=sha256:aeba1cededd732a94099fd939b9ec22c8a4533dda96cbb2973b938480fdfcc77

Observation 1953971f-4bdf-445d-a1dc-cdb47053335a · outbound

This paper cites " " response = query ( prompt ) print ( response ) </ execute - code > <result idx=0 success=True> ’output’:.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models " " response = query ( prompt ) print ( response ) </ execute - code > <result idx=0 success=True> ’output’:

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.911454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.314088Z digest=sha256:b9b01ab72074789bb93dec9b00e140d9b8fdb541d747e735fed39deadc85db9b

Observation 54452e34-cca5-47bf-84db-6e10008c0787 · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:53:34.004403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.325454Z digest=sha256:c6d54372c79c7227a62fdb282d3aca8dc93e030217ffda2818e7f617f1d2fc3a

Observation 26208294-c958-4518-9989-3308f74bc1ef · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 41

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T19:53:33.992721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.353697Z digest=sha256:fec43710d9f9dcc6f5ea8da707832ae0a1a51a83b140bc08b7aba03615e0279b

Observation a264a8e2-edc5-4510-bf98-42b4c350f91e · outbound

This paper cites Example : ‘t = turtle.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Example : ‘t = turtle

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.840388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.388384Z digest=sha256:a23bd063eee65d1609575159d490efffd1ebd48e9c89bce88fa1f28c34b9063a

Observation f768c676-7a52-4781-8769-fc5bad636043 · outbound

This paper cites echo J A G A _ L L M _ G E N E R A T E D _ O S _ C O M M A N D.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models echo J A G A _ L L M _ G E N E R A T E D _ O S _ C O M M A N D

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.794439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.430804Z digest=sha256:1c729b64a04a525df2bb23cce70820c4a730bf5a011e602adf48f5b24e0533e4

Observation 46ed088a-d71e-42bc-a856-3b11fd256ca3 · outbound

This paper cites system () ‘ call , make the turtle ‘t ‘ draw a square of side 10.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models system () ‘ call , make the turtle ‘t ‘ draw a square of side 10

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.783066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.465619Z digest=sha256:61fe909335e6342f82520d413953208d9232496000985dabc7d65bd324717d88

Observation ef0f6acf-4d21-45a9-ba3c-2e6328803dec · outbound

This paper cites " " response = requests . post ( f.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models " " response = requests . post ( f

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.770156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:53:33.486755Z digest=sha256:c37b51324e34c10536953b6160e9ded879c75d0a19ea6a5d046a1b5ba9875e83

Observation c505c07d-fa71-48f9-a6f4-360ef26a4f7d · outbound

This paper cites LLM4Decompile: Decompiling Binary Code with Large Language Models.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLM4Decompile: Decompiling Binary Code with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.640743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.640743Z digest=sha256:a17bec6437b800cf5f25bc45a485e2c99f75cb2bdb9fddcf57a998954e59a1fa

Pith citing papers

Observation d2a4ce5b-b37f-44f8-b429-c30af3ddcabc · inbound

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security cites this paper.

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:25:15.771731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:25:15.771731Z digest=sha256:c8af7a7c47944ae61b582b045299ce4a26e0428fb84a189015a81da7abe7ac7b

Observation 3dff4f9f-4458-46a4-955f-7c6083e29f80 · inbound

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts cites this paper.

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:09:33.025834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:09:33.025834Z digest=sha256:805ff945262682675ef2f60a0c8fc56fc89f20402c5ca40e1ae53103f7512e5a

Observation 898cf871-8391-4070-931f-3fb691354ef7 · inbound

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours cites this paper.

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:51:45.056581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T16:06:18.057868Z digest=sha256:6022173cc62b9f6d6cdfa8376d6145b41a5738eaac5abe85af604d6155d4cd1b