Pith. sign in

Paper Citation Record · LEDGER

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

As of 17 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2506.14682.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14682 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:53:33.486755Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:25:15.771731Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T23:51:45.015967Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved25
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5786a459-44e7-455b-ae16-95ad0238f8f1 · outbound

This paper cites IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.591918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.591918Z digest=sha256:f8b7397db638598412e3be9330a5b338c8a036ee657b829bec1d28bdf03d4a70

Observation d6d8958e-1ec8-4462-9ae4-d15fa45b3c4b · outbound

This paper cites LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.622526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.622526Z digest=sha256:649a26c9f60c990e1623d53253cf32bb529854d145dd7b91e397a7cf4923eade

Observation a2bdf34d-1c59-44e8-ba29-1d6dd7929034 · outbound

This paper cites An Empirical Evaluation of LLMs for Solving Offensive Security Challenges.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.627164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.627164Z digest=sha256:80f2db2f122a5ebd2c60bea23f85ee7d608fcce432e1afc429c4328460589c7a

Observation 8cdf4899-13d1-4341-9b92-d3897291a8c3 · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLM Agents can Autonomously Hack Websites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.631740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.631740Z digest=sha256:f41e4463c78d252c182e104afa24c630c048fa698f5a4b687bda1370da351e47

Observation b8956961-49fd-4354-b263-91394b2e0055 · outbound

This paper cites Llm4decompile: Decompilingbinarycodewithlargelanguagemodels,.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Llm4decompile: Decompilingbinarycodewithlargelanguagemodels,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.298021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:32.637031Z digest=sha256:05b281313159400ceb7f208a0be9b28c024926828058b6cff44c07500b1ceb98

Observation f6cc0e37-7723-497a-a17c-3d861bd3a300 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.644504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.644504Z digest=sha256:c834d81f967f13decbe23307c81907b5e9e7dd9b6f8a5fa2730fdd61942ac89b

Observation b17f629e-780e-4ca8-b4e4-0430f937be60 · outbound

This paper cites ATLAS - Adversarial Threat Landscape for Artificial-Intelligence Systems.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models ATLAS - Adversarial Threat Landscape for Artificial-Intelligence Systems

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.287871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:32.674375Z digest=sha256:91c5c2422a824dfd43f6a44575c27ae6cbdb60fb173bc53ee979bed874cd5fac

Observation a10781cf-654c-441a-ac0e-f37410e03917 · outbound

This paper cites OWASP Top Ten for Large Language Model Applications.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OWASP Top Ten for Large Language Model Applications

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.276513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:32.679927Z digest=sha256:a3d5d0671b73c846a1cbc7592bb9dc1f4f8f848de748df146d140d40fd7ef43a

Observation cd5fef5e-240c-472a-953b-0c62d239a656 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Measuring Massive Multitask Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.684251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.684251Z digest=sha256:e652d0469c817e1e9d913746e3460f4f5ad8f4f2158b8f588767ce5fc29eaf1a

Observation ee12daaa-f90a-435c-a0a1-70b9f18a4511 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.689085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.689085Z digest=sha256:5062a6050de3bd1b7c6b15e4b01592f6d1a6d12d0ab0672448b7be3778b66cef

Observation 786fffbf-ad22-4c29-9c6b-fc7f941845ed · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.693437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.693437Z digest=sha256:b55a568917f9dfd70e9436f4510851678dd8dd65dca9df47dd38f7fc4204ed60

Observation 1aec31e9-929e-4d0b-91d4-37a68ad00980 · outbound

This paper cites OpenAI o1 System Card.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OpenAI o1 System Card

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.262382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:32.697337Z digest=sha256:8d7bb30b40909da4d175d65c89e27f8a77e79aa788f8c999158fbc8a732d365d

Observation 648e0d2c-4834-4711-b100-5d48e8b50c1e · outbound

This paper cites Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.701223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.701223Z digest=sha256:36e52316c7bb17c09b52f95db218680ba3d910701770d927bfdd875ed10cf971

Observation 37b18dbc-7422-4518-9dc3-0f954283512b · outbound

This paper cites Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.705945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.705945Z digest=sha256:9660e7eb5375cf242c47bad57a914e354b48658dc688c45ba86722dcae5b1ee0

Observation 415e0e21-ec74-46e4-b7a1-af7ae258fcce · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.725506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.725506Z digest=sha256:d69467e17024e0fb5de241ed31e2a5037cb2a982b205bf1ae0c357372f3dd25e

Observation 08056cae-2d1e-4ed9-962f-f7d42c40eb71 · outbound

This paper cites Jimenez, John Yang, Kai Liu, and Aleksander Madry.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Jimenez, John Yang, Kai Liu, and Aleksander Madry

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.249705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:32.756599Z digest=sha256:7f56681caa51c3e20db8f7bfd20c6df759a21c897fc8aeb24459502e2d4eb3d5

Observation 113acd9c-9a00-4443-9b82-41d4088c02d2 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.783357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.783357Z digest=sha256:b2a37405018b21f588193651ca3e02946be7a45c0f4a14f07720c4da8067d621

Observation a39166fe-2e49-45b5-96c4-593488160ad8 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models AgentBench: Evaluating LLMs as Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.791730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.791730Z digest=sha256:9916764e06ffa0f4fd4b29e469bd8cb8b1330986c677a3b8a9b320dca706701c

Observation 1ac4ff36-5be3-4839-8875-091815b144b3 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.795634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.795634Z digest=sha256:d0f3c4d55c51a5f0bfe3f276607a431e91ecdfeff6d07df505c6ce4833e36ad9

Observation b8383111-e426-45e0-b10a-14028c21132d · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Mind2Web: Towards a Generalist Agent for the Web

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.803200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.803200Z digest=sha256:f126f75c412300509627c94078e56d57c33cbf99596b4c9fcc3beec27ecffd8a

Observation 1608e2a3-8a37-42b2-9663-684db5f7e669 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.807624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.807624Z digest=sha256:b96c1526f4c4b686f9f8decf32add92e4151cff9a43959d5621e111d8c9b42f5

Observation dd6e4679-7a17-4487-8954-34b71f13326a · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:53:34.149152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:32.814010Z digest=sha256:0e11a0a3ae3c7fdd43ba64785dccafa6e60efd8e45f225e96ae24a22497d0cbf

Observation 0a2b9505-3ecd-448a-985a-de94353d94c7 · outbound

This paper cites InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.840618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.840618Z digest=sha256:6fd143b1c0f00108e05277b3f797bafa2d085ae8d68548921802a58e1c88146f

Observation ef0eb5b6-ae4b-416e-a836-0b3f445ab6ca · outbound

This paper cites NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.863240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.863240Z digest=sha256:faf1ecea8c6e136d7343c5b4f911f20b8c4f915680ac94db9a496663638fbe44

Observation 8cbe8746-024a-480d-8048-9548adcda81f · outbound

This paper cites EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.885270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.885270Z digest=sha256:a11b10a473d171f45e2405a8e62691a66431ff1a23fb030a17c50185de923817

Observation 24625837-5955-4a37-a26f-c90c28372f39 · outbound

This paper cites AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.995201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.995201Z digest=sha256:1a09c2b63bf74757533f0a995084ed17705aa17ed5d77e104594049427d504e8

Observation 38c7678a-0a84-41de-9196-0ece445f7681 · outbound

This paper cites Picoctf learning.https://www.picoctf.org/, 2024.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Picoctf learning.https://www.picoctf.org/, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.077377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.103438Z digest=sha256:c3be9a502a3cd620ee7d0be00aea62a8ace30b23cc8d894f2c33946e6c1e028d

Observation 54ec6960-1b49-4018-b144-1c0b1af26992 · outbound

This paper cites Ai village capture the flag @ defcon31.https://kaggle.com/ competitions/ai-village-capture-the-flag-defcon31, 2023.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Ai village capture the flag @ defcon31.https://kaggle.com/ competitions/ai-village-capture-the-flag-defcon31, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.026552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.122420Z digest=sha256:09824451de5dcee449134dca14046d06bcfa0922403493abbd0edc36d1aa8429

Observation d877e2e1-e64c-4c89-8521-0132a06329a0 · outbound

This paper cites Jupyter datascience notebook docker image, 2025.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Jupyter datascience notebook docker image, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:34.016063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.131197Z digest=sha256:6264abc5baa45ccda1a0d1b8407dc69185c5a3e62e0b5ab6a4eaaf788fe6d95e

Observation 72825d09-5f60-42c1-8844-985a27434988 · outbound

This paper cites Optimizing Large Language Model Hyperparameters for Code Generation.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Optimizing Large Language Model Hyperparameters for Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:33.154793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:33.154793Z digest=sha256:b391948fd004444e6c50c9516c0725b512922c5482b4d1839d24b1ae7b7b01ad

Observation dbdd6fc6-cfb3-4bf4-9e40-81635bd17218 · outbound

This paper cites The Automation Advantage in AI Red Teaming.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models The Automation Advantage in AI Red Teaming

Reference 31

Resolution
malformed identifier
no resolver link, observed 2026-08-15T19:53:33.168800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:33.168800Z digest=sha256:d043c7daeb26b7e8c11f030e8483ee989a50a1cbd9762f6c5dc6eeccb5996331

Observation 4d833146-c9cd-4247-aa55-ccfec3df252c · outbound

This paper cites For example : ‘t = turtle.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models For example : ‘t = turtle

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.981683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.281100Z digest=sha256:5669d68431aad95ad357c288a5c253d175cd1b8403004d656f37513e73b84302

Observation 69d706f8-82b6-4039-bad0-b471a24bdfc3 · outbound

This paper cites S Y S T E M _ C O M M A N D _ E X E C U T E D _ V I A _ E X E C.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models S Y S T E M _ C O M M A N D _ E X E C U T E D _ V I A _ E X E C

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.970997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.295834Z digest=sha256:2d8f1825d86d6d4e32559031a650cf7b6f509a64d9785fe7a4f86b9bfae9635c

Observation ec659ead-6b3c-40ff-8c8f-0b0fbabd57e1 · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:53:33.959915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.306754Z digest=sha256:a393e0cd1b0c8af453d0ae79d642b54bd439581a1ee7b74f3fef729b0ea90832

Observation d19cd96b-6aff-4075-9e4c-415cf88b77b7 · outbound

This paper cites For example : ‘t.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models For example : ‘t

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.949128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.310734Z digest=sha256:5448c8b9a62f6271efb402011a6915b37ce3f90cddcbc86c31910d0cf39d73ce

Observation 1953971f-4bdf-445d-a1dc-cdb47053335a · outbound

This paper cites " " response = query ( prompt ) print ( response ) </ execute - code > <result idx=0 success=True> ’output’:.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models " " response = query ( prompt ) print ( response ) </ execute - code > <result idx=0 success=True> ’output’:

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.911454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.314088Z digest=sha256:4d53d76f9853de3f49444f1d2abde70da1dbd7b0795ad320713a85adfde1a408

Observation 54452e34-cca5-47bf-84db-6e10008c0787 · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:53:34.004403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.325454Z digest=sha256:93987df48bf4ec2c232ffcfa56890744eaf212926d43b9f89f9692ceff33ab37

Observation 26208294-c958-4518-9989-3308f74bc1ef · outbound

This paper cites an unresolved cited work.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work

Reference 41

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T19:53:33.992721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.353697Z digest=sha256:9e595f58efa50991a379ca3b6c7f8940e9007f969b45c0d5ab1da82951663fdb

Observation a264a8e2-edc5-4510-bf98-42b4c350f91e · outbound

This paper cites Example : ‘t = turtle.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Example : ‘t = turtle

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.840388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.388384Z digest=sha256:0245e4c1713a5d3756377abffb65906e6179a8eabf96de2bdc91e6e66831b055

Observation f768c676-7a52-4781-8769-fc5bad636043 · outbound

This paper cites echo J A G A _ L L M _ G E N E R A T E D _ O S _ C O M M A N D.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models echo J A G A _ L L M _ G E N E R A T E D _ O S _ C O M M A N D

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.794439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.430804Z digest=sha256:6a97d525ea521b357b10cb58d3f29afaaec2af6e4dbb204a5bc7c5a2dbbdab4a

Observation 46ed088a-d71e-42bc-a856-3b11fd256ca3 · outbound

This paper cites system () ‘ call , make the turtle ‘t ‘ draw a square of side 10.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models system () ‘ call , make the turtle ‘t ‘ draw a square of side 10

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.783066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.465619Z digest=sha256:802dc92c58bd21621000752fe7bc7bd76eff5b17df2ddc3249af094195342f4b

Observation ef0f6acf-4d21-45a9-ba3c-2e6328803dec · outbound

This paper cites " " response = requests . post ( f.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models " " response = requests . post ( f

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:53:33.770156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:53:33.486755Z digest=sha256:80388079de6a944ec905bc668a5551c40ddaaa3bdfa8082dc6d52fb7a7507622

Observation c505c07d-fa71-48f9-a6f4-360ef26a4f7d · outbound

This paper cites LLM4Decompile: Decompiling Binary Code with Large Language Models.

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLM4Decompile: Decompiling Binary Code with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:32.640743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:32.640743Z digest=sha256:d00d1c56f933b0d181f17cde6237696e914084fa623230352c55f1672fd0211a

Pith citing papers

Observation d2a4ce5b-b37f-44f8-b429-c30af3ddcabc · inbound

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security cites this paper.

Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:25:15.771731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:25:15.771731Z digest=sha256:163bc50e8cd2919bec03d05f5d179dac0507208271e5fb2ad36235293cdc53c6

Observation 3dff4f9f-4458-46a4-955f-7c6083e29f80 · inbound

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts cites this paper.

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:09:33.025834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:09:33.025834Z digest=sha256:bfcaab4fd416e7fd3a7f3f0766d5b62670283bb72ca29d6adb2d1fd087bfe787

Observation 898cf871-8391-4070-931f-3fb691354ef7 · inbound

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours cites this paper.

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:51:45.056581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T16:06:18.057868Z digest=sha256:8d0e67d312399878c1ef7375dec86694c5516becbfdf143477b165b6a2e9bef1