Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:53:33.486755Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2506.14682.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:53:33.486755Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:25:15.771731Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T23:51:45.015967Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5786a459-44e7-455b-ae16-95ad0238f8f1 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6d8958e-1ec8-4462-9ae4-d15fa45b3c4b · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2bdf34d-1c59-44e8-ba29-1d6dd7929034 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cdf4899-13d1-4341-9b92-d3897291a8c3 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLM Agents can Autonomously Hack Websites
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8956961-49fd-4354-b263-91394b2e0055 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Llm4decompile: Decompilingbinarycodewithlargelanguagemodels,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f6cc0e37-7723-497a-a17c-3d861bd3a300 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Evaluating Large Language Models Trained on Code
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b17f629e-780e-4ca8-b4e4-0430f937be60 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models ATLAS - Adversarial Threat Landscape for Artificial-Intelligence Systems
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a10781cf-654c-441a-ac0e-f37410e03917 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OWASP Top Ten for Large Language Model Applications
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cd5fef5e-240c-472a-953b-0c62d239a656 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Measuring Massive Multitask Language Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee12daaa-f90a-435c-a0a1-70b9f18a4511 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 786fffbf-ad22-4c29-9c6b-fc7f941845ed · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aec31e9-929e-4d0b-91d4-37a68ad00980 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OpenAI o1 System Card
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 648e0d2c-4834-4711-b100-5d48e8b50c1e · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37b18dbc-7422-4518-9dc3-0f954283512b · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 415e0e21-ec74-46e4-b7a1-af7ae258fcce · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08056cae-2d1e-4ed9-962f-f7d42c40eb71 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Jimenez, John Yang, Kai Liu, and Aleksander Madry
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 113acd9c-9a00-4443-9b82-41d4088c02d2 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a39166fe-2e49-45b5-96c4-593488160ad8 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models AgentBench: Evaluating LLMs as Agents
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ac4ff36-5be3-4839-8875-091815b144b3 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8383111-e426-45e0-b10a-14028c21132d · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Mind2Web: Towards a Generalist Agent for the Web
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1608e2a3-8a37-42b2-9663-684db5f7e669 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd6e4679-7a17-4487-8954-34b71f13326a · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0a2b9505-3ecd-448a-985a-de94353d94c7 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0eb5b6-ae4b-416e-a836-0b3f445ab6ca · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cbe8746-024a-480d-8048-9548adcda81f · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24625837-5955-4a37-a26f-c90c28372f39 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c7678a-0a84-41de-9196-0ece445f7681 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Picoctf learning.https://www.picoctf.org/, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 54ec6960-1b49-4018-b144-1c0b1af26992 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Ai village capture the flag @ defcon31.https://kaggle.com/ competitions/ai-village-capture-the-flag-defcon31, 2023
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d877e2e1-e64c-4c89-8521-0132a06329a0 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Jupyter datascience notebook docker image, 2025
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 72825d09-5f60-42c1-8844-985a27434988 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Optimizing Large Language Model Hyperparameters for Code Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbdd6fc6-cfb3-4bf4-9e40-81635bd17218 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models The Automation Advantage in AI Red Teaming
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d833146-c9cd-4247-aa55-ccfec3df252c · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models For example : ‘t = turtle
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 69d706f8-82b6-4039-bad0-b471a24bdfc3 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models S Y S T E M _ C O M M A N D _ E X E C U T E D _ V I A _ E X E C
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ec659ead-6b3c-40ff-8c8f-0b0fbabd57e1 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d19cd96b-6aff-4075-9e4c-415cf88b77b7 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models For example : ‘t
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1953971f-4bdf-445d-a1dc-cdb47053335a · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models " " response = query ( prompt ) print ( response ) </ execute - code > <result idx=0 success=True> ’output’:
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 54452e34-cca5-47bf-84db-6e10008c0787 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 26208294-c958-4518-9989-3308f74bc1ef · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a264a8e2-edc5-4510-bf98-42b4c350f91e · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models Example : ‘t = turtle
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f768c676-7a52-4781-8769-fc5bad636043 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models echo J A G A _ L L M _ G E N E R A T E D _ O S _ C O M M A N D
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 46ed088a-d71e-42bc-a856-3b11fd256ca3 · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models system () ‘ call , make the turtle ‘t ‘ draw a square of side 10
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ef0f6acf-4d21-45a9-ba3c-2e6328803dec · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models " " response = requests . post ( f
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c505c07d-fa71-48f9-a6f4-360ef26a4f7d · outbound
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models LLM4Decompile: Decompiling Binary Code with Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2a4ce5b-b37f-44f8-b429-c30af3ddcabc · inbound
Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dff4f9f-4458-46a4-955f-7c6083e29f80 · inbound
PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 898cf871-8391-4070-931f-3fb691354ef7 · inbound
Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.