Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T00:13:38.850570Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2607.26314.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T00:13:38.850570Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0e80cefe-c5ec-46ef-9ce3-e925f7742d4a · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ReAct: Synergizing Reasoning and Acting in Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2440576b-c5ce-4ad2-a8c2-a45c00c72c7f · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9198dcf8-bd12-4b6a-a7ec-0677a79acbe1 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e11cae00-7518-4d02-8a11-e96b8912a597 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ARACNE: An LLM-Based Autonomous Shell Pentesting Agent
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08feda12-dbee-4387-a49f-dfe6fbafe964 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ARTEMIS: Automated red teaming engine with multi-agent intelligent supervision.https://github.com/Stanford-Trinity/ARTEMIS, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c75942d-277b-4329-a925-2a8d7343f2de · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05388b2e-8105-4398-86e9-51c3a55adb32 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 777c2d43-5634-4cbb-84bf-9866060db016 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents AI Control: Improving Safety Despite Intentional Subversion
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da79857c-91b2-4f6b-8768-3c0551638f7f · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents The first confirmed instance of an LLM going rogue for real.https://www.lesswrong.com/posts/XRADGH4BpRKaoyqcs/ the-first-confirmed-instance-of-an-llm-going-rogue-for, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a698fc1-2c92-4c01-b15c-85e23f7d030e · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Security incident disclosure — July 2026.https: //huggingface.co/blog/security-incident-july-2026, 2026
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b8e5f5-8469-4d68-80f0-55952a6b0647 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Hugging Face model evaluation security incident.https://openai.com/index/ hugging-face-model-evaluation-security-incident/, 2026
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f980912-ee82-4f2e-97ba-596d77b2fe0c · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd87db3-b8be-480c-b0fc-8eee94595811 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3924c511-0f88-4f10-9253-33db62c09527 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Acoefficientofagreementfornominalscales.Educational and Psychological Measurement, 20(1):37–46, 1960
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42314e25-6e06-468e-b2d9-1cae28b46b99 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45dd4561-280a-4deb-9da0-5abef8acc08c · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Richard Landis and Gary G
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9267c136-32a2-45fd-bbcc-21459b50cce3 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Measuring Progress on Scalable Oversight for Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4affbfa-247f-4f03-a1b2-6b27d9230b78 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Securing AI Agents with Information-Flow Control
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86876a7f-f4df-452d-ae8a-2cf2f825eee1 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9493587-95d8-4caf-86f3-d8e75d47c0ca · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Agent-Sentry: Bounding LLM Agents via Execution Provenance
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b83292-3333-4b34-a7b5-d45186900daf · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Reachability Across the NL/PL Boundary: A Taxonomy-Driven Dataflow Model for LLM-Integrated Applications
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5238f3-7892-4ff2-8de0-9bdc6fe62b34 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 641b64b0-470f-40e0-92a6-ca1272a37e96 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents AgentLeak: A bench- mark for internal-channel privacy leakage in multi-agent LLM systems.arXiv preprint arXiv:2602.11510, 2026
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2fa30d1-67cd-48f7-aa8e-2640c8efc208 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents R-Judge: Benchmarking safety risk awareness for LLM agents
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e5d497d-fcc8-4626-9813-2f09af110ba5 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents AgentRaft: Automated detection of data over-exposure in LLM agents.arXiv preprint arXiv:2603.07557, 2026
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df107a2f-a6d3-475a-b07c-c8e00779e2f2 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ToolSafe: Enhancing tool invocation safety of LLM-based agents via proactive step-level guardrail and feedback.arXiv preprint arXiv:2601.10156, 2026
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9abeb8cc-1606-46b7-a7bc-4ac28bb28e2e · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fed7f13-aa82-4871-8cc2-b238786c407e · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents The state of secrets sprawl 2026
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddbcb8ed-12d3-4099-8824-aec0b2e6093d · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents PentestJudge: Judging Agent Behavior Against Operational Requirements
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fde0c26-8131-4a68-aff5-cb8490ca0615 · outbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Unresolved cited work
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.