Pith. sign in

Paper Citation Record · LEDGER

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

As of 17 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2607.26314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26314 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T00:13:38.850570Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e80cefe-c5ec-46ef-9ce3-e925f7742d4a · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:36.762325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:36.762325Z digest=sha256:5c206a3033e7383edac4291a8b943c3a5973c1d48c0a42ee33ad1a1e6b71f692

Observation 2440576b-c5ce-4ad2-a8c2-a45c00c72c7f · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:36.832447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:36.832447Z digest=sha256:20625a957f3cd8b081149f8336477c9db2f7915e0a888a958c3d6bd737adc31a

Observation 9198dcf8-bd12-4b6a-a7ec-0677a79acbe1 · outbound

This paper cites HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:36.891218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:36.891218Z digest=sha256:f6bcb838919a01ba741a1248c9208baca2010089f1339f0a3676887b8723c884

Observation e11cae00-7518-4d02-8a11-e96b8912a597 · outbound

This paper cites ARACNE: An LLM-Based Autonomous Shell Pentesting Agent.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ARACNE: An LLM-Based Autonomous Shell Pentesting Agent

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:36.977658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:36.977658Z digest=sha256:6279b448fc38bb0f35ec384309e367534b449621fe87de22919009b0f468eaae

Observation 08feda12-dbee-4387-a49f-dfe6fbafe964 · outbound

This paper cites ARTEMIS: Automated red teaming engine with multi-agent intelligent supervision.https://github.com/Stanford-Trinity/ARTEMIS, 2025.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ARTEMIS: Automated red teaming engine with multi-agent intelligent supervision.https://github.com/Stanford-Trinity/ARTEMIS, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.065267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.065267Z digest=sha256:000e3fc04346679438831a32cb14325c524c1ab92799963be09cd639d5e2acd1

Observation 8c75942d-277b-4329-a925-2a8d7343f2de · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.291705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.291705Z digest=sha256:f1f6f6c5bae86cfe9ab93e42079eb36e3da038a5037dec7763f576ab950f5453

Observation 05388b2e-8105-4398-86e9-51c3a55adb32 · outbound

This paper cites ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.364940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.364940Z digest=sha256:b6390561b90e2a79762e4637767d49c07bcace9c9a2304533c46d7814dc52b14

Observation 777c2d43-5634-4cbb-84bf-9866060db016 · outbound

This paper cites AI Control: Improving Safety Despite Intentional Subversion.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents AI Control: Improving Safety Despite Intentional Subversion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.456860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.456860Z digest=sha256:f5baee2e55bafb1dd1338a784f1e8b082bc6113830688eb989135768392e86d3

Observation da79857c-91b2-4f6b-8768-3c0551638f7f · outbound

This paper cites The first confirmed instance of an LLM going rogue for real.https://www.lesswrong.com/posts/XRADGH4BpRKaoyqcs/ the-first-confirmed-instance-of-an-llm-going-rogue-for, 2025.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents The first confirmed instance of an LLM going rogue for real.https://www.lesswrong.com/posts/XRADGH4BpRKaoyqcs/ the-first-confirmed-instance-of-an-llm-going-rogue-for, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.532216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.532216Z digest=sha256:457d14eac63976d54c31ed941d2127a5d5c31439de9ff3f84c501e5d2c084c3e

Observation 2a698fc1-2c92-4c01-b15c-85e23f7d030e · outbound

This paper cites Security incident disclosure — July 2026.https: //huggingface.co/blog/security-incident-july-2026, 2026.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Security incident disclosure — July 2026.https: //huggingface.co/blog/security-incident-july-2026, 2026

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.604754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.604754Z digest=sha256:0d0920011480f54bce1eed67a7b0945ca9347e9b2022c20e2901945e71a0af25

Observation 05b8e5f5-8469-4d68-80f0-55952a6b0647 · outbound

This paper cites Hugging Face model evaluation security incident.https://openai.com/index/ hugging-face-model-evaluation-security-incident/, 2026.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Hugging Face model evaluation security incident.https://openai.com/index/ hugging-face-model-evaluation-security-incident/, 2026

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.662490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.662490Z digest=sha256:0de9ce437a3827d6a3a1e30419c64310c1aad012f476838cee17db98bcb1f754

Observation 0f980912-ee82-4f2e-97ba-596d77b2fe0c · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.719006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.719006Z digest=sha256:78fb35b00bd299a735f8354201e7b0fe4e4a5086e9fe4cd9a7c45e83e5a79df6

Observation 4fd87db3-b8be-480c-b0fc-8eee94595811 · outbound

This paper cites ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.772309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.772309Z digest=sha256:4323208abce78a02a88c37f7dc82ec03a9472e112793309685b164bca935fba0

Observation 3924c511-0f88-4f10-9253-33db62c09527 · outbound

This paper cites Acoefficientofagreementfornominalscales.Educational and Psychological Measurement, 20(1):37–46, 1960.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Acoefficientofagreementfornominalscales.Educational and Psychological Measurement, 20(1):37–46, 1960

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.844121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.844121Z digest=sha256:45dc60c6995d204d672133cbe4349ad5b2d670987a88dc616e8dca4b486a25f9

Observation 42314e25-6e06-468e-b2d9-1cae28b46b99 · outbound

This paper cites an unresolved cited work.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.921821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.921821Z digest=sha256:5653425a164f12934ec7d103fc50db6b638db4d2e391ab5d2710957023aa7bd0

Observation 45dd4561-280a-4deb-9da0-5abef8acc08c · outbound

This paper cites Richard Landis and Gary G.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Richard Landis and Gary G

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.986366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.986366Z digest=sha256:d681331cedfc640c359aea274e8095b28088f041e44190ed62c39b389607fa67

Observation 9267c136-32a2-45fd-bbcc-21459b50cce3 · outbound

This paper cites Measuring Progress on Scalable Oversight for Large Language Models.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Measuring Progress on Scalable Oversight for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.052542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.052542Z digest=sha256:0197f0879f6452d7a24bfdabd7f9f87b771f395312d9cc9d841ab7073c47fbb0

Observation e4affbfa-247f-4f03-a1b2-6b27d9230b78 · outbound

This paper cites Securing AI Agents with Information-Flow Control.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Securing AI Agents with Information-Flow Control

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.246732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.246732Z digest=sha256:6fd279320127f5c785b8dae243238a56b01d119b5c16731316d0a9433a7f84de

Observation 86876a7f-f4df-452d-ae8a-2cf2f825eee1 · outbound

This paper cites Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.185153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.185153Z digest=sha256:9a9fb9ec0e940048ba82542a901168f738fae9c45a1416c6b19b3e8c75865c4b

Observation d9493587-95d8-4caf-86f3-d8e75d47c0ca · outbound

This paper cites Agent-Sentry: Bounding LLM Agents via Execution Provenance.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Agent-Sentry: Bounding LLM Agents via Execution Provenance

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.358900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.358900Z digest=sha256:96012757415fe5e2593174fd80b6d5a32738a739ca656dfbacb621d70e2d7853

Observation 29b83292-3333-4b34-a7b5-d45186900daf · outbound

This paper cites Reachability Across the NL/PL Boundary: A Taxonomy-Driven Dataflow Model for LLM-Integrated Applications.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Reachability Across the NL/PL Boundary: A Taxonomy-Driven Dataflow Model for LLM-Integrated Applications

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.298962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.298962Z digest=sha256:bb2d7b11feeded5e3c129e2a5c41f557ec73211f762c14bcac79e708c24d2b86

Observation 9d5238f3-7892-4ff2-8de0-9bdc6fe62b34 · outbound

This paper cites How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.473059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.473059Z digest=sha256:14af08607f780914e0b3b992d6a15cfe15c134ae4f8d93425dbfc2d46f1a8f84

Observation 641b64b0-470f-40e0-92a6-ca1272a37e96 · outbound

This paper cites AgentLeak: A bench- mark for internal-channel privacy leakage in multi-agent LLM systems.arXiv preprint arXiv:2602.11510, 2026.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents AgentLeak: A bench- mark for internal-channel privacy leakage in multi-agent LLM systems.arXiv preprint arXiv:2602.11510, 2026

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.423705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.423705Z digest=sha256:f9b16e3daab9f5e7e2fedf1b4f5566a76a81e5b838a78e6d5aeda4d044216b47

Observation a2fa30d1-67cd-48f7-aa8e-2640c8efc208 · outbound

This paper cites R-Judge: Benchmarking safety risk awareness for LLM agents.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents R-Judge: Benchmarking safety risk awareness for LLM agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.605015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.605015Z digest=sha256:50434dbb7fb49759c8061ba13ff396b342e2c7ffaeacd4d9a37d79cd79416520

Observation 0e5d497d-fcc8-4626-9813-2f09af110ba5 · outbound

This paper cites AgentRaft: Automated detection of data over-exposure in LLM agents.arXiv preprint arXiv:2603.07557, 2026.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents AgentRaft: Automated detection of data over-exposure in LLM agents.arXiv preprint arXiv:2603.07557, 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.543533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.543533Z digest=sha256:8d6c76933b1fb0affe88cf781e1f161f6c68fddfadbf42ee3471b937bc6159e6

Observation df107a2f-a6d3-475a-b07c-c8e00779e2f2 · outbound

This paper cites ToolSafe: Enhancing tool invocation safety of LLM-based agents via proactive step-level guardrail and feedback.arXiv preprint arXiv:2601.10156, 2026.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ToolSafe: Enhancing tool invocation safety of LLM-based agents via proactive step-level guardrail and feedback.arXiv preprint arXiv:2601.10156, 2026

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.723773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.723773Z digest=sha256:7a691bbfe7d5134a7a68242d96a0fa56278cc6c499771eea0d2114085fefd22d

Observation 9abeb8cc-1606-46b7-a7bc-4ac28bb28e2e · outbound

This paper cites CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.664431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.664431Z digest=sha256:127532cdc9087239db8df816a6cf33d9ca4f70b2e9e740b9fb12b237fd87c7cb

Observation 6fed7f13-aa82-4871-8cc2-b238786c407e · outbound

This paper cites The state of secrets sprawl 2026.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents The state of secrets sprawl 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.850570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.850570Z digest=sha256:4421158a2ba0d76595008038316f883526e47246485bbff8d2bba5b3989c0125

Observation ddbcb8ed-12d3-4099-8824-aec0b2e6093d · outbound

This paper cites PentestJudge: Judging Agent Behavior Against Operational Requirements.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents PentestJudge: Judging Agent Behavior Against Operational Requirements

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:38.793264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:38.793264Z digest=sha256:ee225c16ef5b183d2ef73e0320bbc9f1b1b010c70aa900483672d0d7996d6370

Observation 2fde0c26-8131-4a68-aff5-cb8490ca0615 · outbound

This paper cites an unresolved cited work.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Unresolved cited work

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.205624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.205624Z digest=sha256:f42f68a17b84a37ac43ae5f6ef894c5ac4630b8b2c52bbf46cfd52c36b91ea22

Pith citing papers

No inbound Pith citation observations are available.