Pith. sign in

Paper Citation Record · LEDGER

StepShield: When, Not Whether to Intervene on Rogue Agents

As of 16 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2601.22136.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.22136 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:48:33.255423Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:12:50.686474Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:15:04.163349Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 64bd849a-d9fa-43f1-b34e-893d1edc2cf7 · outbound

This paper cites AgentHarm : A benchmark for measuring harmfulness of LLM agents.

StepShield: When, Not Whether to Intervene on Rogue Agents AgentHarm : A benchmark for measuring harmfulness of LLM agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:30.823358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:30.823358Z digest=sha256:b717a95309450f23e50cc3fd9d1665bb9bb0e376a7fdb7a5c13538b1600f71f0

Observation f5dfa6f7-bfcd-413f-b3c3-513ac6e6929f · outbound

This paper cites ShieldAgent : Shielding agents via verifiable safety policy reasoning.

StepShield: When, Not Whether to Intervene on Rogue Agents ShieldAgent : Shielding agents via verifiable safety policy reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:30.930571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:30.930571Z digest=sha256:5034c2bb99c2b62f954fa27885244b18a3cf6f7b18e1f1169613bcb99907b58b

Observation fcc911a9-6e9a-4465-b246-efe51ee96d5e · outbound

This paper cites AI coding tool wiped our database, says startup in catastrophic failure.

StepShield: When, Not Whether to Intervene on Rogue Agents AI coding tool wiped our database, says startup in catastrophic failure

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:31.111845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:31.111845Z digest=sha256:079314ebb8378e398155087b2f0c03afa75a22e08dc333838cef2876dd31d31d

Observation e4ad13b9-97c3-4c95-95a8-58816ee72d03 · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

StepShield: When, Not Whether to Intervene on Rogue Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:31.340685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:31.340685Z digest=sha256:38d21167b07951a98a99cfb35cd0943cc93bf0440800e3362df3ce54c1d81c59

Observation 9349ea88-b7a4-46f3-85e0-9ed3621696d6 · outbound

This paper cites On the computational complexity of self-attention.

StepShield: When, Not Whether to Intervene on Rogue Agents On the computational complexity of self-attention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:31.452397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:31.452397Z digest=sha256:909e3cfab0b902025101cd9a0443883d602f883c6cc378520a0a0b5a7d7c04ff

Observation 912cb6c6-a9e0-4919-9895-00af93d91ec1 · outbound

This paper cites Specification gaming: the flip side of AI ingenuity.

StepShield: When, Not Whether to Intervene on Rogue Agents Specification gaming: the flip side of AI ingenuity

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:31.595193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:31.595193Z digest=sha256:286ac0f27883065964b656f4b16596613f985531a6290f38c7134c9e84237274

Observation f2c1726a-0a9b-49b7-98d6-e8aacb93e39d · outbound

This paper cites SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents.

StepShield: When, Not Whether to Intervene on Rogue Agents SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:31.664952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:31.664952Z digest=sha256:2e83c7342c5585544b25585f8cb0efb22fa493a837ac5c768ff2da6191af45c3

Observation 8b2c6108-81db-4274-bf75-d3f63ad0e6d0 · outbound

This paper cites AgentBench : Evaluating LLMs as agents.

StepShield: When, Not Whether to Intervene on Rogue Agents AgentBench : Evaluating LLMs as agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:31.755426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:31.755426Z digest=sha256:1650f362b60d271d2c7d38167bdfefebb7e15de5e9f13bcccf9c749798b31be8

Observation 5fc275b6-2049-4bbe-9bbc-8bece699ce32 · outbound

This paper cites GAIA : A benchmark for general AI assistants.

StepShield: When, Not Whether to Intervene on Rogue Agents GAIA : A benchmark for general AI assistants

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:31.823861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:31.823861Z digest=sha256:379c8a1f61af3820e4f6c6c0768befb5922f1b0b6e74c18ae06112f9d8934d78

Observation 3ca2469a-38e3-4fd2-9e75-7eeda60a5d84 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

StepShield: When, Not Whether to Intervene on Rogue Agents Discovering Language Model Behaviors with Model-Written Evaluations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:31.914327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:31.914327Z digest=sha256:d8f85ed6fda54db579cde693f28272c7011ee4746d7a3494804eb99b31c7bd92

Observation 738cb261-02f8-40c9-9be6-f77e14aaba5a · outbound

This paper cites Identifying risks of LM agents with an LM -emulated sandbox.

StepShield: When, Not Whether to Intervene on Rogue Agents Identifying risks of LM agents with an LM -emulated sandbox

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:31.979160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:31.979160Z digest=sha256:737d90bb159c6fd6377fd92e59c9c5ec9517240bfb614f4960673d1a6c4d1645

Observation c277faaa-63a7-472b-a685-0c478c99983c · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

StepShield: When, Not Whether to Intervene on Rogue Agents Toolformer: Language models can teach themselves to use tools

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:32.034384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:32.034384Z digest=sha256:1322ce9073d66ad8251f27370e5fba1bc81149b942150c76861147219e09c4e4

Observation a4928575-9cf5-4ef3-81db-a4143043e496 · outbound

This paper cites SafeArena : Evaluating the safety of autonomous web agents.

StepShield: When, Not Whether to Intervene on Rogue Agents SafeArena : Evaluating the safety of autonomous web agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:32.094844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:32.094844Z digest=sha256:e0c9e297ea5d38860898b446cb3482e3600f3362bdd957cd4eccab7988454a92

Observation d9a8c388-065a-4cd9-8fd9-f94e79554392 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

StepShield: When, Not Whether to Intervene on Rogue Agents Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:32.145784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:32.145784Z digest=sha256:27fa13b31bf0599bcfbc9eb0b617213232b0807cdf712778e4c40871c7d147c1

Observation a2a70930-fa0e-405e-9a3c-f74b81c056fc · outbound

This paper cites GuardAgent : Safeguard LLM agents via knowledge-enabled reasoning.

StepShield: When, Not Whether to Intervene on Rogue Agents GuardAgent : Safeguard LLM agents via knowledge-enabled reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:32.322808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:32.322808Z digest=sha256:8ef3ea9a0b03ada73c5bc00615121b7654b6818d4bbc612cf9abab1b908425ea

Observation d97c84ab-9c89-45ed-be05-c2d32f2cf919 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

StepShield: When, Not Whether to Intervene on Rogue Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:32.428948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:32.428948Z digest=sha256:a6017287828a94559e2f642024320289d5a4d5d4b0c7e23746aee8a217814042

Observation a63ded48-0b2e-4a79-8353-b5b4e0056a43 · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

StepShield: When, Not Whether to Intervene on Rogue Agents TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:32.516604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:32.516604Z digest=sha256:586e55b143b856b3f13d7832a8c1bd231746d504c770e33edcbfb12e268ca641

Observation 03ed57ae-ac9b-4d76-9445-412142a8fcb8 · outbound

This paper cites ReAct : Synergizing reasoning and acting in language models.

StepShield: When, Not Whether to Intervene on Rogue Agents ReAct : Synergizing reasoning and acting in language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:32.646810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:32.646810Z digest=sha256:92afb23f01f68d37027ba0824d48024a6cfd6d2505d4093ba226af5b2671757e

Observation 852df4b5-06c0-47ef-bc11-7862bb19d395 · outbound

This paper cites SafeAgentBench : A benchmark for safe task planning of embodied LLM agents.

StepShield: When, Not Whether to Intervene on Rogue Agents SafeAgentBench : A benchmark for safe task planning of embodied LLM agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:32.798048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:32.798048Z digest=sha256:8308fff2058b4404656773e67fb8f5ac6f8d02e56a9443b1858f18c7a163abfb

Observation 6758dd8f-abbc-4e88-b0fa-2696c8c7aa84 · outbound

This paper cites R-Judge: Benchmarking Safety Risk Awareness for LLM Agents.

StepShield: When, Not Whether to Intervene on Rogue Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:32.932827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:32.932827Z digest=sha256:fd786dda2fbcd21f48081038941c41a40357d1faf1e51d1f916267f6b696923d

Observation a84379fe-a046-422c-8e31-03cd8df422ac · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

StepShield: When, Not Whether to Intervene on Rogue Agents Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:33.022088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:33.022088Z digest=sha256:085ca6aea46433caa6f50236f03086b05392a688ab836b14a326fd8ea806a9a7

Observation c33a4aee-cdf4-4ea5-9312-67612fb98833 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

StepShield: When, Not Whether to Intervene on Rogue Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:33.095817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:33.095817Z digest=sha256:4af403efa4c46c3738fcff30b3537d8e330427e409324b37a9a4ebcd090425e5

Observation 85d149df-4e38-4b02-80a9-8311d60642dd · outbound

This paper cites WebArena : A realistic web environment for building autonomous agents.

StepShield: When, Not Whether to Intervene on Rogue Agents WebArena : A realistic web environment for building autonomous agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:33.184945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:33.184945Z digest=sha256:d4d5cd72d13759f2f53926649b956470597788b1bcc74b5c8e73b2395a4d9c20

Observation bd92b36d-2730-4f00-9d37-cc90ffe96690 · outbound

This paper cites Agent-as-a-judge: Evaluate agents with agents.

StepShield: When, Not Whether to Intervene on Rogue Agents Agent-as-a-judge: Evaluate agents with agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:33.255423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:33.255423Z digest=sha256:2b57e663fcd22296c74d67f50ada9e6740566b844c11f8d9ddcaf8759e362aaf

Pith citing papers

Observation e3ccaed7-18a8-46ed-8ecc-dae40298f0c5 · inbound

ProjGuard: Safety Monitoring for Computer-Use Agents via Low-Dimensional Projections cites this paper.

ProjGuard: Safety Monitoring for Computer-Use Agents via Low-Dimensional Projections StepShield: When, Not Whether to Intervene on Rogue Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:10.878626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T21:12:50.686474Z digest=sha256:bc2e86f915b35bce4bb71b54473ba0e2cce4271d9de3511e9d884641d6b842e3