Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T06:48:33.255423Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2601.22136.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T06:48:33.255423Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T21:12:50.686474Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T21:15:04.163349Z
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 64bd849a-d9fa-43f1-b34e-893d1edc2cf7 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents AgentHarm : A benchmark for measuring harmfulness of LLM agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5dfa6f7-bfcd-413f-b3c3-513ac6e6929f · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents ShieldAgent : Shielding agents via verifiable safety policy reasoning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc911a9-6e9a-4465-b246-efe51ee96d5e · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents AI coding tool wiped our database, says startup in catastrophic failure
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ad13b9-97c3-4c95-95a8-58816ee72d03 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9349ea88-b7a4-46f3-85e0-9ed3621696d6 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents On the computational complexity of self-attention
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 912cb6c6-a9e0-4919-9895-00af93d91ec1 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents Specification gaming: the flip side of AI ingenuity
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c1726a-0a9b-49b7-98d6-e8aacb93e39d · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b2c6108-81db-4274-bf75-d3f63ad0e6d0 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents AgentBench : Evaluating LLMs as agents
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc275b6-2049-4bbe-9bbc-8bece699ce32 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents GAIA : A benchmark for general AI assistants
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ca2469a-38e3-4fd2-9e75-7eeda60a5d84 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents Discovering Language Model Behaviors with Model-Written Evaluations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738cb261-02f8-40c9-9be6-f77e14aaba5a · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents Identifying risks of LM agents with an LM -emulated sandbox
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c277faaa-63a7-472b-a685-0c478c99983c · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents Toolformer: Language models can teach themselves to use tools
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4928575-9cf5-4ef3-81db-a4143043e496 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents SafeArena : Evaluating the safety of autonomous web agents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9a8c388-065a-4cd9-8fd9-f94e79554392 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a70930-fa0e-405e-9a3c-f74b81c056fc · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents GuardAgent : Safeguard LLM agents via knowledge-enabled reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d97c84ab-9c89-45ed-be05-c2d32f2cf919 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a63ded48-0b2e-4a79-8353-b5b4e0056a43 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ed57ae-ac9b-4d76-9445-412142a8fcb8 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents ReAct : Synergizing reasoning and acting in language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 852df4b5-06c0-47ef-bc11-7862bb19d395 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents SafeAgentBench : A benchmark for safe task planning of embodied LLM agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6758dd8f-abbc-4e88-b0fa-2696c8c7aa84 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a84379fe-a046-422c-8e31-03cd8df422ac · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents Agent-SafetyBench: Evaluating the Safety of LLM Agents
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c33a4aee-cdf4-4ea5-9312-67612fb98833 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85d149df-4e38-4b02-80a9-8311d60642dd · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents WebArena : A realistic web environment for building autonomous agents
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd92b36d-2730-4f00-9d37-cc90ffe96690 · outbound
StepShield: When, Not Whether to Intervene on Rogue Agents Agent-as-a-judge: Evaluate agents with agents
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3ccaed7-18a8-46ed-8ecc-dae40298f0c5 · inbound
ProjGuard: Safety Monitoring for Computer-Use Agents via Low-Dimensional Projections StepShield: When, Not Whether to Intervene on Rogue Agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.