Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T04:08:05.893904Z
Paper Citation Record · LEDGER
As of 22 July 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2605.10779.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T04:08:05.893904Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 25347ef8-f034-433d-82f4-0317f3b8a608 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Agentharm: A benchmark for measuring harmfulness of llm agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 7a04aab9-ac02-4a30-8e31-b425cedff576 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments System card: Claude Sonnet 4.6
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fda9ff58-de75-42f5-861d-ad4916e81e84 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments OpenClaw 2026 security crisis: Protect your API keys now
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f17b7352-a605-41d7-be2f-cf80c44c2d32 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Windows agent arena: Evaluating multi-modal os agents at scale
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a8b99f5f-08e3-4e3e-ba1f-60660968ef3a · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Mind the gap: Text safety does not transfer to tool-call safety in llm agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 75b9a77b-fd56-4ca3-b28c-7d1710bb0bf0 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Teleai-safety: A comprehensive llm jailbreaking benchmark towards attacks, defenses, and evaluations
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d50e995d-1a9a-4250-a896-33517ebcf31d · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e03808e8-26fc-438b-ab90-759db19a0ec6 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments DeepSeek-V4: Towards highly efficient million-token context intelligence
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d1507b5f-e09d-4e89-b4fb-ec738fcf91a0 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 287fef8e-2eef-47bf-8305-0f94ee4cfeb9 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Sok: The attack surface of agentic ai–tools, and autonomy
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8865d338-5b76-4706-900b-6df475dbfcef · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Wasp: Benchmarking web agent security against prompt injection attacks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation ef8ede4c-3c49-4dbb-b7d2-898cb6af804d · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments OpenClaw security issues include data leakage & prompt injection
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 6f355e14-b6ac-42bf-a936-a9f769c8279d · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Security advisories for OpenClaw
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation cd7f7159-18a2-4d29-995f-9c2232da4f9d · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments SSRF in image tool remote fetch in OpenClaw.https://github .com/openclaw/openclaw/security/advisories/GHSA-56f2-hvwg-5743
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d843f2e7-b45a-44e7-b698-9b06b453eb5f · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments OpenClaw nostr privateKey config redaction bypass leaks plaintext signing key via config.get
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 34dacf78-8f10-45d8-9809-62bbfb04e622 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Gemini 3.1 Pro model card
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 98ec8e2c-151d-46c8-b4d3-4be3a69c242c · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fbe1e613-c4e2-4a2c-aeb0-8ca6f54c08a3 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Researchers reveal six new OpenClaw vulnerabilities
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 591d3d8c-5b9a-45f5-82c2-16605617d6b7 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Jiang, Y
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e9ca5bb6-b05d-44e4-a8bd-79a41dbbeec4 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Os-harm: A benchmark for measuring safety of computer use agents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f0a8a938-e0b2-43b7-bda1-42eb5809d47f · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Sec-bench: Automated bench- marking of llm agents on real-world software security tasks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2db58369-1310-4a60-83d4-3cbf9dc76cbc · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0686fc54-fca5-421e-95d6-79a9bb105864 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation ee976a74-c2e8-4c40-9e88-6af6a0b9c889 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Besafe-bench: Unveiling behavioral safety risks of situated agents in functional environments
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 6ef0844a-8a1b-4c52-a106-74b3fc8819aa · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 912a0731-be66-4907-ac9f-bd6b05868032 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 955d3ad5-a082-4171-b960-c862ed561857 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments CVE-2026-26322: OpenClaw SSRF vulnerability in gateway tool via unrestricted gatewayUrl parameter
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8e7cb622-bc14-437a-8acc-5c1aaf8ce550 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments CVE-2026-43528: OpenClaw security vulnerability
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 93f3511b-3c72-4da4-8445-113789dfcede · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments CVE: Common vulnerabilities and exposures
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 090fed7d-4139-465e-a357-9c52fa2a8e21 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments PinchBench: An independent benchmark for OpenClaw agent performance on complex real-world tasks.https://pinchbench.com/
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fe37f0e3-d05b-4c4b-9de4-d24f2c0b8871 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments GPT-5.3-Codex system card
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 36b6e734-bfbe-4d30-9147-c1dee5d9bcac · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 35994066-7dd9-4da6-a124-593cc4c88d49 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Known vulnerabilities.https://clawdocs.org/securit y/known-vulnerabilities/
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4b7a497d-b483-471a-8a3d-9174d9177d2c · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments The OpenClaw prompt injection problem: Persistence, tool hijack, and the security boundary that doesn’t exist
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 02d19e5a-b245-4305-ab78-ebbfc86e1139 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Qwen3.6-Plus: Towards real world agents
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 02591845-6c30-4b65-918c-0ab860ccfaf7 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Identifying the risks of lm agents with an lm-emulated sandbox
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0d635b06-c562-462c-8943-ee37d329e3fa · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 15ecae64-40df-4897-a1b1-2bc47fd6673c · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments A Security Analysis of the OpenClaw AI Agent Framework
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 19ec58f4-ac51-41f0-8ae1-07931c2a39b6 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Meta is having trouble with rogue AI agents
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c3d819d0-9c6c-4239-a17e-c7b28bac36f1 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments MITRE ATT&CK: Adversarial tactics, techniques, and common knowledge.https://attack.mitre.org/
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 75c0bbf1-234b-4d75-8b57-1b27d456dd5e · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments ClawSafety: "Safe" LLMs, Unsafe Agents
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation dea6e0f3-4bd3-42da-8494-666d80c8ba0d · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Safetoolbench: Pioneering a prospective benchmark to evaluating tool utilization safety in llms
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 10f89a0e-15b8-4ba6-b48f-04b89abc1c07 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation dcf30c38-2ed2-482b-b851-7d516d2bec5c · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b6df349a-4ce0-4f00-929f-43da6817fedf · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d4a718db-00e6-412c-92e4-3089f24083d9 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c10ab230-deea-4f9e-86f4-1c726cf894d9 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments ClawBench: Can AI Agents Complete Everyday Online Tasks?
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d2f40add-ce62-4dc1-95c7-babccf75ff46 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 98a215d5-998e-4099-b9e3-368a6c3dac27 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments destructive actions
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fc44a2ab-a9c3-4bd1-bc1e-6352418eb984 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0d23b74f-6cd4-47ae-915e-fb617963f808 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Are you sure?
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2e20461e-dec2-4ee2-937d-1a8c6b7a9c87 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Are you sure?
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 48f05cc4-00d7-4f4c-b030-067cc233f7f4 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0b489a9f-1005-4e02-98ca-289d0c0170e8 · outbound
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments https://api.eve.com/v1
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
No inbound Pith citation observations are available.