Pith. sign in

Paper Citation Record · LEDGER

R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2401.10019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.10019 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:52:20.546491Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T02:47:50.885307Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0abc19b7-3461-4e85-b56e-dadc5c37c35a · inbound

LLM Multi-Agent Systems: Challenges and Open Problems cites this paper.

LLM Multi-Agent Systems: Challenges and Open Problems R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:23:05.082632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T10:23:05.028790Z digest=sha256:528ca9b4fcd049c35ac32be8ee8fa727a64947f21708ecb7ab3655eab6a0bc5e

Observation f2c018c3-9266-4a2f-af8f-fe2126da702e · inbound

Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents cites this paper.

Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:36:57.091607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T13:36:57.011451Z digest=sha256:06e84343ac27b3e393d41f6dfd1542505872a157e0a61ff4fb521a9f3c057d88

Observation df6d1152-6f3f-4f2a-94bd-40b6b89500f8 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:57:38.220209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:82bdc711466275a668208cae8dd2e6a10da4b3fd530dd71afb228cfdb46cd1b7

Observation 5d89de49-1012-4e06-be02-7e1de5e1f376 · inbound

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents cites this paper.

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T09:52:20.546491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:52:20.546491Z digest=sha256:a577882a427249293cedd01fa7e8e9323f745dc05888f99f1fd954b49f94dcf6

Observation 3f281932-5d65-49e8-8d7a-b745b7c28f33 · inbound

Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels cites this paper.

Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T07:06:38.144746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:06:38.144746Z digest=sha256:49be6b6d4d18fbda099c8959d41e33fdc450f1d0607a3b60caec60881e2750f2

Observation 6758dd8f-abbc-4e88-b0fa-2696c8c7aa84 · inbound

StepShield: When, Not Whether to Intervene on Rogue Agents cites this paper.

StepShield: When, Not Whether to Intervene on Rogue Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:32.932827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:32.932827Z digest=sha256:f7d44da8d3c48e2ce804e88f6827db47563e91eecf6daf9bbb1ae4881c245620

Observation 6c937a7c-2055-4db7-94e4-fc49830a7256 · inbound

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis cites this paper.

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:18:17.059142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T21:17:15.523370Z digest=sha256:5c85837f66f713de3ab311e451c96d640c4944686f0c813628de3855c0c6dcbe

Observation 520afc4f-e6cc-4ad4-ae8f-3319f22556ad · inbound

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis cites this paper.

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:03:02.674323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T22:02:00.638549Z digest=sha256:9d1b0fa12e0d8714dbf49a28843eb4ab2e33ff43f5522a3d4f4b83fb94ee2546

Observation 4049a741-54cd-4c61-a3b7-8da8564ed31e · inbound

DRAFT: Task Decoupled Latent Reasoning for Agent Safety cites this paper.

DRAFT: Task Decoupled Latent Reasoning for Agent Safety R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:27:09.842266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:23:32.266509Z digest=sha256:fe8befc092cbaac7992a977f80e6be631764e8e9ccaa3be35a91af60447efee3

Observation d2f6e009-fa7a-4f8f-b378-29afc43db615 · inbound

SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement cites this paper.

SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T11:36:41.748710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T11:36:41.748710Z digest=sha256:e47ffe89de0e3a23301eb0058fe3fc0cb3ee08de453c5fb1bd6e00b2b9307e13

Observation bb63e9ed-6dc7-40f1-b67f-caa1942984e8 · inbound

Conversations Risk Detection LLMs in Financial Agents via Multi-Stage Generative Rollout cites this paper.

Conversations Risk Detection LLMs in Financial Agents via Multi-Stage Generative Rollout R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:51:10.933125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:54:41.534985Z digest=sha256:360137b4083d5236117c0ac20e5d569e52dbd7174af8bf93294c543f35412f54

Observation 9c081e7e-5055-46d8-a312-fba893267525 · inbound

Policy-Invisible Violations in LLM-Based Agents cites this paper.

Policy-Invisible Violations in LLM-Based Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:00:59.873412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:20:14.720123Z digest=sha256:ab2d6fe8b38fad37ea9f47dca1fc5c893936c5fc503165d4bb3485e271dcc4f5

Observation db891510-4d5e-4900-b3cb-67eabced0515 · inbound

Beyond Autonomy: A Dynamic Tiered AgentRunner Framework for Governable and Resilient Enterprise AI Execution cites this paper.

Beyond Autonomy: A Dynamic Tiered AgentRunner Framework for Governable and Resilient Enterprise AI Execution R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:27.753401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:49:37.972403Z digest=sha256:3d32ab67b5651ac0764b85b90d0d1a0a95ed7c8a3107749b3fe2c8886a9e83fa

Observation 967c5c0a-2f0d-4fd8-b5e8-bcd54bff5ea2 · inbound

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents cites this paper.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.667888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:13b3ca5194bce5259dfab927f065382da60ebae2431bcda8ef8d6d94635070ba

Observation a97f5a52-8e9e-41a7-9639-4fd1cbb5d36b · inbound

What You Approve Is What Executes: Consent Integrity for Black-Box LLM Agents cites this paper.

What You Approve Is What Executes: Consent Integrity for Black-Box LLM Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:36:22.963107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:12:38.815852Z digest=sha256:7b01f368b731792965163a7ed630669fd7b314314ad3e187328fb3c66c41aec1

Observation adbe4bee-02be-4761-a05e-2c81c0712b0e · inbound

Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification cites this paper.

Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:28.180056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:41:57.594760Z digest=sha256:41988663ecc778e76017213103e7e93c500df7c35858a4198f56b35aeef2db7f

Observation 81136a4f-1ad6-44bb-bb8d-4fe17553ea2c · inbound

RPO-PDT: Demonstrating Role-Play-Based Knowledge Adaptation for Student Support Dialogue (Demonstration System) cites this paper.

RPO-PDT: Demonstrating Role-Play-Based Knowledge Adaptation for Student Support Dialogue (Demonstration System) R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.406892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T16:45:27.352655Z digest=sha256:5a9d41b7e011369a120a9a8f26276831a2b06af760ca3dac76d7af6e1b82e205

Observation 876c6ca4-f5d5-47f6-af89-492a463d9772 · inbound

AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments cites this paper.

AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.460941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T12:53:33.800969Z digest=sha256:315263c8fbd4454719009e36a7f8c231bec8c1745a279d1231b3e742bf39292d

Observation 0c204640-a2d8-418d-a808-64b9884a0547 · inbound

The Gate Is Only as Honest as Its Contracts: ContractGuard for the Contract Layer of Risk-Aware Causal Gating cites this paper.

The Gate Is Only as Honest as Its Contracts: ContractGuard for the Contract Layer of Risk-Aware Causal Gating R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:49:18.277503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T20:58:24.056435Z digest=sha256:bbe96b5d7a01f054e6c0aa4a89dbddb6c38304699237e4662c2752872b8f4fbb

Observation b9648ffe-e43d-478a-aa3f-7c0accafecbe · inbound

ClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent Failures cites this paper.

ClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent Failures R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:38:44.293868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T03:55:09.415657Z digest=sha256:31311db784d7fc661d4b8db20699bccc4e794b05965f7a916b0a07c7db57927b

Observation f47a0459-d913-46bd-9bc7-ab1b30fed566 · inbound

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming cites this paper.

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:09:34.788571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:11:40.088809Z digest=sha256:c14a90e2cf43de34f613b2445377b74627a6b8d151c76d69c55dd3ce4c1b8980

Observation aa6f9bec-9208-4351-be60-225610dcf38e · inbound

The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities cites this paper.

The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T02:47:50.905903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T02:41:24.813416Z digest=sha256:04c9b16885cd4651612caebd52dedba4ff844d464623ffe5e9b912d6c51a6f1a

Observation 2fdfbdb3-23a1-43e4-9190-4e1caab379fc · inbound

SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents cites this paper.

SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T03:05:54.867875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T03:01:02.743557Z digest=sha256:56f1fe355da59311750741cdb11062885f0c4b49bb8fcce4f4c6ef0f93a2bdfa

Observation 56127026-aa2a-44e8-86c7-485e89c9fef1 · inbound

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks cites this paper.

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T06:49:30.678939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:49:30.678939Z digest=sha256:10abd033d1f49b59b5806e25658f651669f452ca4a0df3c7fcb7f8d351167d65

Observation e2638466-da69-4762-afad-f4a6fa0209d3 · inbound

Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines cites this paper.

Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T12:39:58.329709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:39:58.329709Z digest=sha256:1b907337ea07b9875c6d2729c233b35acb96e36df948ad72c93dbd43fb7eb9ec

Observation 06b0426c-f018-45aa-ac6d-7537ad121c6b · inbound

How Benchmarks Mis-Score Computer-Use Agents cites this paper.

How Benchmarks Mis-Score Computer-Use Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T09:36:43.813643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T09:36:43.813643Z digest=sha256:41ae192f264139aaab8baad6e5abaa5a101812b8cb1472ac39e46581bb17c2c7