Pith. sign in

Paper Citation Record · LEDGER

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components

As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.02357.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02357 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:22.498212Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c731d977-eaff-4ee6-a600-57b1d766d91f · outbound

This paper cites and Vichy, L.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components and Vichy, L

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:23.601617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:28:20.088809Z digest=sha256:3223c654e6452e56ffaa8175ef7dcbd00d20df142274d462a4980e576fd0e8c8

Observation e01a18ed-85a8-4d31-90b3-c9742917d96d · outbound

This paper cites AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:28:22.875538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:28:20.217361Z digest=sha256:843cffa6e31d8aefb9cf2f8f34a488e174e81bdb34ae38b20bb2c0df0e7e3d25

Observation 5c6c9aba-0b6c-4ebc-aeb0-3687fcedd11e · outbound

This paper cites S., and Terry, J.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components S., and Terry, J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:23.418350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:28:20.369273Z digest=sha256:62ca60c158ff2d9a46d1e7d25d84389214407210b36ef51ccf1372ad983b4918

Observation bfae4b35-f4a7-4ab0-b4be-56ad08290888 · outbound

This paper cites and Jaffer, I.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components and Jaffer, I

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:23.213872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:28:20.485938Z digest=sha256:9aecf2e850339e766e3f14b604ceaf442f845ebd0d8cb7096b127bc7d4c2192b

Observation d92a8644-8493-4c6b-99b3-ccb8f4a3cd73 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:20.578837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:20.578837Z digest=sha256:ddf730e5a8c5ad5880980587b9ae386454a285f5daf326f6547d3bed84c06d15

Observation d99caefb-5ec6-40dc-bab6-a222343171bf · outbound

This paper cites FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:20.709513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:20.709513Z digest=sha256:0e267bb9225fc4c89c4f166573462e2508fafc2ee842faab1990b438b2632013

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:20.868221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:20.868221Z digest=sha256:8ded3623d77cd665777fafa901e25dcb506647a2c5d210df91d5a87f1e67b2d7

Observation 1150a597-0e19-4efa-955e-c9ea5c03a4b7 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components AgentBench: Evaluating LLMs as Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:20.990736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:20.990736Z digest=sha256:7a42b21f74fdc3e3db877de3d7a0c5ed1a7dd2f1cc9180cf01f299ed0bcdbf0d

Observation 8e0d26c6-80e1-46c1-9551-0cc53537e299 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components The Alignment Problem from a Deep Learning Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.158123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.158123Z digest=sha256:b7f6620eb35900f9020de4c09c84f925dc3698379d8782ff8adf3eb22e2e38d3

Observation 2cb1cb4b-00b5-4c05-91ac-66e1b4859e42 · outbound

This paper cites S., O'Brien, J., Cai, C.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components S., O'Brien, J., Cai, C

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.294656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.294656Z digest=sha256:f3c9c5ba093b9b9f1cede37f345cf5ec138e08cd6925ba0fa646b8f62874a2e3

Observation bba966d5-1688-442e-be81-61fac0c0b936 · outbound

This paper cites Open Problems in Technical AI Governance.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Open Problems in Technical AI Governance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.484300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.484300Z digest=sha256:08918fd5d51fb6c0a9a90e77572402c00f2f18aae51fbfb59e3da50356b64665

Observation 24a86f83-83d2-4712-9042-e6e893aa06df · outbound

This paper cites Model evaluation for extreme risks.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Model evaluation for extreme risks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.649808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.649808Z digest=sha256:dc3f841f66a401519f6007fd45eeb54ae644d5a98fae597e36da0cb7405a8ba8

Observation f64e927d-60da-4507-910d-227805fa25bd · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.793229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.793229Z digest=sha256:19147ff6e025648e4901af2b071c53d31f9872a1cf40a611a1d7d6a124b4ed8a

Observation b9384660-b1cd-4cb1-aed0-a026cb12bb6d · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.929672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.929672Z digest=sha256:6041de37ac07ac5e10f37265b4d9df1ca2034add8a4a6a8e560d042a96aa3954

Observation 54b5fbf1-be2d-46ee-8c4e-19235f995d69 · outbound

This paper cites Benchmarking Complex Instruction - Following with Multiple Constraints Composition.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Benchmarking Complex Instruction - Following with Multiple Constraints Composition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:23.013534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:28:22.121414Z digest=sha256:4fc6dc6b006ef24d69e4c07c987073301207fca71171a0dec93ba0196ee45058

Observation 41bb9b5d-f5f4-49f8-af4f-d7e22349bff6 · outbound

This paper cites S., Shah, A., and Tellex, S.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components S., Shah, A., and Tellex, S

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:22.287533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:22.287533Z digest=sha256:5d268bbb7e20d00983ae93df1f9b2ec9a4dcddc1dcf56cf73e0a8d1319e2c627

Observation 0cc63e0e-1775-43ac-bf38-d6870c090c41 · outbound

This paper cites InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:22.414764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:22.414764Z digest=sha256:19653a5fd07034ece0ecbd30e16d43c878c9306651d580a7e146de4c69855687

Observation 6b8009c5-20f0-4790-96fa-03d4034b8ec4 · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:22.489474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:22.489474Z digest=sha256:f8abc97aa951869ae74fa11ee6957d32a049d25b54e861a9de21a1dfa0119fdb

Observation bc09d1d5-80d4-422a-a37a-b0276b53b150 · outbound

This paper cites write newline.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components write newline

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:22.498212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:22.498212Z digest=sha256:db8ab8d1fccf6396b9930821590a7b4d57e0b295cec0d61e509a78b20b948661

Pith citing papers

No inbound Pith citation observations are available.