Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2406.14598.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:55:56.585528Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation cf7b16cb-dcf4-49a5-bcad-a26f9edf9c18 · inbound
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 148
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ee8055b1-1b77-4894-a888-15c2f0e1eb1d · inbound
A Survey on LLM-as-a-Judge SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c0276105-7099-4721-b601-143ac667269f · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 258
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5fb4048d-bc8e-49e2-b7ce-73c6c5f11954 · inbound
YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 974542d6-0b78-4b71-92dd-5c58fadb3373 · inbound
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 158347e9-6d12-4fe2-bbca-df8c4297d188 · inbound
MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ab51e0-3848-44b1-95eb-7a933152af0c · inbound
$C$-$\Delta\Theta$: Circuit-Restricted Weight Arithmetic for Selective Refusal SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9454a312-c9dd-4365-8b21-bcf958eeea5b · inbound
MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab7dbaa-68b6-48f4-87f7-35e96e9886b8 · inbound
Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3deb2e8f-27cb-400c-be3d-2239f6a4af34 · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb51319b-d4d7-4e55-81f7-3331eb471a22 · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b580803-0f28-4bb6-9bac-a812f763645e · inbound
VoxSafeBench: Not Just What Is Said, but Who, How, and Where SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a051cc05-28bc-43f9-a159-250633e5e6f4 · inbound
Reasoning Structure Matters for Safety Alignment of Reasoning Models SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 66f106a5-f58d-4a15-b4e7-77231dad9779 · inbound
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 023a6b59-0ace-4f14-b849-07e9ce053959 · inbound
PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6a9c689d-4263-40dd-a3d4-7dd306a4c755 · inbound
MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3863060-2459-4805-a883-3af81d3d0a8d · inbound
Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4fae6c5d-e014-4d68-9b6d-429bcb2548c2 · inbound
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation adbb2505-209d-4310-bee7-3bd379216a5a · inbound
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98678dea-e9e8-42c1-8fb3-106423668e54 · inbound
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92bacfc6-ff43-4ce3-88c4-e004278317c7 · inbound
Before the Last Token: Diagnosing Final-Token Safety Probe Failures SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65ef1288-e9c1-4c3e-8f3d-83fcc8138e28 · inbound
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6b539532-c25e-4445-ad80-74ef7a44f0c4 · inbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3bc4c532-d970-4940-bee3-b93c29046270 · inbound
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2d07b2ba-8fc4-4367-aca1-c377a5a35717 · inbound
Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a7864e0-598b-4367-b31d-4ca3a104b58f · inbound
Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3b4a5b4-c87c-4933-9cfb-bd5cf21446bb · inbound
What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations? SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50313c44-6044-420c-8078-5018a7021242 · inbound
Efficient Safety Benchmarking via Item Response Theory SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6f120d8-6269-4093-8125-48c31a9e5266 · inbound
Discriminatory Compliance: How LLMs Answer Queries from Protected Groups SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b860b720-92ce-4383-9e4d-59ea998a0212 · inbound
PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bcc533d1-39a6-4439-9ad0-f61a47f1db0b · inbound
Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 19a1a96f-8173-463b-ac50-a83ed3eff9b3 · inbound
Agentic Abstention: Do Agents Know When to Stop Instead of Act? SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4816bf2f-87a9-4924-8233-5347664420a4 · inbound
Addressing Over-Refusal in LLMs with Competing Rewards SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0aeee549-91e6-40e4-a8b4-d131e713f3eb · inbound
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2579c87e-8a8f-4c25-bc4e-2702148f01ed · inbound
BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d12b4db-f03f-48e6-83eb-facc8d89715b · inbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1805f2bd-03c7-4e0c-856d-5a072c199b0c · inbound
Fence: Specialized SLM Guardrails for LLM Applications SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9066c1b2-ab38-40eb-9c4c-92a390a3db51 · inbound
Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb055c66-c166-43b2-9b36-00df8cf1e36f · inbound
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b050bf87-bc25-450d-993d-c098885d63cf · inbound
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.