Pith. sign in

Paper Citation Record · LEDGER

SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2406.14598.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.14598 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:55:56.585528Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cf7b16cb-dcf4-49a5-bcad-a26f9edf9c18 · inbound

Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents cites this paper.

Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 148

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T13:36:57.194715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T13:36:57.011451Z digest=sha256:434d3c462a9bfff9728146596bd158224b1f47d50f05a53888bee880631f8b2c

Observation ee8055b1-1b77-4894-a888-15c2f0e1eb1d · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.166011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:197ae9033e2338189d940d58ae2a93ff91a716560cb8576646a589c4449727ea

Observation c0276105-7099-4721-b601-143ac667269f · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 258

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.739874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:2a2d5dae8cb9268c9e3554a1c5cd29c80d3cf4e68b249aaa457517547d95f7df

Observation 5fb4048d-bc8e-49e2-b7ce-73c6c5f11954 · inbound

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models cites this paper.

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T19:55:56.585528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:55:56.585528Z digest=sha256:5fc1eaa9f1cad472d92f8b9675c7c1474126dc5ac4668ce29f11132349029f49

Observation 974542d6-0b78-4b71-92dd-5c58fadb3373 · inbound

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal cites this paper.

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:56:45.857797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T18:56:13.680353Z digest=sha256:d6dce6e12d4920a04d3d425d4d0f33a6ff7c9031dd7c57e48360b661796867c1

Observation 158347e9-6d12-4fe2-bbca-df8c4297d188 · inbound

MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation cites this paper.

MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T10:22:08.785775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:22:08.785775Z digest=sha256:7eb0c61fa6b5a5e97cdb325746dd8c89d883f8050ec7fbd398d08700ce99f69d

Observation 99ab51e0-3848-44b1-95eb-7a933152af0c · inbound

$C$-$\Delta\Theta$: Circuit-Restricted Weight Arithmetic for Selective Refusal cites this paper.

$C$-$\Delta\Theta$: Circuit-Restricted Weight Arithmetic for Selective Refusal SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:40:20.100340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:40:20.100340Z digest=sha256:69c67e458399b834c25b968d20aa74d9e8702ac494119eda92b50372a0ba28cc

Observation 9454a312-c9dd-4365-8b21-bcf958eeea5b · inbound

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence cites this paper.

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T17:19:39.109017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:19:39.109017Z digest=sha256:5678ec660f633b7ea4877ba871970d71d61b13ee22e73f3e84eb930301d52824

Observation 8ab7dbaa-68b6-48f4-87f7-35e96e9886b8 · inbound

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules cites this paper.

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:28:09.987310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T19:24:54.381722Z digest=sha256:b07a60e5f6407681cca2e170930998290952bbb848feb2b2e583adf15c432099

Observation 3deb2e8f-27cb-400c-be3d-2239f6a4af34 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:52.512867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T18:25:53.037936Z digest=sha256:77b17feb03b572b39f3cb81b1d7f3aa2fc368a8ec596c4e825e759d9a84cd33b

Observation cb51319b-d4d7-4e55-81f7-3331eb471a22 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T00:19:33.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:19:33.861692Z digest=sha256:b66d44b1a15d093576df4a2e9dcd57808743325c9c2d7313973ed5f79ab0c026

Observation 0b580803-0f28-4bb6-9bac-a812f763645e · inbound

VoxSafeBench: Not Just What Is Said, but Who, How, and Where cites this paper.

VoxSafeBench: Not Just What Is Said, but Who, How, and Where SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:24:22.069363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T10:19:28.041282Z digest=sha256:470e52f246b9031752ee7377149d45326b6305cb37fd10cf92499d61e86ff171

Observation a051cc05-28bc-43f9-a159-250633e5e6f4 · inbound

Reasoning Structure Matters for Safety Alignment of Reasoning Models cites this paper.

Reasoning Structure Matters for Safety Alignment of Reasoning Models SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:56:03.984376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T02:36:57.093584Z digest=sha256:76e37a7d0aaa17d0fdcdf004f8f306f699c9bd4ec1225f8b831ca8ff4012a023

Observation 66f106a5-f58d-4a15-b4e7-77231dad9779 · inbound

FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios cites this paper.

FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:51:43.870558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T19:06:29.779618Z digest=sha256:689fc9999065e70a199b5aa27dff8943f15ae9aa8bc8294fa7b909cb97cd0b4f

Observation 023a6b59-0ace-4f14-b849-07e9ce053959 · inbound

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs cites this paper.

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:42.601099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T19:09:07.557773Z digest=sha256:5962c3d7d364d4fb7b70ab3938ef06f6b7a213d66b43fa1667d1fd5817cec23a

Observation 6a9c689d-4263-40dd-a3d4-7dd306a4c755 · inbound

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety cites this paper.

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:31:00.988350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T16:00:32.413225Z digest=sha256:f8970d9f75d8497fac10a4ab3e8455c059a128e74d3ba2b4ac84ce2ba3d73d19

Observation b3863060-2459-4805-a883-3af81d3d0a8d · inbound

Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses cites this paper.

Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:46:31.235979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T02:16:49.785596Z digest=sha256:d801793c970dba2155355f1260dd51433704b035c67e3c15dd4998df8583dddf

Observation 4fae6c5d-e014-4d68-9b6d-429bcb2548c2 · inbound

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts cites this paper.

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:43.579732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:11:29.066362Z digest=sha256:db1e50b8c614d01fe38502a1d693642cbc36d2d21dfb845422c02cc853c7b534

Observation adbb2505-209d-4310-bee7-3bd379216a5a · inbound

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization cites this paper.

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:45:44.083441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T18:08:16.051117Z digest=sha256:5788c77aea0b87c27ca414954a58164be62fce95080221015d0f62b99e393b92

Observation 98678dea-e9e8-42c1-8fb3-106423668e54 · inbound

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs cites this paper.

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:07:27.020952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T07:06:46.387088Z digest=sha256:6bb7d92516c1597b7e67d53a73865bbf28768ec5b342c3a3cd0c1965b001dc5b

Observation 92bacfc6-ff43-4ce3-88c4-e004278317c7 · inbound

Before the Last Token: Diagnosing Final-Token Safety Probe Failures cites this paper.

Before the Last Token: Diagnosing Final-Token Safety Probe Failures SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:38:00.077747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:37:02.439454Z digest=sha256:733fabe61cb919ebc37a407f4e4226ca5e76dcc83ceb100471cea7d92739720f

Observation 65ef1288-e9c1-4c3e-8f3d-83fcc8138e28 · inbound

LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs cites this paper.

LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:22:55.720651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T20:20:04.306202Z digest=sha256:9e67a93bb5ee6c9379c449dc40d1b0ba1653fdbddbc3e2f5e8bc448400c39142

Observation 6b539532-c25e-4445-ad80-74ef7a44f0c4 · inbound

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts cites this paper.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:51.907305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:8add8f9b4b4b3ac064290bf701d7e6cd8341104cb72cd88cda8d354d7e44966f

Observation 3bc4c532-d970-4940-bee3-b93c29046270 · inbound

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content cites this paper.

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T09:13:15.944664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T09:11:58.843585Z digest=sha256:027d606b9a526cb73b735e41a650e37b468e083eda9aac6c79cca67740cae7f5

Observation 2d07b2ba-8fc4-4367-aca1-c377a5a35717 · inbound

Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges cites this paper.

Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:47:30.681311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T17:01:03.040933Z digest=sha256:52d9df8014527ef0d30e5d6881127c4ae5df3cfa02a08a3faa051c23df281c36

Observation 0a7864e0-598b-4367-b31d-4ca3a104b58f · inbound

Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts cites this paper.

Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:57:38.388352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T13:32:18.368158Z digest=sha256:0e50ec1e323030152ba40c6e12c87b0ea97391c4756c6c79001020c623990bf4

Observation f3b4a5b4-c87c-4933-9cfb-bd5cf21446bb · inbound

What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations? cites this paper.

What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations? SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:39:31.052108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:43:17.000740Z digest=sha256:c5c60e733388de167648a89f42de9e3c71524e09c994b21eabf0c154846aa615

Observation 50313c44-6044-420c-8078-5018a7021242 · inbound

Efficient Safety Benchmarking via Item Response Theory cites this paper.

Efficient Safety Benchmarking via Item Response Theory SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:55:48.976540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T15:51:00.484343Z digest=sha256:3f6f4de643a2bcc7a53a1677c3f223dbce4bcddc0522cbe4f6d094569cd85574

Observation b6f120d8-6269-4093-8125-48c31a9e5266 · inbound

Discriminatory Compliance: How LLMs Answer Queries from Protected Groups cites this paper.

Discriminatory Compliance: How LLMs Answer Queries from Protected Groups SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-06-26T12:59:29.405506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T12:58:09.227648Z digest=sha256:21984e79e9c9c3411d5ef477b09f4b433d8684b2b67c6ddc4d491a91eb859acb

Observation b860b720-92ce-4383-9e4d-59ea998a0212 · inbound

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models cites this paper.

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:09:59.500402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T23:52:02.327522Z digest=sha256:41eeebf5ffcd81896a64b4a1974c3897054cc6adf52725a79e8f35d4ddfac883

Observation bcc533d1-39a6-4439-9ad0-f61a47f1db0b · inbound

Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation cites this paper.

Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:07.985799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T20:55:44.549142Z digest=sha256:12c98b211b0949dbf2df70abffad5cb738ede1b675fdffe695d942721d318a69

Observation 19a1a96f-8173-463b-ac50-a83ed3eff9b3 · inbound

Agentic Abstention: Do Agents Know When to Stop Instead of Act? cites this paper.

Agentic Abstention: Do Agents Know When to Stop Instead of Act? SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:54:34.380647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T09:54:07.138157Z digest=sha256:0f115a75726b9287b7dd436eb34ae88b900ec86efaea04b9cfad856062f21479

Observation 4816bf2f-87a9-4924-8233-5347664420a4 · inbound

Addressing Over-Refusal in LLMs with Competing Rewards cites this paper.

Addressing Over-Refusal in LLMs with Competing Rewards SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:29.223049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T06:59:12.695984Z digest=sha256:34b0d5be61d008090d1f80117a9c8660f365e35ec63dff803d183759f68a3e29

Observation 0aeee549-91e6-40e4-a8b4-d131e713f3eb · inbound

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale cites this paper.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:8dabd5f138df039649001c0723cb1b42e63f0a8382da1c56c44db8ca6a256e46

Observation 2579c87e-8a8f-4c25-bc4e-2702148f01ed · inbound

BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation cites this paper.

BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T02:00:44.842416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:00:44.842416Z digest=sha256:3e609c81670e92fb2de798741c13cc0aba63d52d5e2491ace20b9863914e46e3

Observation 0d12b4db-f03f-48e6-83eb-facc8d89715b · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.199932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.199932Z digest=sha256:8d16f138b2abb9c37530cd686e58672aab63de8b27ceca67404c9da1c0cb6c03

Observation 1805f2bd-03c7-4e0c-856d-5a072c199b0c · inbound

Fence: Specialized SLM Guardrails for LLM Applications cites this paper.

Fence: Specialized SLM Guardrails for LLM Applications SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T13:21:55.083013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:21:55.083013Z digest=sha256:07e8ba7ce066e4b536036c5867a5831ddc2eb09bb3a2d493968daa58c0782f7f

Observation 9066c1b2-ab38-40eb-9c4c-92a390a3db51 · inbound

Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering cites this paper.

Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T02:23:11.890673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:23:11.890673Z digest=sha256:d4792755bc29fe776ed8f0e80bf61895b8b481344d45ca08bc0477937245ed21

Observation bb055c66-c166-43b2-9b36-00df8cf1e36f · inbound

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications cites this paper.

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T02:04:06.864702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:04:06.864702Z digest=sha256:56f0683affa1fa590ed6a0507bbca792397ab96a019c01de9f8820610dff524b

Observation b050bf87-bc25-450d-993d-c098885d63cf · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:22.240966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:22.240966Z digest=sha256:4f13bcc051de37ecfc8fd4ec43f3bf2024699cc9f94f067d2d365d7fd5ab5e62