Pith. sign in

Paper Citation Record · LEDGER

SafetyBench: Evaluating the Safety of Large Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 58 inbound Pith citation observations for arXiv:2309.07045.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.07045 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 58 of 58 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:30:35.326756Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

19
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9af1f249-4f1a-4f1d-b322-b462cff38150 · inbound

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools cites this paper.

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools SafetyBench: Evaluating the Safety of Large Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:08:09.724948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T08:08:09.444352Z digest=sha256:cafa2b308fae4c45d7de264825016dcc30bf1bd1f0beeeb14de27036acbc1953

Observation fffc98aa-9840-491a-89b7-8920516259bb · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey SafetyBench: Evaluating the Safety of Large Language Models

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.493758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:b46bb24ee5e2691e29c9d80084be688b2e44f780d7eef2bf2f003eaac6b7b86a

Observation 484257d8-a543-48f8-8f52-60cdaa5fde61 · inbound

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain cites this paper.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain SafetyBench: Evaluating the Safety of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.612898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.612898Z digest=sha256:fd7c61ee32892cea3a90e0b24b0e3edbb9c2080a81f349fc6c6b4c55958de473

Observation 5cf03c89-189d-4d19-9b85-cb257f81d139 · inbound

Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models cites this paper.

Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models SafetyBench: Evaluating the Safety of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:40:22.183979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:40:22.183979Z digest=sha256:44c0d2036fd9e86755e79c84443709ff2e31d4e751f9bdc5477deb72fd370e40

Observation d48f5628-26f5-4b15-9151-28a563e9a88b · inbound

Usage Governance Advisor: From Intent to AI Governance cites this paper.

Usage Governance Advisor: From Intent to AI Governance SafetyBench: Evaluating the Safety of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T00:05:28.217269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:05:28.217269Z digest=sha256:800fe41676471ade00bab6256ef33b7aaf60dfedd3b18e1191fa08e5e2a0cb44

Observation 67a91d28-e450-48ff-8ae4-d393bde4d4db · inbound

SafeWorld: Geo-Diverse Safety Alignment cites this paper.

SafeWorld: Geo-Diverse Safety Alignment SafetyBench: Evaluating the Safety of Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T19:41:42.296751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:41:42.296751Z digest=sha256:c9c6983452b27427d06f68933f75a627287e386436aefa9a467ca987c6e6acad

Observation 090df110-1039-4efd-a3e2-b1b391f763eb · inbound

Large Action Models: From Inception to Implementation cites this paper.

Large Action Models: From Inception to Implementation SafetyBench: Evaluating the Safety of Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T16:29:57.048566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:29:57.048566Z digest=sha256:1838bc9808c7365e6b58cef47d8c518fe8d49ac788cc452c305c748ef3ac5b5b

Observation 8856b602-8ee2-4cb0-9932-312bbddc852d · inbound

Observing Micromotives and Macrobehavior of Large Language Models cites this paper.

Observing Micromotives and Macrobehavior of Large Language Models SafetyBench: Evaluating the Safety of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T18:24:05.980853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:24:05.980853Z digest=sha256:3950a89cc70c01cd32961fe05fef9c79aa9d76d6cf3107b26d54bf1115a1368c

Observation 32945e92-6f28-441e-affa-60ddc0015473 · inbound

Unanswerability Evaluation for Retrieval Augmented Generation cites this paper.

Unanswerability Evaluation for Retrieval Augmented Generation SafetyBench: Evaluating the Safety of Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T14:17:45.755457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:17:45.755457Z digest=sha256:f9428e629d87de2a0a5dfa0ed0df929d6051299459af4dcc17ae5db620140712

Observation 2ee2dd48-4fdd-42c4-8a65-8b70f6609029 · inbound

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models cites this paper.

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models SafetyBench: Evaluating the Safety of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:52.285186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:52.285186Z digest=sha256:e0c9e3a9f8ea46ba1977e379c4deea83d2ceb302c30b69b793dddf224c2f3e6d

Observation a82aabb7-4b40-4daa-a556-e8e0d0903a6b · inbound

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense cites this paper.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense SafetyBench: Evaluating the Safety of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.710337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.710337Z digest=sha256:80d5c77346c75466f31be34c73895f8fd03977e1cbbd03afbdea076ec32f0903

Observation b4df0278-4455-45e0-9d3b-66eac63bc6c5 · inbound

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values cites this paper.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values SafetyBench: Evaluating the Safety of Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.906558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.906558Z digest=sha256:0da4b7c4447173baf0bb2b572694b128c2f774de183cd1ddc24605caab68d2d3

Observation d2badddb-fd5d-4b8c-bf4b-474b0047ce56 · inbound

A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy cites this paper.

A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy SafetyBench: Evaluating the Safety of Large Language Models

Reference 243

Resolution
unresolved
no resolver link, observed 2026-08-10T20:05:12.809379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:05:12.809379Z digest=sha256:9d530b8c21d4f6342f51ef1662b712659ef262075e2468b0b817568a8357df7b

Observation 662bdd8f-9230-439a-8a93-9bda8a182b86 · inbound

The Dual-use Dilemma in LLMs: Do Empowering Ethical Capacities Make a Degraded Utility? cites this paper.

The Dual-use Dilemma in LLMs: Do Empowering Ethical Capacities Make a Degraded Utility? SafetyBench: Evaluating the Safety of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T18:31:03.503934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:31:03.503934Z digest=sha256:d8f3fd9ddd0bcd28714474a255f5c4be8ab5e27337e3924906eec54691ae823e

Observation 2c31a27d-bbb8-49fb-8920-65d290d457a6 · inbound

Global Perspectives of AI Risks and Harms: Analyzing the Negative Impacts of AI Technologies as Prioritized by News Media cites this paper.

Global Perspectives of AI Risks and Harms: Analyzing the Negative Impacts of AI Technologies as Prioritized by News Media SafetyBench: Evaluating the Safety of Large Language Models

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:20.516102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:20.516102Z digest=sha256:a5c7eb1e61b3f374f353ecab8f61b05e88b5f6d8b952da90c7ad92648228399e

Observation a44391cd-317e-447f-a0d8-44b23ac6a679 · inbound

Data-adaptive Safety Rules for Training Reward Models cites this paper.

Data-adaptive Safety Rules for Training Reward Models SafetyBench: Evaluating the Safety of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.211691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.211691Z digest=sha256:536f472749cf9c8b0b31579fcfc04a7d7a376267c334242c678478491a18d69f

Observation 10b0be7e-f833-4089-a919-4ac79cdbfcff · inbound

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation cites this paper.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation SafetyBench: Evaluating the Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.952819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.952819Z digest=sha256:b65d8d78141c134069f6c4a9720583ad5f6b8445296c5fc08377136901fad2c0

Observation 144eaf94-db3e-4dbb-a7c2-81c034d99905 · inbound

o3-mini vs DeepSeek-R1: Which One is Safer? cites this paper.

o3-mini vs DeepSeek-R1: Which One is Safer? SafetyBench: Evaluating the Safety of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T23:32:16.759645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:32:16.759645Z digest=sha256:282636421473377b27044955cf1b045c24aa7c14eb6938bd736c3668ef52cf3d

Observation 4b353e2b-e070-4a16-860a-ec0b2ebac3a2 · inbound

How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers cites this paper.

How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers SafetyBench: Evaluating the Safety of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:12.395417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:12.395417Z digest=sha256:cfb6fd222ea7e74234632a2fd66c5fe915d6c029a7f297bc0f7cdec4a4448bd7

Observation ae8ba6a1-a12d-43ae-8844-28b66c412ca0 · inbound

ELAB: Extensive LLM Alignment Benchmark in Persian Language cites this paper.

ELAB: Extensive LLM Alignment Benchmark in Persian Language SafetyBench: Evaluating the Safety of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:30:35.326756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:30:35.326756Z digest=sha256:02b9881191287519de64bcb16b3a45a4c7e6036baf433dd44e278552194430ce

Observation 9c51b3f1-abd1-40af-bdb8-059fd1d5938f · inbound

DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain cites this paper.

DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain SafetyBench: Evaluating the Safety of Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T12:04:24.413358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:04:24.413358Z digest=sha256:dc07cf0ad13d26ebd4699f52741f02779a9d6fdc46bdc8c6db8ec6a8c705256d

Observation dd57880a-7a84-4a34-9a16-6b4f19f005ab · inbound

Tinkering Against Scaling cites this paper.

Tinkering Against Scaling SafetyBench: Evaluating the Safety of Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:03:31.858045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:03:31.858045Z digest=sha256:fe65370a5115040216c33b10a3068754d9c95495f8cd6c73a9fc2fd3b2f96f42

Observation fa7d71ed-2fc2-4077-9b08-806c3c3ef1b1 · inbound

Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers cites this paper.

Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers SafetyBench: Evaluating the Safety of Large Language Models

Reference 175

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:35.643591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:35.643591Z digest=sha256:e6c1c967410f883850897d59ebd0622f31fb5990223eaa3628948ec788fedd1d

Observation b0ddeb8c-676e-4f02-a3ba-96630d077b1b · inbound

GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection cites this paper.

GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection SafetyBench: Evaluating the Safety of Large Language Models

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:39.973804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:39.973804Z digest=sha256:3f16fc9547e703e3946647dd064db2d2cbd4340f114ab6d03e04b5a89e18cc15

Observation a444735e-47da-4991-9cee-16ece478ec49 · inbound

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use cites this paper.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use SafetyBench: Evaluating the Safety of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:31.148227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:31.148227Z digest=sha256:e187ab273a32d24a31a700a9fb51bf8d4ae61d88e194504abc5ec5029c2eaa64

Observation 61a96958-5215-45cd-a6be-9a719353b435 · inbound

Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models? cites this paper.

Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models? SafetyBench: Evaluating the Safety of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:42.777127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:42.777127Z digest=sha256:2df0545fcfd9efd4fd712e016e9ac79976cbc98b4ce2f636c5a828e6cb1e1c99

Observation fad597c3-b1dd-45f1-b0ff-93fac7ab0d79 · inbound

LLM-based HSE Compliance Assessment: Benchmark, Performance, and Advancements cites this paper.

LLM-based HSE Compliance Assessment: Benchmark, Performance, and Advancements SafetyBench: Evaluating the Safety of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.091799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.091799Z digest=sha256:2b90c2aee888512c50132fc7df5ee871f57da7f36a9ec4351c1f80e24acea1d7

Observation fe0f0a71-571a-48fb-be9b-dc635c57f7cb · inbound

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models cites this paper.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models SafetyBench: Evaluating the Safety of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:16.805123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:16.805123Z digest=sha256:29dd773077377691e9d42be8d9e48dd28227763b42c18fc3784476bc389e410b

Observation 8d2f4893-47b3-4137-af3c-f0aa5f4dec6e · inbound

Large Language Models Often Know When They Are Being Evaluated cites this paper.

Large Language Models Often Know When They Are Being Evaluated SafetyBench: Evaluating the Safety of Large Language Models

Reference 3

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:16:15.735454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.735454Z digest=sha256:2f424cdeca04c40421d785c9c7899c2422c6cd47f330963a4a66e218603d39d5

Observation e8ab284a-3778-40f8-9bf3-8bd0b1e2834d · inbound

MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine cites this paper.

MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine SafetyBench: Evaluating the Safety of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:54.874110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:54.874110Z digest=sha256:2ccd687e6d16ad7bdcfd4a9a877a0756e56b9177bc46ea48730cab1f32952808

Observation a58398f4-79b8-4b32-8d86-a01a50d3b514 · inbound

SafeCoT: Improving VLM Safety with Minimal Reasoning cites this paper.

SafeCoT: Improving VLM Safety with Minimal Reasoning SafetyBench: Evaluating the Safety of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:02.714368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:02.714368Z digest=sha256:7105988f76518ed6977a6e9548b8ee766f524a4d3667ae7ad4ac5c1ee7f752fa

Observation 5152cbc0-bd5b-4048-866a-a5f5ef1b9b59 · inbound

PL-Guard: Benchmarking Language Model Safety for Polish cites this paper.

PL-Guard: Benchmarking Language Model Safety for Polish SafetyBench: Evaluating the Safety of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:54.958475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:54.958475Z digest=sha256:7c39073e6927280574157f6c6e513c2014e94af6ce4d3750efb743c604de4a00

Observation ff373dd7-bf4a-417b-ab03-8b441f3d62af · inbound

Informing AI Risk Assessment with News Media: Analyzing National and Political Variation in the Coverage of AI Risks cites this paper.

Informing AI Risk Assessment with News Media: Analyzing National and Political Variation in the Coverage of AI Risks SafetyBench: Evaluating the Safety of Large Language Models

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-08-06T10:31:06.665914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:31:06.665914Z digest=sha256:a175a233899e95c445795b130543007bc86fb6a1d05a67a7a46b802b47d6f8c4

Observation 99b6e058-53b0-43ab-ba9a-6c69bde88bcc · inbound

Observation of momentum dependent charge density wave gap in EuTe4 cites this paper.

Observation of momentum dependent charge density wave gap in EuTe4 SafetyBench: Evaluating the Safety of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T22:45:17.057354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:45:17.057354Z digest=sha256:296ad396ddb579324c2eedc03de70859153b4da976ffe1204de87988daafb739

Observation 20151091-1b09-439b-8963-3e7fdbb5a3cc · inbound

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models cites this paper.

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models SafetyBench: Evaluating the Safety of Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:50:08.558363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T17:50:08.399160Z digest=sha256:96bbfb94596a0a343a08250e24a860b518b2374b3814bc9e05e1d899152ebf7d

Observation 13039f77-573b-488c-935f-c28a25313167 · inbound

Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models cites this paper.

Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models SafetyBench: Evaluating the Safety of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T05:29:20.767922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:29:20.767922Z digest=sha256:dd65317fcb4b75b3e80a858173da1861f6f520db0dfea5c954845a4161eb1c05

Observation 0b4b09e4-fb46-4bf9-9358-e27503e3e903 · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm SafetyBench: Evaluating the Safety of Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:25.711468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:25.711468Z digest=sha256:7df410462650f90c71fa9275457a205c3c983567ef044d25160532ff35129106

Observation 4d32f22f-ce66-49cd-b9fd-ac8d1e2ec432 · inbound

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models cites this paper.

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models SafetyBench: Evaluating the Safety of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T19:55:56.607844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:55:56.607844Z digest=sha256:088d39eb9ecde30c6e4ff45765d05a070a960f19c14a4b533ff6198d84ac95cb

Observation 31ff1476-4f45-4a69-b5f3-fc965d7af2ae · inbound

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting cites this paper.

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting SafetyBench: Evaluating the Safety of Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T12:45:26.981751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:45:26.981751Z digest=sha256:d663c3961800ecdd56274fc00dd7cff437cf6dda82ac7d7d5dece9c1370b7318

Observation 29ed8296-e9a6-4df2-a9b4-323305286cc8 · inbound

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization cites this paper.

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization SafetyBench: Evaluating the Safety of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:58.482981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:58.482981Z digest=sha256:68b296da4cac7ebef239bd1893593adca62f827fbe2be662ce5675b5a62b4cf5

Observation a7f5c285-c784-4bbb-ace5-834a5f3a7e3e · inbound

Beyond Context: Large Language Models' Failure to Grasp Users' Intent cites this paper.

Beyond Context: Large Language Models' Failure to Grasp Users' Intent SafetyBench: Evaluating the Safety of Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:11:13.678082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T20:09:25.827452Z digest=sha256:da5b6dab31dc7132a943ebe0ba114b682284194037a90d1817ed1e49a2274b51

Observation f5fbc10e-70a4-445a-8b2a-2e2e99205d45 · inbound

Breakdowns in Conversational AI: Interactional Failures in Emotionally and Ethically Sensitive Contexts cites this paper.

Breakdowns in Conversational AI: Interactional Failures in Emotionally and Ethically Sensitive Contexts SafetyBench: Evaluating the Safety of Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.316619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T19:46:56.456625Z digest=sha256:8c9784b70af9073caf21d9583a05e55bb43a68aa73b7f17477a16282d80d94c4

Observation c903e576-4032-44b3-b529-1bafc14e20cf · inbound

VoxSafeBench: Not Just What Is Said, but Who, How, and Where cites this paper.

VoxSafeBench: Not Just What Is Said, but Who, How, and Where SafetyBench: Evaluating the Safety of Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:22.098823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T10:19:28.041282Z digest=sha256:4eb10ea1e0ebb9f71e2caafa534067873e12ed53498a981cd099f022f6bd28b4

Observation f1c6253d-ff1e-4731-a693-4590896ba80f · inbound

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs cites this paper.

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs SafetyBench: Evaluating the Safety of Large Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:24.878897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T08:01:25.938248Z digest=sha256:7eee6e369b9a59ee9de91216986719c4302c8a748243b18c58a9eb96a13da9e5

Observation d67dfdc2-8794-4a7b-a329-e7920a0585f6 · inbound

How Sensitive Are Safety Benchmarks to Judge Configuration Choices? cites this paper.

How Sensitive Are Safety Benchmarks to Judge Configuration Choices? SafetyBench: Evaluating the Safety of Large Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:56:35.407538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T03:41:32.338529Z digest=sha256:4d249f0728a3443febff72a9011eb7cf4b7fc8d80f24437b67c3c658cbbdcc7c

Observation 210eaf0c-c35b-4cd5-bb8a-6d26961a24bd · inbound

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels cites this paper.

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels SafetyBench: Evaluating the Safety of Large Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T21:39:24.707089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T12:07:02.778631Z digest=sha256:5b74e4c6091ff780cee03f53ad5407c27c196095d5e9dc74e7c828a66c63069c

Observation e61ae913-d454-4b05-90b8-92a50e6ce9fd · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks SafetyBench: Evaluating the Safety of Large Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.839851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:52ae428dc6df4f97b34863086d0dad6bee550f4a669f96c89b5e04a41742d819

Observation d7668e7c-4541-4545-b656-62b918729385 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety SafetyBench: Evaluating the Safety of Large Language Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.043818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:65dda13014700ffeeda52823d1daa3919a5ce51878c1973f5201e3dd03c29702

Observation 4d77745e-3327-4ef7-8c07-7c5138ffe310 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety SafetyBench: Evaluating the Safety of Large Language Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:42.916232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:9c642717f01df070642ad75cc866a2ee44e1c73b17048b5a5a07fefba573ad74

Observation 2c26ebce-5e82-406e-aaa1-46c1469485b6 · inbound

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation cites this paper.

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation SafetyBench: Evaluating the Safety of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T23:33:25.778345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:33:25.778345Z digest=sha256:65c0776059874aee582655e1a6927b3345626a0264d705a214f4eff86797b470

Observation 7859f42a-5e2b-46f0-acb9-e58f687c05b2 · inbound

Efficient Safety Benchmarking via Item Response Theory cites this paper.

Efficient Safety Benchmarking via Item Response Theory SafetyBench: Evaluating the Safety of Large Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:55:48.988875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T15:51:00.484343Z digest=sha256:8ae05f5f1228d57f91fdf538b25306f13f1b0781421ef431629f088e493cae9d

Observation 874e7ef1-95b4-4fd1-a4cd-b738bf512552 · inbound

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety cites this paper.

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety SafetyBench: Evaluating the Safety of Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:52:55.751302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T00:46:03.210076Z digest=sha256:e6c6b52209645e1f3eda090bc0e01c551042576959781088aec63ac905550e19

Observation 4114ba8d-56ca-476e-bbb9-665ba29f2bbd · inbound

Two AI Metrics Diverged: Will it Make All the Difference? cites this paper.

Two AI Metrics Diverged: Will it Make All the Difference? SafetyBench: Evaluating the Safety of Large Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:36:56.167767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-02T12:29:24.439779Z digest=sha256:7dbbf59d3880ef1284c1b8923d7d0bc66e325bc63bdee7bb124b9b5a8802a9df

Observation 501a2773-fc1d-4157-ae95-cde76bae5975 · inbound

Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI cites this paper.

Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI SafetyBench: Evaluating the Safety of Large Language Models

Reference 176

Resolution
unresolved
no resolver link, observed 2026-08-02T02:22:11.309343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:22:11.309343Z digest=sha256:a90bdc3cf67dc94435c3be20b38ded9349b8afeafb8009197c4d7a72dfab2afb

Observation 35546b27-3658-41db-ad14-ecb340c7aaee · inbound

What AI Red-Team Evaluations Can and Cannot Prove cites this paper.

What AI Red-Team Evaluations Can and Cannot Prove SafetyBench: Evaluating the Safety of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.369854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.369854Z digest=sha256:b7564e40575970b4db289a6fb692756f053ba369d01b134a3d4ee9db072cdb20

Observation e6acd021-14d3-4938-a9b6-830ddaa8c29f · inbound

AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models cites this paper.

AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models SafetyBench: Evaluating the Safety of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:30:02.855362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:30:02.855362Z digest=sha256:4a4ecd91efc67b6c7d924b2d5fe4b56c9781dd491f21190b9e46c16ab726732f

Observation 2b2a3d9a-93ff-40a9-b21e-f596346e3570 · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges SafetyBench: Evaluating the Safety of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:22.021201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:22.021201Z digest=sha256:e4f168817ee254b9027324dc78028a4a4e64ed010acc1c61b6c9a003140bc688

Observation ace38749-87d3-437c-9f1f-b634cb471d62 · inbound

CockpitHAT: Dependency-Graph-Driven Hierarchical Attribution for Embodied Multi-Agent Cockpits cites this paper.

CockpitHAT: Dependency-Graph-Driven Hierarchical Attribution for Embodied Multi-Agent Cockpits SafetyBench: Evaluating the Safety of Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T15:09:53.296703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:09:53.296703Z digest=sha256:bb229afa9f3bfdb077b7cd345e4ab12dd2189a530b4dd12b9621c36332fa3a81