Pith. sign in

Paper Citation Record · LEDGER

SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2311.08370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.08370 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:05:17.610727Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:59:32.652255Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4f5f243a-d99e-431f-9ae2-b485bdeb8495 · inbound

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs cites this paper.

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T16:25:14.904441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T16:25:14.744887Z digest=sha256:5113bf6aad821666df7760bc99dd9aa7eeffe9ee28677ca5c3a3b98a8eb20f91

Observation 26fd4747-a2ad-4cfa-b1e9-139de5cd2713 · inbound

Granite Guardian cites this paper.

Granite Guardian SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.511553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.511553Z digest=sha256:c6cb643b4813a0f1937f454cd0f24cea345d5936c91532534b11222cdb0ecb5a

Observation f2445329-97dd-4927-8001-60aef3eec8ed · inbound

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models cites this paper.

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-10T14:51:58.835765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:51:58.835765Z digest=sha256:1da607cf4e114dccd3b1c6e59359b51ea76790482ec9c4b7c6de48d3466e3bb3

Observation 47aba5ba-9f30-4c3b-9b95-6928a2e9bb0b · inbound

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation cites this paper.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.975584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.975584Z digest=sha256:7e7ea724dac74861122a0e83c69fdbf92ccbe4d855ae8e4e4699a87d1cca47cb

Observation 861366ee-8566-4fb9-8b06-63bba9d9dac5 · inbound

o3-mini vs DeepSeek-R1: Which One is Safer? cites this paper.

o3-mini vs DeepSeek-R1: Which One is Safer? SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T23:32:16.829764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:32:16.829764Z digest=sha256:24d3264262f2ce58072a18a2ee0ead017ddaf9d41178cbf692c6281892052c35

Observation f57258b7-1e11-40f5-881c-523b8d93e231 · inbound

GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning cites this paper.

GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:17.610727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:05:17.610727Z digest=sha256:d1cd4b511c8d4d778f99f123790afe6a5d7a8fe7074db8863e3c860b130d9004

Observation f83ef09f-72de-4e00-92e2-0405adcdb8d1 · inbound

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency cites this paper.

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:33:40.319869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:33:40.319869Z digest=sha256:f224cd68402c987a80252a5ce024cdff1fa8484dae3329062f220d2eb34663fa

Observation 09a1df70-ccc8-4ba9-9cf1-87fe13daffc1 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 290

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:08.309134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:08.309134Z digest=sha256:095ee42e3f5783f44660d81d77dda16588a24521380877e715fa29b871d93745

Observation dd6466c5-a3a4-419c-afb5-52a2f4965a5b · inbound

Scaling behavior of large language models in emotional safety classification across sizes and tasks cites this paper.

Scaling behavior of large language models in emotional safety classification across sizes and tasks SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:25:23.366764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:25:23.366764Z digest=sha256:2f3f4e6486e97218233bf93596c14698cea4de03750e54c9428b90dbd84a4a60

Observation 7923d2dd-4db1-40a6-ae09-5a6959bf521e · inbound

YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models cites this paper.

YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T08:52:34.725666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:52:34.725666Z digest=sha256:2f762e285876049f30821702cb718d7152334f5642143d00a41272ea59e26bf4

Observation 4a269938-0b2e-41e3-89cb-813df70be459 · inbound

Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming cites this paper.

Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T11:43:31.089245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T11:43:31.089245Z digest=sha256:1f2f2e68690c4a3de49795be0562a80277ec89b455c54a077920e1bfc8533feb

Observation 541b6e26-2efd-412e-afee-8da6d76b1fcd · inbound

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules cites this paper.

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:28:09.962511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T19:24:54.381722Z digest=sha256:98738a1fc88604aa387e9295b29874e7716abfc944c717acfff64427bb7c66e6

Observation cc973d17-47c9-489b-82d9-1ab742d3aec2 · inbound

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs cites this paper.

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:46:35.378670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:27:13.339411Z digest=sha256:20c6d88c0ff0249ef713a96c02b46b678b31ebae7090762401570b3d4aa660d7

Observation 3e4d273c-0bd7-403f-bafe-ec9e1246d42a · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:20:23.111469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:c40d61f68002440c36433c1ef9116d397cc71706753f83739e69b7e25f8d8966

Observation 9ea79794-54cc-4934-974d-0f9707be9d06 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.476793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:d6791c6b5a81f425e409611a5aef35a1220361c662befb6cbded64665bf26c33

Observation bb71f8f2-96a5-477f-90a6-42c06ca75125 · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.880901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:13:16.896486Z digest=sha256:deebe7453a34781e6f22d775e7594c67f353f128e71b18d0b226e2bc66aa407d

Observation b388b35e-96a5-4604-a763-b991bc3f650c · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T14:49:14.212795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:49:14.212795Z digest=sha256:cd242d2393e1d2dda27434627fb78d609a98b4214e52724a820bfcdec0951002

Observation 93da4356-bb7f-465f-8d4c-42dab9e7f2f1 · inbound

GLiGuard: Schema-Conditioned Classification for LLM Safeguard cites this paper.

GLiGuard: Schema-Conditioned Classification for LLM Safeguard SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:20:55.261341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T03:19:11.495230Z digest=sha256:2904d7b5a8f5cbeb663fcd9cddff43ff7b658b5a9a906dbdc7be102615895cbb

Observation 33dfac57-ee03-403b-82b1-6f1f32950cde · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:23.905179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:e3216498ebfb928eb88eb1f0da0d7d475693287ab809a33645918c69a91a406b

Observation d56d2ff7-7585-4262-bde0-e4d5dfb61862 · inbound

Assisted Counterspeech Writing at the Crossroads of Hate Speech and Misinformation cites this paper.

Assisted Counterspeech Writing at the Crossroads of Hate Speech and Misinformation SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:11:12.991681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-22T07:07:36.726431Z digest=sha256:e6ef68af5bb0d89b8c0cd1979ea5b42e73e42845d5f116dbc1d501af903da04e

Observation 98611d6b-ce7f-4d3c-a713-58155635b318 · inbound

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content cites this paper.

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T09:13:15.953047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T09:11:58.843585Z digest=sha256:92d70d284123692846ab2dcdd4440c8e9314d7d82ab59e9053cc1f7649951301

Observation 038c79bb-5e1e-48f9-99d3-ef3b851d43f2 · inbound

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection cites this paper.

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:09:18.912001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T20:40:15.506976Z digest=sha256:c3dc727194286e522dd002b2cec5da6bc6c582bc718c67e2776a8b29fd79fd06

Observation 3136f6e0-4c54-4f1d-9371-5d9978f37405 · inbound

CATCH-ME if you RAG: a dataset of Contextually Annotated multi-Turn Counterspeech against Hate and Misinformation Exchanges cites this paper.

CATCH-ME if you RAG: a dataset of Contextually Annotated multi-Turn Counterspeech against Hate and Misinformation Exchanges SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:59:32.654360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T17:30:07.053955Z digest=sha256:60b2d943d1188f921f9089dadeedb8b83abc26e4693f5c7e222667f3ed06d865

Observation e7cfaa08-b924-4810-9d1c-e6f39525d57e · inbound

Efficient Safety Benchmarking via Item Response Theory cites this paper.

Efficient Safety Benchmarking via Item Response Theory SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:55:48.964480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T15:51:00.484343Z digest=sha256:c7544c097b24bf72c761360f45e12081860f2d363a59115b053e6e1105dae119

Observation b6c5d441-8210-43b1-b07c-0172b1668320 · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:35:51.347673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-29T03:35:34.594617Z digest=sha256:85b6924a1b17be455f5714e45d190cd11c705e3af140f81063b22a3798113d54

Observation 88e4652b-bc78-4315-926b-f45ec0be14ba · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:34:34.459004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-30T09:32:59.824110Z digest=sha256:0c21e4ef957c6986b9dccb7c80123e158a7e9e12a22718745190d72681488fb7

Observation 431a847b-be93-46c0-b3b2-f91be2408787 · inbound

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation cites this paper.

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:36:57.443689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:36:57.443689Z digest=sha256:dd1533c8a8d59921d0233ed08eddd837b81c9c9e7c8eeaecb85a1e808cdc6dba

Observation 9f208a3a-3e98-4988-bd5a-a9a9a43da872 · inbound

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B cites this paper.

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T14:52:12.602287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:52:12.602287Z digest=sha256:210b5bb2fdf63a38606f50eb69a10274222a544cda6b1c9ede4d3db9a576aa25