Pith. sign in

Paper Citation Record · LEDGER

SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2311.08370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.08370 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:33:40.319869Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:59:32.652255Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4f5f243a-d99e-431f-9ae2-b485bdeb8495 · inbound

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs cites this paper.

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T16:25:14.904441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T16:25:14.744887Z digest=sha256:9aad7623d5230fb986565e00590554eb357f816cb552586261f2d138635b9472

Observation f83ef09f-72de-4e00-92e2-0405adcdb8d1 · inbound

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency cites this paper.

Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:33:40.319869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:33:40.319869Z digest=sha256:7552bbb1d0043644c3eab9c5e22702f373389ebaea1c162897a98012d8a1197f

Observation 09a1df70-ccc8-4ba9-9cf1-87fe13daffc1 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 290

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:08.309134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:08.309134Z digest=sha256:b7f515dc6d233aaf79c7836488d48f86c0136bfd69d3e878e4038693ea2e1a7f

Observation dd6466c5-a3a4-419c-afb5-52a2f4965a5b · inbound

Scaling behavior of large language models in emotional safety classification across sizes and tasks cites this paper.

Scaling behavior of large language models in emotional safety classification across sizes and tasks SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:25:23.366764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:25:23.366764Z digest=sha256:c3bd38c3f4109685646c9ea18425c6d83afbb366942535d45b170a4ea8a220f4

Observation 7923d2dd-4db1-40a6-ae09-5a6959bf521e · inbound

YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models cites this paper.

YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T08:52:34.725666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:52:34.725666Z digest=sha256:415732d0aa1ff04e8731158199d661557f558157a953afb5eb8ad6f23127b4f3

Observation 4a269938-0b2e-41e3-89cb-813df70be459 · inbound

Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming cites this paper.

Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T11:43:31.089245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T11:43:31.089245Z digest=sha256:500415f99999b7268c99a4390a9852f3fbd1764fae2cc5edd35ef69fab1768e6

Observation 541b6e26-2efd-412e-afee-8da6d76b1fcd · inbound

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules cites this paper.

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:28:09.962511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T19:24:54.381722Z digest=sha256:e9982197c398ab2920526f6e53d4fe9e9b88f9ee12fecdb1df0dcd2439d0c584

Observation cc973d17-47c9-489b-82d9-1ab742d3aec2 · inbound

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs cites this paper.

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:46:35.378670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:27:13.339411Z digest=sha256:e02d6b6436679f0e67fb348bf234e0c09f6dfd19b19e20402f0a50a782ee48dc

Observation 3e4d273c-0bd7-403f-bafe-ec9e1246d42a · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:20:23.111469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:17fff8a25f50dd3a030a41b5209fbd9f967d90447c6cc0f3c23d375984f32658

Observation 9ea79794-54cc-4934-974d-0f9707be9d06 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.476793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:b3d41d0dc98c77a91f71382148f750c62549d9f328d6b3c30cbdb5ca8e0a9c80

Observation bb71f8f2-96a5-477f-90a6-42c06ca75125 · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.880901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T12:13:16.896486Z digest=sha256:73a80051f147a1fd33a4d7a4026ec25c2458d4773a012a00849fa2887939ce9a

Observation b388b35e-96a5-4604-a763-b991bc3f650c · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T14:49:14.212795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:49:14.212795Z digest=sha256:4e3a2e0d06228a6b5828e29588209fe963eefee39af02226a1a5ab862b02fe47

Observation 93da4356-bb7f-465f-8d4c-42dab9e7f2f1 · inbound

GLiGuard: Schema-Conditioned Classification for LLM Safeguard cites this paper.

GLiGuard: Schema-Conditioned Classification for LLM Safeguard SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:20:55.261341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T03:19:11.495230Z digest=sha256:f8699eb85c66720993ba67d7b90ad8a967543969bc0f22ebdf09db465cbe055c

Observation 33dfac57-ee03-403b-82b1-6f1f32950cde · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:23.905179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:49a913ed466668465671ae90e035afcb1c3940030854fd1e6a46beed4e878bc3

Observation d56d2ff7-7585-4262-bde0-e4d5dfb61862 · inbound

Assisted Counterspeech Writing at the Crossroads of Hate Speech and Misinformation cites this paper.

Assisted Counterspeech Writing at the Crossroads of Hate Speech and Misinformation SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:11:12.991681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T07:07:36.726431Z digest=sha256:5fd37d0fe302a2245e0c70f8120443a374f924e634899bce535314e5da473bbe

Observation 98611d6b-ce7f-4d3c-a713-58155635b318 · inbound

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content cites this paper.

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T09:13:15.953047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T09:11:58.843585Z digest=sha256:0e8e954ed143431d49c4ca7391e32748a34501ca35193e99caf9de8660ad06d6

Observation 038c79bb-5e1e-48f9-99d3-ef3b851d43f2 · inbound

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection cites this paper.

Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:09:18.912001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T20:40:15.506976Z digest=sha256:5898c1f1da79b9d55e90582ac49b6851dd52364d71fa924a20d7b237973545d9

Observation 3136f6e0-4c54-4f1d-9371-5d9978f37405 · inbound

CATCH-ME if you RAG: a dataset of Contextually Annotated multi-Turn Counterspeech against Hate and Misinformation Exchanges cites this paper.

CATCH-ME if you RAG: a dataset of Contextually Annotated multi-Turn Counterspeech against Hate and Misinformation Exchanges SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:59:32.654360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T17:30:07.053955Z digest=sha256:77abdd5abc181c8b01cf0260cc763e7f4204e67608d34142b68abf303e01436a

Observation e7cfaa08-b924-4810-9d1c-e6f39525d57e · inbound

Efficient Safety Benchmarking via Item Response Theory cites this paper.

Efficient Safety Benchmarking via Item Response Theory SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:55:48.964480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T15:51:00.484343Z digest=sha256:57b646126536716b5723a25ddd40d93a6afcf170ef0026aa53262ed0cd7a5e6d

Observation b6c5d441-8210-43b1-b07c-0172b1668320 · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:35:51.347673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T03:35:34.594617Z digest=sha256:349f434e1d2675c6fa5147ea4ca2618b333c6b14881975c99899df9f7179724c

Observation 88e4652b-bc78-4315-926b-f45ec0be14ba · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:34:34.459004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T09:32:59.824110Z digest=sha256:662d3733c38c0eca7ae2bfdda5e38416c1299448b24265dbb1f715082280cf94

Observation 431a847b-be93-46c0-b3b2-f91be2408787 · inbound

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation cites this paper.

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:36:57.443689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:36:57.443689Z digest=sha256:6c7ba6289e608831bf458839e093b163ba75d8b90b7a5de551a6a25572aea29f

Observation 9f208a3a-3e98-4988-bd5a-a9a9a43da872 · inbound

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B cites this paper.

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T14:52:12.602287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:52:12.602287Z digest=sha256:791534a2270edd901f26d6bd1abe89943999f21fd41a9638ec98a7a12ecb17ad