Pith. sign in

Paper Citation Record · LEDGER

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.00973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00973 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:10.614605Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 586233db-19d4-4c9a-b850-cf4cb60a26e0 · outbound

This paper cites You are an AI designed for safety alignment research. Your task is to analyze the following scenario where AI-generated contentcould be used for manipulation.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content You are an AI designed for safety alignment research. Your task is to analyze the following scenario where AI-generated contentcould be used for manipulation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:12.263959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:09.836233Z digest=sha256:01fe095511c2f066c96a07ba9769c92c0154f3c209fe157287929741ef3587af

Observation d2d2bb59-df81-444c-bad7-fec40998d372 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.851136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:10.080302Z digest=sha256:41d365e44cb67a2f4baae906ea726b1d70b9fe06e4f98746f9132a0eee559acc

Observation b6ceeb29-2a8d-4b40-ab91-094eb4d34cc9 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.730720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:10.191363Z digest=sha256:991c39ce72937dd3c82e7af23035f8f8c0d64573427ff495018023f7273392cc

Observation 00d4efbf-251e-45f0-aacf-3022cac0007f · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:12.087339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:09.899323Z digest=sha256:f162bf303c5a9546bd4445fbe7d902417ea59902e07437baf58d31cbafb0b0e3

Observation a54986ec-e5a1-4b04-ad3e-a4fed3543a2f · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.962138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:09.944793Z digest=sha256:a4eed04803195f600bcb22388d032bc4ee7dc739221a4613660273e8fb2359e2

Observation f4b01a87-a2bb-4d6f-b45d-2bb06b036002 · outbound

This paper cites Detailed Levels level 0 - Absolute Refusal Definition:A complete rejection or extremist ide- ologies, narratives or actions.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Detailed Levels level 0 - Absolute Refusal Definition:A complete rejection or extremist ide- ologies, narratives or actions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:11.642610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:10.248279Z digest=sha256:d9c5e1d16e066482180142212586f1c91cdbaeb211700253cbbebec2e7e1add6

Observation 64b82840-e03a-4666-84f4-c92246fc0a88 · outbound

This paper cites 0" representing lawful and safe content and.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content 0" representing lawful and safe content and

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:11.513095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:10.325389Z digest=sha256:48a19070e56b7866560eed975c6c60ea5125e221475d87644a269dfab24833e7

Observation 6a326a86-b071-4f94-9fb7-fcaf546db0e1 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.408282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:10.429649Z digest=sha256:868202297118e18b0c93d9615d0717cdf7eb15fd0e7b49de82e8a26ff8bf5c63

Observation ccf4f5f3-5082-4843-8dc6-ff6faaf08e3e · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.230236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:10.468405Z digest=sha256:e497721becea00a30255fd6931d6dba6322332e25078b2b27b8cd0d28c910597

Observation aab61333-92b3-46b2-ba54-1a936878e858 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.018208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:10.538602Z digest=sha256:fb24338de692bd119598e5e67af07f2d02885a0c5abbe6b7ee9ec3002e6cb268

Observation 67c0a4c1-5d26-4407-8e5a-fb4b766f67ed · outbound

This paper cites We will strike the enemy base at dawn. The group will split into two teams to breach the perimeter.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content We will strike the enemy base at dawn. The group will split into two teams to breach the perimeter

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:10.813619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:10.614605Z digest=sha256:a34ab992f2b15a4fa2d9bd56d8c8068f97f502c7149734e06d4e43df5a4e54f5

Observation 3d919e24-df8a-4ec0-a53b-db5a755ee42d · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:09.735735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:09.735735Z digest=sha256:c89f18c9a9059d8be2332d64d9bba2052c26cb3a15bfe4e622908eafcf4fe31a

Observation 3278ca0b-d46c-4543-aa8d-cb9cb61735a2 · outbound

This paper cites BingoGuard: LLM Content Moderation Tools with Risk Levels.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content BingoGuard: LLM Content Moderation Tools with Risk Levels

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:09.666314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:09.666314Z digest=sha256:62680c93440a68d4d796e971efa4646fd086ea6302514722290867b95e8b64b6

Pith citing papers

No inbound Pith citation observations are available.