Pith. sign in

Paper Citation Record · LEDGER

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.00973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00973 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:10.614605Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 586233db-19d4-4c9a-b850-cf4cb60a26e0 · outbound

This paper cites You are an AI designed for safety alignment research. Your task is to analyze the following scenario where AI-generated contentcould be used for manipulation.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content You are an AI designed for safety alignment research. Your task is to analyze the following scenario where AI-generated contentcould be used for manipulation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:12.263959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:57:09.836233Z digest=sha256:5be190fb16a4e33797e8730ddd1ecda0b90b11c9ff2e797d5735d64195783321

Observation d2d2bb59-df81-444c-bad7-fec40998d372 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.851136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:57:10.080302Z digest=sha256:7b9e1f505866f55729ef378c3990a39cb958403beed96507804789638a2c8a47

Observation b6ceeb29-2a8d-4b40-ab91-094eb4d34cc9 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.730720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:57:10.191363Z digest=sha256:dd2243e4d8d4029c8381d0bece2434add26f257f7845729dcf3f557e09a820a9

Observation 00d4efbf-251e-45f0-aacf-3022cac0007f · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:12.087339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:57:09.899323Z digest=sha256:9afbea605c8ed280ec7e7ee23d37e1b846c5ec549f7068c4e7a1227dc47697e3

Observation a54986ec-e5a1-4b04-ad3e-a4fed3543a2f · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.962138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:57:09.944793Z digest=sha256:f56384d16c79b7af58dd79dace8d4c9580349953e3e69957a663346773b952f6

Observation f4b01a87-a2bb-4d6f-b45d-2bb06b036002 · outbound

This paper cites Detailed Levels level 0 - Absolute Refusal Definition:A complete rejection or extremist ide- ologies, narratives or actions.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Detailed Levels level 0 - Absolute Refusal Definition:A complete rejection or extremist ide- ologies, narratives or actions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:11.642610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:57:10.248279Z digest=sha256:8a4c21efcf4e4ded3b97250cf025971076d2f7d485c8b568e930a7c2781a1114

Observation 64b82840-e03a-4666-84f4-c92246fc0a88 · outbound

This paper cites 0" representing lawful and safe content and.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content 0" representing lawful and safe content and

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:11.513095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:57:10.325389Z digest=sha256:afe10444942b5c91a4dead6a47c1939bdd342785ead236d07b9d1716cccae72a

Observation 6a326a86-b071-4f94-9fb7-fcaf546db0e1 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.408282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:57:10.429649Z digest=sha256:7359b479e1f325d0734586c53aa7e67941d36a50ad432f50106fa5c4342eb0c7

Observation ccf4f5f3-5082-4843-8dc6-ff6faaf08e3e · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.230236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:57:10.468405Z digest=sha256:5e3c70f9b16330dbb9d967be1b1b7ce2171d35b0990bb60f71dc7c2fd2f06e13

Observation aab61333-92b3-46b2-ba54-1a936878e858 · outbound

This paper cites an unresolved cited work.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:11.018208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:57:10.538602Z digest=sha256:b77577a2c7e29b2da8a1658ca5efcab7e5be097c5dec9059a44615a3add5cc73

Observation 67c0a4c1-5d26-4407-8e5a-fb4b766f67ed · outbound

This paper cites We will strike the enemy base at dawn. The group will split into two teams to breach the perimeter.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content We will strike the enemy base at dawn. The group will split into two teams to breach the perimeter

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:10.813619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:57:10.614605Z digest=sha256:9d46dc608528f5e9d9017cf9179444062a040da4afb199dcdae34616df4fb918

Observation 3d919e24-df8a-4ec0-a53b-db5a755ee42d · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:09.735735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:09.735735Z digest=sha256:c89f18c9a9059d8be2332d64d9bba2052c26cb3a15bfe4e622908eafcf4fe31a

Observation 3278ca0b-d46c-4543-aa8d-cb9cb61735a2 · outbound

This paper cites BingoGuard: LLM Content Moderation Tools with Risk Levels.

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content BingoGuard: LLM Content Moderation Tools with Risk Levels

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:09.666314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:09.666314Z digest=sha256:62680c93440a68d4d796e971efa4646fd086ea6302514722290867b95e8b64b6

Pith citing papers

No inbound Pith citation observations are available.