Pith. sign in

Paper Citation Record · LEDGER

ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2310.17389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.17389 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:28.909375Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.654187Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4ae480b6-9c61-470c-a90d-6399246e790d · inbound

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs cites this paper.

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:25:14.866587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T16:25:14.744887Z digest=sha256:ba0b831d384c479a8df338336b3822f19af3a12c923f745b515f674bceceafd8

Observation a6049c2b-aeaa-44a8-8685-d434ac56dbb4 · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.523130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:1deb7d255d218ee337e8b6ee1f81f975efc2b97abc00ffdf0ebf52abeae853ab

Observation e773cb32-b65a-4db0-8695-1cd4132b6e08 · inbound

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use cites this paper.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.909375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.909375Z digest=sha256:32c980ff08e9f3e24465b8a69c3b5cd4a41d65de35d1e256df0fc2afa66703a4

Observation 96d8ff20-356b-4be2-9aba-bb9b38955dea · inbound

Large Language Models Often Know When They Are Being Evaluated cites this paper.

Large Language Models Often Know When They Are Being Evaluated ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.795060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.795060Z digest=sha256:4bdebc64aaf034cc6774ae5fbb7946573b84e6bfbd1d898cb773f17c16f5f2d5

Observation c0f879a3-92f5-496f-abec-9b8f05638b4d · inbound

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law cites this paper.

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:54.802416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:54.802416Z digest=sha256:1f16adca23e36e8bd583fce799c269d05c1bb86cef8cbfe22c184e5dbc54e9b1

Observation 899debdd-8dbc-4652-b50f-8fe1d3245bd9 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.647489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.647489Z digest=sha256:3b00593451c8669d2cbe2413b31f96c2f8515d2b9813791e43795e6cd3897c8f

Observation a59b22d2-ab12-433d-8fca-fcf93f1968f3 · inbound

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments cites this paper.

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:02.699513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:02.699513Z digest=sha256:dc7a802f50d842cd34b66fcb65946ef374895edfe227b625c1aea7f8e7514dd6

Observation 8945f4db-2a41-4c61-9a11-ef0ac1d66d3e · inbound

LLMs Encode Harmfulness and Refusal Separately cites this paper.

LLMs Encode Harmfulness and Refusal Separately ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:03:15.999906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:03:15.999906Z digest=sha256:247796c6e56ab771d68acb786e5cf86718a523b3eeb57276cb607f3fcc5ba81d

Observation 7d08457f-5a13-4da3-af0c-6a428b8406fa · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:44:22.044971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:914608930615cd232537949239de4af90ebb686112843336a89bdb826daa7156

Observation fe7c1cb3-98e9-4202-af1b-118f250474bf · inbound

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio cites this paper.

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T05:07:09.235559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:07:09.235559Z digest=sha256:ac90c3d8529d83de6d4b2391fc5c8dea4fccedd85acd4525aa0ddc2925befde2

Observation 86848d71-d41c-4229-aecd-7ea9e6c58cbc · inbound

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders cites this paper.

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:12.009201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:12.009201Z digest=sha256:e191c895c155f6d21af736182044b30f7d4a2b037dcb10ee197f1b5bcdf9678d

Observation f7e3cf6a-63e5-4fe4-b5bc-3396ff7d194c · inbound

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization cites this paper.

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:01.782343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:22:40.613937Z digest=sha256:d6c81a2e5c2e03660839403c6dd069073668efbb8a5ed78ab52738ed75a8631a

Observation 14c65156-fe1a-4726-8f6c-d6fff637a2ff · inbound

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs cites this paper.

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:37.658425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:27:13.339411Z digest=sha256:225bb77b26fcd2f5735dce3dd4bfe92fcee726c63d367e450ee465561fc364fe

Observation f3e2d616-a3f3-4103-a614-aa060b0a2e05 · inbound

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts cites this paper.

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:08:25.512197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T09:07:57.713675Z digest=sha256:a45c9be38b8981413aa69fcd74e50f26e7d4ed56bd7c9d048f557e9a95129371

Observation d5b2a56b-da7c-4d7e-99ac-83fb41c9db39 · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:ca78d0ca39b397c7ec35b219093d48b754ab36b70f1a918bca0d081878a2bfd9

Observation e8af5d2c-560f-4e6f-aa37-e9d5cb70633c · inbound

Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs cites this paper.

Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:46:05.501341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:32:00.789721Z digest=sha256:5859ec3a41537b260019c1c5f30e2d6d2a8baa7c5403ea6cd0fa610776557886

Observation 2d758f06-8adc-4c14-8c4d-2f363f1f20ad · inbound

Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems cites this paper.

Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T19:06:09.741391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:41:53.950049Z digest=sha256:5f3fbf682a43f9326207e45f7b43479f6b08c981ca2077a77d695fd90faaf219

Observation 01f579b9-361c-4712-a23f-9db28b5a9661 · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.946872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:11d2138ab9b81551d2da4f16539bf05549f24ea0972f14db495ecfd943c25c19

Observation 700c0b50-026d-4736-b3d9-31e4fa83de6f · inbound

When Youth Enter the Algorithmic Wild: Discovering and Understanding Potentially Harmful Teen Videos on Douyin and Kwai cites this paper.

When Youth Enter the Algorithmic Wild: Discovering and Understanding Potentially Harmful Teen Videos on Douyin and Kwai ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:25:19.788374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T04:20:27.456558Z digest=sha256:091a14e9603148f7e5294c44b539921f4991111ec10a49f73ac5a737c1bfa1e2

Observation c61f411c-1406-4b77-8914-6f10225690b0 · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.897253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:0236c8e24c0084667a132ea36f8217e7c5735c516adb95193f7249abd24d4194

Observation 59558a22-481c-44f1-a762-506f6272a304 · inbound

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety cites this paper.

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.655714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T14:38:55.045628Z digest=sha256:57dffce10078eaa436e5c804ea4767e83eb906cde35782bb5f3db3dd28f6f3f6

Observation 048e4edc-5e2d-46cd-aa71-38c0bc50b6ab · inbound

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B cites this paper.

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T14:52:12.591220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:52:12.591220Z digest=sha256:b081539fea95eb6b8235b0b76c23ac6bb30e781b021d2b94a5e7f0feb3615c36

Observation 2fd0eeb0-ba99-451f-a1ed-ebdbb1a87b24 · inbound

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) cites this paper.

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:07:16.653669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:07:16.653669Z digest=sha256:218af94d5388e8b09715098ed27309b82e2c2d85e614e3e11390cc38ea168866