Pith. sign in

Paper Citation Record · LEDGER

ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2310.17389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.17389 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:00:23.903994Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.654187Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4ae480b6-9c61-470c-a90d-6399246e790d · inbound

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs cites this paper.

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:25:14.866587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T16:25:14.744887Z digest=sha256:642f96186534e2fe77728a23d58bc99c77665ba8db3c15b7616ff25a82aa3c8e

Observation a6049c2b-aeaa-44a8-8685-d434ac56dbb4 · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.523130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:9c8e0e60f7bc1c602a8d932dd496740e08b19cc4b2321c0eeb83d4cfab2169b2

Observation 116905c3-65b8-486b-8be9-7045ea8274e6 · inbound

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models cites this paper.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.475947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.475947Z digest=sha256:8e644fe2e4429a91af58d319f7ee67ff40999edca9d2fd6aebc6f14b11925b8a

Observation 1272fbd4-66a2-4e77-b78a-bbad6b853592 · inbound

Improved Large Language Model Jailbreak Detection via Pretrained Embeddings cites this paper.

Improved Large Language Model Jailbreak Detection via Pretrained Embeddings ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:21:32.488307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:21:32.488307Z digest=sha256:fb314efa102a2a7ffbb7b5345c983c9f26878156de02d26fe544539bed0259c0

Observation c8b0fca0-7505-4e7a-b79e-7f55e6f44299 · inbound

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models cites this paper.

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:52.197451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:52.197451Z digest=sha256:e8288a0b45db939218cc3633f6a4e4f9175e3682b4994d0433dc4b3dea18b414

Observation 198ee191-dbd8-4346-be03-fbbedc16531f · inbound

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs cites this paper.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.152264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.152264Z digest=sha256:15c683236c321565d5f371b63854f4d557c15f24dce65297e274c8be20b00bd9

Observation bfd16ff5-243f-48be-8b4e-a9f4127b063a · inbound

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch cites this paper.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.236804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.236804Z digest=sha256:4deaebcce721901e991cc237a5d9e89a9981fe8b74d3677d6c85b64b315bda38

Observation 39a67614-0899-461c-905f-e26885284d9b · inbound

Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails cites this paper.

Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:54.750339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:15:54.750339Z digest=sha256:1202239cc959a34ae8203ec3b2386871039ab40d4d5ca6e24df06ae7af8ac7eb

Observation 57b6ae12-f0d5-406f-b47f-1ed7be08b990 · inbound

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences cites this paper.

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T10:23:06.168953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:23:06.168953Z digest=sha256:fc72eb09a6a05f1445a699340bd75ae2ee6d0367ebe8902e73c60e6cc0f6459f

Observation dc7e57a3-1e0a-4af9-bf72-6eda56d6263f · inbound

Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing cites this paper.

Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T06:00:23.903994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:00:23.903994Z digest=sha256:38f05ae31dca53f0851827d0d9af3014f299e5c7190040215ff6465fc7da7d89

Observation fd904c7b-fa90-4073-ad78-93f07df5ee81 · inbound

GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning cites this paper.

GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:17.470235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:05:17.470235Z digest=sha256:1cb589a10a8817e968ee7f1f2c74c8985eddc725a4a3ec7e49d6fd203f910def

Observation e773cb32-b65a-4db0-8695-1cd4132b6e08 · inbound

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use cites this paper.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.909375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.909375Z digest=sha256:3c2639c102631ec5d39b5b0c75dc77988855bcee3433a2373b91dde0668ba27a

Observation 96d8ff20-356b-4be2-9aba-bb9b38955dea · inbound

Large Language Models Often Know When They Are Being Evaluated cites this paper.

Large Language Models Often Know When They Are Being Evaluated ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.795060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.795060Z digest=sha256:39f0590b61c03b7d889251ea67dd21ef99ad3e1f0d4a7ce56592c208e014bc4f

Observation c0f879a3-92f5-496f-abec-9b8f05638b4d · inbound

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law cites this paper.

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:54.802416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:54.802416Z digest=sha256:2acb0bcebbe8ea8135a6c91fa03d6e0b6a8dfd3b985361dee698b3f305ac0452

Observation 899debdd-8dbc-4652-b50f-8fe1d3245bd9 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.647489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.647489Z digest=sha256:87a8d20e7f4004bf7864590551953cb234557a4af8526d52c23541ae54df8b0f

Observation a59b22d2-ab12-433d-8fca-fcf93f1968f3 · inbound

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments cites this paper.

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:02.699513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:02.699513Z digest=sha256:71bd0b75f3d61150cfbe55ba30c5b4708c703062ba4c3d28fe1e86d20f9b1bfb

Observation 8945f4db-2a41-4c61-9a11-ef0ac1d66d3e · inbound

LLMs Encode Harmfulness and Refusal Separately cites this paper.

LLMs Encode Harmfulness and Refusal Separately ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:03:15.999906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:03:15.999906Z digest=sha256:fad2f8199636d265c51ac4d5e91ff92673aff42592f560b85fe2f372953034b1

Observation 7d08457f-5a13-4da3-af0c-6a428b8406fa · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:44:22.044971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:f81949e912fefcd431cda6f226d582732560be08d66b5dc498a03b074755438c

Observation fe7c1cb3-98e9-4202-af1b-118f250474bf · inbound

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio cites this paper.

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T05:07:09.235559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:07:09.235559Z digest=sha256:b4d0f91b67454e8f6321bbf5851193163c0dcda6b45a59c3b36fa57c6ee5adaf

Observation 86848d71-d41c-4229-aecd-7ea9e6c58cbc · inbound

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders cites this paper.

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:12.009201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:12.009201Z digest=sha256:dd45c86abbb8b2e7838162f44aee3ad48f5a8b4f05418c7ee035654dde66ad90

Observation f7e3cf6a-63e5-4fe4-b5bc-3396ff7d194c · inbound

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization cites this paper.

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:01.782343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:22:40.613937Z digest=sha256:85ae67729c4bfd07dc996f8f697353d1bbad927e46895de8b2d7df08a6940bf8

Observation 14c65156-fe1a-4726-8f6c-d6fff637a2ff · inbound

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs cites this paper.

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:37.658425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:27:13.339411Z digest=sha256:fcd2432d0860deaf31eff515bcd66b5a3bf2234fe2c79a47e6499922cfaaad5c

Observation f3e2d616-a3f3-4103-a614-aa060b0a2e05 · inbound

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts cites this paper.

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:08:25.512197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T09:07:57.713675Z digest=sha256:9c2d38ef98e35892d8c5dbcafde15c5311492c1707bb11d7311068f7deea628f

Observation d5b2a56b-da7c-4d7e-99ac-83fb41c9db39 · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:53dde3b5e4b9bd45fad361af2b51cb050e89f492d365bea0e8caeb8c4fee98d4

Observation e8af5d2c-560f-4e6f-aa37-e9d5cb70633c · inbound

Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs cites this paper.

Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:46:05.501341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T00:32:00.789721Z digest=sha256:b0cbc62d4319c888beb95b4e424ac3f6553f49e619eca2f104311591defb4c7f

Observation 2d758f06-8adc-4c14-8c4d-2f363f1f20ad · inbound

Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems cites this paper.

Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T19:06:09.741391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T12:41:53.950049Z digest=sha256:0c7b815a7f78ed11df069d1d5ef9c55896042e9dfafcb3a842f91ffa2efeca7e

Observation 01f579b9-361c-4712-a23f-9db28b5a9661 · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.946872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:ff2ccadda36017cd2fa51b617b318b1fa0d87e9d484adbae08d678c91cc7609c

Observation 700c0b50-026d-4736-b3d9-31e4fa83de6f · inbound

When Youth Enter the Algorithmic Wild: Discovering and Understanding Potentially Harmful Teen Videos on Douyin and Kwai cites this paper.

When Youth Enter the Algorithmic Wild: Discovering and Understanding Potentially Harmful Teen Videos on Douyin and Kwai ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:25:19.788374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:20:27.456558Z digest=sha256:e1b01a1964a2f8f3fba184fde5504a6f045703bc2bfbbd1a7dca8f5f41d71e20

Observation c61f411c-1406-4b77-8914-6f10225690b0 · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.897253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:d56234d1f670b5cefc5d536abcc5edeb0094a2850436a1dca585a34fc9dde2b6

Observation 59558a22-481c-44f1-a762-506f6272a304 · inbound

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety cites this paper.

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.655714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T14:38:55.045628Z digest=sha256:e6de8896b3c36caf7df392f0861c5dcb6625e7b2c3d2c85a549c307375c084ff

Observation 048e4edc-5e2d-46cd-aa71-38c0bc50b6ab · inbound

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B cites this paper.

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T14:52:12.591220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:52:12.591220Z digest=sha256:666daae2c13f31cf59e3341dff285beaca16649a85d958bb25f27c24387086b9

Observation 2fd0eeb0-ba99-451f-a1ed-ebdbb1a87b24 · inbound

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) cites this paper.

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:07:16.653669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:07:16.653669Z digest=sha256:aba62f9166d0e0c5a9ad86a06a583735a0398bf102d035ab97b14a0b6411c5c9