Pith. sign in

Paper Citation Record · LEDGER

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns

As of 19 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 1 inbound Pith citation observation for arXiv:2501.16750.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16750 v1

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T11:03:29.370944Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T01:02:51.934157Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T01:02:53.997511Z

Reference resolution

100 of 100 outbound references displayed

  • verified exact3
  • verified fuzzy53
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ad7f745-cbff-421d-9ec5-fc59bfb5fbbb · outbound

This paper cites https://en.wikipedia.org/wiki/ Coleman-Liau_index.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://en.wikipedia.org/wiki/ Coleman-Liau_index

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.902823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.902823Z digest=sha256:6b11086470ebb73d360e3c11db060683cb05da087ae380fa04a3df65cb2fd7ef

Observation 03b1fd37-1304-4c6b-ad09-a10ebfb26c4c · outbound

This paper cites https://github.com/unitaryai/detoxify.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://github.com/unitaryai/detoxify

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.907255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.907255Z digest=sha256:01e6a58f98071a31c9890430e4057de22f4ea2a927d1ac8503e79a2553be180a

Observation d4f50355-d933-4a08-88d7-26d878d1b1b5 · outbound

This paper cites https://gdpr-info.eu/.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://gdpr-info.eu/

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.911411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.911411Z digest=sha256:37a74b21f53bfc284fd669266b6e445dba25be13cdb4902cf7b5ac661b1cc061

Observation 6dcf60a8-89d3-48d1-939a-e30de2bde689 · outbound

This paper cites https://chatgpt.com/g/g-w0y3CvDM 9-freddy-griffin.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://chatgpt.com/g/g-w0y3CvDM 9-freddy-griffin

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.915052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.915052Z digest=sha256:d1f50cc7ac6f35601e00e66c899cb05d5d7c55e8dd81d65ff2a56cf28facb9eb

Observation a63bdbfd-cad7-412f-9a4b-8f93bc1429cd · outbound

This paper cites https://chatgpt.com/g/g-JlQ9WBdHB-hate /.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://chatgpt.com/g/g-JlQ9WBdHB-hate /

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.919260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.919260Z digest=sha256:17d8f80a023cb789fe2165c9aa31d4633cc55b7d3c9eb6aabe272fa9eed26a98

Observation fd1fe658-5211-4357-b5f7-f4cc5a9b7645 · outbound

This paper cites https://chatgpt.com/g/g-87uTmBE65-ru de-gpt.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://chatgpt.com/g/g-87uTmBE65-ru de-gpt

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.937731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.937731Z digest=sha256:fb55b636236603116441f0d0c74ecc612fd41fcadc5f2715f4cb8ca962ba2048

Observation 0bbde68a-1663-43f4-a4ce-e354f24e7abe · outbound

This paper cites https://www.perspectiveapi.com.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://www.perspectiveapi.com

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.967561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.967561Z digest=sha256:2f031039e89fde2ea2effe00c191010fec09915b769ea0be1ef0de7687c242a7

Observation 3ac119e8-d12a-4ca5-a0b5-a24c13c232be · outbound

This paper cites https://osf.io/edua3/.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://osf.io/edua3/

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:27.987546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:27.987546Z digest=sha256:ab6b258ad411bdf80b98c61ed6f37f42feea3fb7ca8c7e027007fe7685f765e3

Observation 8edefd31-f421-45fc-b7bc-f66f21388a73 · outbound

This paper cites https://lmsys.org/blog/2023-03-30-vicuna/.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns https://lmsys.org/blog/2023-03-30-vicuna/

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.018050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.018050Z digest=sha256:31ffea6f4aba51f7c45977527328450cd7118fa373f23b46e5094504c045a00e

Observation ebaeb6ad-bb19-4144-8aed-535777dd7068 · outbound

This paper cites ADL Task Force Issues Report Detailing Widespread Anti-Semitic Harassment of Journalists on Twitter During 2016 Campaign.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns ADL Task Force Issues Report Detailing Widespread Anti-Semitic Harassment of Journalists on Twitter During 2016 Campaign

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.048010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.048010Z digest=sha256:a7e5b06fb79c04d27a82d829ac896309cc3a718f32bab280319600f12a811340

Observation e659f29b-ae11-4dec-bf41-e49ed0e947be · outbound

This paper cites Online Hate and Harassment: The American Experi- ence 2023.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Online Hate and Harassment: The American Experi- ence 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.056730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.056730Z digest=sha256:84e759f97782b247985cba9097d0eaf9699e713d57e6aeec7e335c5bebe4d2dc

Observation 350e5bd4-1ec6-429b-b67f-34892d7950fc · outbound

This paper cites Aunties, Strangers, and the FBI: Online Privacy Concerns and Experiences of Muslim- American Women.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Aunties, Strangers, and the FBI: Online Privacy Concerns and Experiences of Muslim- American Women

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.094242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.094242Z digest=sha256:5922adfa2839766cfd20706c5df406926018273256b706f5809454ab0e2c09a1

Observation ef305288-9568-414e-bd57-75a81094b2ab · outbound

This paper cites Google’s Jigsaw was trying to fight toxic speech with AI.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Google’s Jigsaw was trying to fight toxic speech with AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.127008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.127008Z digest=sha256:95abdb91816b4347800af6614c81b339e4d75bcb8a8ca1560be7a4a435a86c28

Observation 8a0f5ff9-abec-4355-a5b8-1cdd73640df6 · outbound

This paper cites Robust Hate Speech Detection in Social Media: A Cross-Dataset Empirical Evaluation.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Robust Hate Speech Detection in Social Media: A Cross-Dataset Empirical Evaluation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-10T11:03:29.690850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.161185Z digest=sha256:710d976dd9e0ff4393a8383a9223897fc8eca234347592a528744fcaae5ac537

Observation 085a2934-d5cf-43a2-977e-d98ce24af45b · outbound

This paper cites METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.189110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.189110Z digest=sha256:6d49ef2f9c14bb3c3c2745d57b65fe02b9905db08671cacb42dea2cb245a9440

Observation 20622d68-2f80-4d1b-bf9d-8f72e7490f70 · outbound

This paper cites The Pushshift Reddit Dataset.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns The Pushshift Reddit Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.199948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.199948Z digest=sha256:b990abc126a52491f90b799a98110e44761c66d5dfd22c121a5f21a18613f6fb

Observation b2799bf0-0fce-4029-a73a-5d097b6c6e68 · outbound

This paper cites Nuanced Metrics for Measuring Unin- tended Bias with Real Data for Text Classification.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Nuanced Metrics for Measuring Unin- tended Bias with Real Data for Text Classification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.218992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.218992Z digest=sha256:ec535ed0297961b628fcfa85a80049592a969c4a3b950e74bd2189a63d54fd39

Observation 7793d8ca-21ff-48cf-997d-30f68d8a965d · outbound

This paper cites HateGAN: Adversarial Generative-Based Data Augmentation for Hate Speech Detec- tion.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns HateGAN: Adversarial Generative-Based Data Augmentation for Hate Speech Detec- tion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.232006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.232006Z digest=sha256:438625888be09a982a08e84f833d46c44f4cc70a50a2caa9c2309ae181fa9b42

Observation e8e5201c-059b-43ca-bd76-21bab6bfc547 · outbound

This paper cites John, Noah Constant, Mario Guajardo- Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns John, Noah Constant, Mario Guajardo- Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.241154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.241154Z digest=sha256:4cc801b1757e8f00b60c534d3e776ae53fabc7476f6ef9a10d6bdecd244899b4

Observation 9d5aa832-e0aa-453f-b517-7e8ed0cccfec · outbound

This paper cites Hate is not Binary: Studying Abusive Behavior of #GamerGate on Twitter.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Hate is not Binary: Studying Abusive Behavior of #GamerGate on Twitter

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.254179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.254179Z digest=sha256:444cee2a93ee3719a96a45a0882022896f591f199f4ab1a13d734f244c921174

Observation a93cd6a6-24b0-4a66-8eb4-3a6261c457f8 · outbound

This paper cites Christiano, Jan Leike, Tom B.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Christiano, Jan Leike, Tom B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.269724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.269724Z digest=sha256:0fd5104ac21aaa7f3aa7616c6cb3897025453927a17a46071f7b7a04903e0eb0

Observation 2df97a1e-c6ba-4887-bd58-49efbe1fe76c · outbound

This paper cites Hate Campaign.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Hate Campaign

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.380797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.284063Z digest=sha256:0128a07ca2d3f3e73245506786d4d7193981cf76715db2db2300b806a93cd6a4

Observation 99b942b8-788a-41f9-a1d3-8136f5181f24 · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.368765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.296688Z digest=sha256:c5b810840871f12fc7ee4c893a5c052279bdd239d0286b1e35d88882f770d6bd

Observation 6bc594ad-b278-462e-b48c-96687bb8aec7 · outbound

This paper cites Toxicity in ChatGPT: Analyzing Persona-assigned Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Toxicity in ChatGPT: Analyzing Persona-assigned Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.304717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.304717Z digest=sha256:593769092bceb1ce8feebcb21f44f5d41a3fd7324a79d3134cda5417fc686fb8

Observation 28de4e90-701c-4e55-b3dc-17e14de7a540 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Trans- formers for Language Understanding.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns BERT: Pre-training of Deep Bidirectional Trans- formers for Language Understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.357619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.309122Z digest=sha256:ae785cb8ca9a67a13e8223f55f138331b8fc93448ddbdeb0c7993225f4a00685

Observation eea97b84-2711-4c36-8827-81c04e587bf6 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:31.347120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.315782Z digest=sha256:806e1a1297800a233d44c6d8acab7bc47ee907b10871d2e4ae2dd401ede7dccd

Observation 613a99d5-4c1b-4084-b042-e22588cb4e20 · outbound

This paper cites Paraphrase a text.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Paraphrase a text

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.337317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.327174Z digest=sha256:8b890743fce07672c8fff78f302260a1d9fb0084b5678f1d7f33a280f7255d12

Observation 76ba4c38-f8d0-461e-b25b-b4f8cff0cc8f · outbound

This paper cites Guide: Large Language Models-Generated Fraud, Malware, and Vulnerabilities.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Guide: Large Language Models-Generated Fraud, Malware, and Vulnerabilities

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.326970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.333912Z digest=sha256:e829fa30004fafff3f376413ef4f34fc8aaafa0fe35f4b2b21e487975d57803e

Observation 04699b60-c2af-4c3b-b9f2-8c5f1663c1fe · outbound

This paper cites Black-Box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Black-Box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.315813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.350922Z digest=sha256:18ae2b7d48edfea1b08ab1bf0e31f4d71355ca4172d7b79e4c0ced1fd40662c3

Observation c7948ace-3c0f-4252-a6a9-4621348310dc · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.360191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.360191Z digest=sha256:6721eada002a4702936c42f8b96002be786edbd6e87de216e03d30d847836273

Observation 7352fa66-8c9f-435c-a322-6dd973b7e909 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:31.304561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.364370Z digest=sha256:02e38848bfceb43dd7eee03e5bc897ed9439802e151f949e3c19087ea11664a5

Observation 8a497de1-fa71-4e2f-a6e4-cbf5a18d9def · outbound

This paper cites Hancock, and Zakir Durumeric.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Hancock, and Zakir Durumeric

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.289569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.368006Z digest=sha256:c9884ac5d2f5a1903ab0443f2c0b75c9bdf860c2321678d19320cfdb98c18932

Observation 3fe4fc75-27d8-45af-ab7c-d413cf709c09 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.267866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.372041Z digest=sha256:08294902a0dd6ceca367cff6f66548cfc82501cd53a521140cb97781f2c8f4f1

Observation c02e3e85-f651-41d7-aad1-4cec723a7953 · outbound

This paper cites Hess, Kelley P.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Hess, Kelley P

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.213406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.378471Z digest=sha256:64c1994164219cf8e7d39a62a50ae1981a8dc55fb7399ea3534cae8c13274f2c

Observation cdf58a07-63fb-49cd-8294-29708f51959c · outbound

This paper cites Deceiving Google's Perspective API Built for Detecting Toxic Comments.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Deceiving Google's Perspective API Built for Detecting Toxic Comments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.386994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.386994Z digest=sha256:22bcc133f0e79fddbecd9d08ccad1362937192af78e54e5e85b6ba6406d7f182

Observation c7bd1e3f-72df-40d0-b70b-8a873d84cdfb · outbound

This paper cites Adversarial Example Generation with Syntactically Controlled Paraphrase Networks.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Adversarial Example Generation with Syntactically Controlled Paraphrase Networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.149871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.390871Z digest=sha256:4a197c8e51de2d6adf202ba4c8d6b9e2b25aaaf3433ed438810c4624261a72ac

Observation 867775fb-fa87-4bb4-ab05-25d298584f0d · outbound

This paper cites High Accuracy and High Fi- delity Extraction of Neural Networks.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns High Accuracy and High Fi- delity Extraction of Neural Networks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.082157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.416115Z digest=sha256:01d2170dc2c95af348258861287a65faa9613016cacb0ebef524f74a19d5f412

Observation 18b20293-f9a8-4252-9dcc-d9f66218911d · outbound

This paper cites Is BERT Really Robust? A Strong Baseline for Natural Lan- guage Attack on Text Classification and Entailment.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Is BERT Really Robust? A Strong Baseline for Natural Lan- guage Attack on Text Classification and Entailment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.447493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.447493Z digest=sha256:86597dd580a60694a54b8b0d4e1215542f9669c2f85bd4d8c783024c43ba9f2e

Observation 68f67a60-d589-40c7-854f-bb23bb8592e4 · outbound

This paper cites Toxic Comment Classification Challenge, 2017.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Toxic Comment Classification Challenge, 2017

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.055167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.484402Z digest=sha256:3d0eb8794c655d3eeab024f6b2e1e252e93a903b6fca0a0fc7f3ea19e413b4ee

Observation 850c24bd-9013-43db-a4cb-0eb5d49bbb12 · outbound

This paper cites Jigsaw Unintended Bias in Toxicity Classification,.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Jigsaw Unintended Bias in Toxicity Classification,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.042736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.491766Z digest=sha256:f5ce5f4fa1455120d10ed57bc3c070edc408e2bffbf8cb0406fba4783b9a2499

Observation 3560f2f0-ae31-4139-b05c-a2b323d259dc · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:31.030756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.537604Z digest=sha256:14b20dbed90d0e0803ebbe687a58f51ea01d3baf4f8bb484336f3adcb1a54ace

Observation 1119b65d-bc8f-44bd-80f4-5b0d50529dbe · outbound

This paper cites Content Analysis: An Introduction to Its Methodology.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Content Analysis: An Introduction to Its Methodology

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.020120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.567684Z digest=sha256:d32abc632c6b9667e967d8379e21352ad69e46aa4b3c414c39acff4a7ed51f57

Observation fd7a6e8a-f462-44dc-b6ce-68e12e98ba1b · outbound

This paper cites Parikh, Nico- las Papernot, and Mohit Iyyer.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Parikh, Nico- las Papernot, and Mohit Iyyer

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:31.006299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.590059Z digest=sha256:723a19a44ab0c840caf7ce679b2e569f6b036d4e7c152933fee283d927f74b56

Observation af4b1e44-f981-4f32-92e8-7b956d279b36 · outbound

This paper cites TweetBLM: A Hate Speech Dataset and Analysis of Black Lives Matter-related Microblogs on Twitter.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns TweetBLM: A Hate Speech Dataset and Analysis of Black Lives Matter-related Microblogs on Twitter

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-10T11:03:29.570718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.620351Z digest=sha256:0330259e0ff6d8a5cef9eef372bf144bdb4fbff5c23ecd60f00379d96e58c9f5

Observation 5ba05853-35ad-47aa-80bb-f9575da20453 · outbound

This paper cites TextBugger: Generating Adversarial Text Against Real-world Applications.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns TextBugger: Generating Adversarial Text Against Real-world Applications

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.995160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.652938Z digest=sha256:23947420d86cde17fe77d9130225563287d418d28f7ac271b520087875d5bdd3

Observation 4d76440f-0804-46ca-aff2-72fd0895ebf8 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.689441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.689441Z digest=sha256:59267758317e494df0beeb73981ddb17401ac6afb52268fec68aaf9707a0356b

Observation cd03df65-d1a3-4b55-95ec-c71b069e6fb4 · outbound

This paper cites A Holistic Approach to Undesired Content Detection in the Real World.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns A Holistic Approach to Undesired Content Detection in the Real World

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.983943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.710325Z digest=sha256:bcaa131264604703409340aed6b6fabbcd6aef99bb694374b054983492807419

Observation 0a6cefe5-9337-45eb-991d-8652c1d33b41 · outbound

This paper cites HateX- plain: A Benchmark Dataset for Explainable Hate Speech De- 14 tection.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns HateX- plain: A Benchmark Dataset for Explainable Hate Speech De- 14 tection

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.958534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.727341Z digest=sha256:87bcef63ff035eeca4c168cc89d39e7fd0fd92891547b6b0712c52aa5df07bbf

Observation 515c39ea-1658-4802-a2eb-f09dec649daa · outbound

This paper cites AI Trained on 4Chan Becomes ‘Hate Speech Machine’.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns AI Trained on 4Chan Becomes ‘Hate Speech Machine’

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.930790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.739293Z digest=sha256:bb778af162b3d68724186bf16893eaaa24d27681a3c1bb848a81369edef04c95

Observation fbfdfdfb-1768-4df7-a7e6-fe61556966a1 · outbound

This paper cites Mazurek, Florian Schaub, and Elissa M.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Mazurek, Florian Schaub, and Elissa M

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.892022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.749638Z digest=sha256:4b1b17aef7b6d7d697b992731080b872d8da670e9a5e3a91ecae891197db4ec0

Observation 76e81ca1-999c-42c7-93f7-be9c5dde4df2 · outbound

This paper cites The Challenge of Detecting Hate Speech.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns The Challenge of Detecting Hate Speech

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.796736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.766621Z digest=sha256:86b70823a03ece829d497c1743a03fe485b14eb304c78e6379cbf3c734e8656c

Observation df5f1a68-3d31-4125-893d-ab986c018863 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:30.746708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.783465Z digest=sha256:18930a1f0eb27358ae7fe90d668b9d14377a49087653a4f53edf960fccc13b35

Observation 54135c04-8df8-43c4-97a4-01d6cffc24ee · outbound

This paper cites What is hate speech.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns What is hate speech

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.735138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.800210Z digest=sha256:872250092e481cfe9051bef5313b200546cd44b7dc433cab77cab4be50389048

Observation 4c2f3670-5515-45eb-8bde-5f11cd7172ee · outbound

This paper cites Handling Disagreement in Hate Speech Modelling.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Handling Disagreement in Hate Speech Modelling

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.723835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.805787Z digest=sha256:cbffd3633ab98461c2b5e540e8a6288b07fd62b9275fb0a0c160de8f0d88d5b5

Observation e5466900-68dc-497c-b282-cf1c31e194ff · outbound

This paper cites I Know What You Trained Last Summer: A Survey on Stealing Ma- chine Learning Models and Defences.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns I Know What You Trained Last Summer: A Survey on Stealing Ma- chine Learning Models and Defences

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.712642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.810279Z digest=sha256:ab2240a1de521c90cb2f4aae8069b127fc9c7d3ba3af41cd10655fd1f9cfee3b

Observation dd69ecf4-ebe1-418a-9767-1f520db815f3 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:30.700183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.816793Z digest=sha256:5d39524dd8fc10b40d6c9d2c54aace2aec157ca17de13ade2c71321aac9065d5

Observation 566f0571-34f3-4fe8-8d60-856a0dd5e825 · outbound

This paper cites Introducing GPTs.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Introducing GPTs

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.688919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.830570Z digest=sha256:557e58ea762a7dbd17c548898ade219bd16686cbd328b562cd3f6819a7004fdc

Observation be5b8a04-8b95-4a9e-b688-049aef60b248 · outbound

This paper cites GPT-4 Technical Report.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns GPT-4 Technical Report

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.836125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.836125Z digest=sha256:41eefb8f5ac859d20feb627e275a591701f195c0389cb9b11ccc743f964dbb93

Observation 055fc79f-3b4f-4522-bd68-cddc1c2c34b6 · outbound

This paper cites Offensive and Hateful Text Multiclassification.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Offensive and Hateful Text Multiclassification

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.678911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.847019Z digest=sha256:efaac1b577efba1428ad38e3cf08a0caa872d11c6eecb4a4dc7f872a04cc8c70

Observation 4512db10-1eff-4c35-b2cd-94e30131da9d · outbound

This paper cites Facebook’s race-blind practices around hate speech came at the expense of Black users, new docu- ments show.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Facebook’s race-blind practices around hate speech came at the expense of Black users, new docu- ments show

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.658933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.857094Z digest=sha256:8fa0d75208681ff60cc2ac7f9ad7eea0b058b28c3a5f70ffec713ccf709ce86d

Observation cc9834e9-ce0b-4570-978c-f1483673bc5a · outbound

This paper cites UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.872472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.872472Z digest=sha256:2c7c946dafc9c90afd7a33346a4b54fd6e0e60c654d42a4135dc46f3bdf85604

Observation bc3954aa-717b-4681-abc0-a62a2463d96d · outbound

This paper cites Gener- ating Natural Language Adversarial Examples through Proba- bility Weighted Word Saliency.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Gener- ating Natural Language Adversarial Examples through Proba- bility Weighted Word Saliency

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.626013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.884560Z digest=sha256:1092fb2a7a5d074b3f63c29b8d4eafeba08bb039665eaee4509270d824dfb366

Observation 5d432944-2249-4239-95ba-b70f0ef1aa70 · outbound

This paper cites Schuller.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Schuller

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.571183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.891454Z digest=sha256:776a48a0e857dba2a13336532fd2298d656bb9c227bce1aaee19bfc1014c74fd

Observation 10176289-a493-4eb0-9720-25beae53e1c8 · outbound

This paper cites Margetts, and Janet B.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Margetts, and Janet B

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.505330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.902216Z digest=sha256:9531401ab66a39d14bf5f628cf6967bb738fae6c67359f2606b8bcf8cbdbdfff

Observation 5c772455-2cf9-45d1-aa32-3115a565f95c · outbound

This paper cites Sachdeva, Renata Barreto, Geoff Bacon, Alexander Sahn, Claudia von Vacano, and Chris J.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Sachdeva, Renata Barreto, Geoff Bacon, Alexander Sahn, Claudia von Vacano, and Chris J

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.443534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.906112Z digest=sha256:187f2b0714e924f43a01a026a8d475d5a76e6699e7b73eb2ffc4ddf73faa9562

Observation 574f0623-0853-4351-997b-924c24e0f4f9 · outbound

This paper cites TUBERAIDER: Attributing Coordinated Hate Attacks on YouTube Videos to their Source Communities.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns TUBERAIDER: Attributing Coordinated Hate Attacks on YouTube Videos to their Source Communities

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.373777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.914055Z digest=sha256:24873bc68db1e22b7954f6fe8c4737422be66ee5517c724e97a8a6f693c83c55

Observation 3dd02bca-aeeb-470d-9dc9-b95fc2e2d14b · outbound

This paper cites Generative AI as a Vector for Harassment and Harm.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Generative AI as a Vector for Harassment and Harm

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.361183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.917853Z digest=sha256:929217144019bf9686cc3aa73b6fbc9595dc63352af1406f9006b56cfb5b7b06

Observation b43cb975-e63a-4db4-af32-ce6610a0568f · outbound

This paper cites Man who harassed black student online must deliver ‘sincere’ apology, renounce white supremacy.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Man who harassed black student online must deliver ‘sincere’ apology, renounce white supremacy

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.350491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.922110Z digest=sha256:9dc1e43402adb6b073706a1edcda6f16643eb46f635b96d6aba2bcb7af3d7296

Observation c945fcff-def5-4e11-b494-4b0f1d8c8bb8 · outbound

This paper cites Do Anything Now: Characterizing and Evaluat- ing In-The-Wild Jailbreak Prompts on Large Language Mod- els.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Do Anything Now: Characterizing and Evaluat- ing In-The-Wild Jailbreak Prompts on Large Language Mod- els

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.338709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.926135Z digest=sha256:9f7d655dbaf2cb8d4076cd98f8666c2468145cf681d2b33ad884f7f5882cb089

Observation 8cd9120a-fe98-49f3-845c-8e79e8e9c624 · outbound

This paper cites On Xing Tian and the Perseverance of Anti-China Sentiment Online.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns On Xing Tian and the Perseverance of Anti-China Sentiment Online

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.327376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.930210Z digest=sha256:2901564639ff7db6dc326a83efc0638b8c61e7dc57d94468db5d6ff6585aca42

Observation 50ec7bac-64e6-4f6e-8608-02e913eced0b · outbound

This paper cites Model Stealing Attacks Against Inductive Graph Neural Networks.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Model Stealing Attacks Against Inductive Graph Neural Networks

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.316013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.933692Z digest=sha256:9b754147b05065936a540280501648d3ebcaec364bf7be6cb480280794d9f065

Observation efd0ea7a-3324-4732-bbf3-3ddaf119f503 · outbound

This paper cites Analyzing the Targets of Hate in Online Social Media.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Analyzing the Targets of Hate in Online Social Media

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.304658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.937591Z digest=sha256:80e7e86fc60f49c1e91af321744f7627b70d37fba3c2516c89c6486c0f3e4caa

Observation 4f3f5cb9-3a3b-48a8-877b-806ceca6bd09 · outbound

This paper cites Learning to summarize from human feedback.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Learning to summarize from human feedback

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:28.955630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:28.955630Z digest=sha256:7a407eb4a8a928318f9d71fecfe1988f46bf56cbd66586ffe207d2783fd93557

Observation 2ef07ce3-c05c-4fd6-ac2a-f360ec628db1 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:30.290822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.980546Z digest=sha256:5bb2196377eb5a4f1c76523dd276d28e70f417adbbfbae6dcefb7971bc76c187

Observation 0c967b31-044a-49d2-a191-b7678e5d6db1 · outbound

This paper cites Large-Scale Hate Speech Detection with Cross-Domain Transfer.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Large-Scale Hate Speech Detection with Cross-Domain Transfer

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.210478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.015186Z digest=sha256:07ca07f30ab3642acbbc83b762422371214086d1c0376693d22cfe3df66a987d

Observation e1d50bf6-b164-47f5-9bf6-7978ac6476dc · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns LLaMA: Open and Efficient Foundation Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.053661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.053661Z digest=sha256:5dc2cc76102bc59bbb6d0a545b1ec466e283d189fcfc797b598410e2a0ba7af4

Observation 4913f0b0-d5dc-4d5d-8ad7-f434389b6bda · outbound

This paper cites Reiter, and Thomas Ristenpart.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Reiter, and Thomas Ristenpart

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.194371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.071955Z digest=sha256:23ea239eb3d79a3af14d9db70dc83f0b58fb38e997b00550a1bbe905ef4db742

Observation a76556e8-7611-4441-91cb-99eca8b791e6 · outbound

This paper cites Visualizing Data using t-SNE.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Visualizing Data using t-SNE

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.094145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.094145Z digest=sha256:97d90f8e51c5b599a499ceace7766eb7423dfddbc40d885fcf7f5732e1be20a9

Observation 62f85124-fd18-4fc4-b752-6cbff76e2090 · outbound

This paper cites Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.175553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.123862Z digest=sha256:8f076a1c293b8e8f6be26ede77677461cdfbeea8da94d846cb5ccc7ddbcffb0c

Observation d3ad8727-8c8a-46d9-8470-645e194b289e · outbound

This paper cites Moderating New Waves of Online Hate with Chain-of- Thought Reasoning in Large Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Moderating New Waves of Online Hate with Chain-of- Thought Reasoning in Large Language Models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.164132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.159151Z digest=sha256:bc30c8fffd4744af571def2c67bf2d39851570b185de15a92a9d9fde0216aea4

Observation a7abee6c-1f43-47e3-9d9c-7ef49454d838 · outbound

This paper cites Vu, Alice Hutchings, and Ross J.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Vu, Alice Hutchings, and Ross J

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.152770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.193217Z digest=sha256:ca33a982b40a0f828cfc1bec13e79a80645963803fa387dff7319da78434a513

Observation 3ffcd2df-e5bf-4d0c-ba60-94956a9dc078 · outbound

This paper cites There’s so much responsibility on users right now: Expert Advice for Staying Safer From Hate and Harassment.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns There’s so much responsibility on users right now: Expert Advice for Staying Safer From Hate and Harassment

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.125214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.206641Z digest=sha256:3b08b06dc9717e2d26f71787bbe17ec4525cee89fd572a9d84d41521176c8108

Observation ee32c18a-1f36-4b04-b4aa-bc5d59673f41 · outbound

This paper cites Challenges in Detoxifying Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Challenges in Detoxifying Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.226885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.226885Z digest=sha256:acaca7ae37f9562926b429f4b9528097c359307208a7512cf5637268c7c3da38

Observation 10dcf881-889b-49a2-bdee-0601e963f804 · outbound

This paper cites Not All Asians are the Same: A Disaggregated Approach to Identify- ing Anti-Asian Racism in Social Media.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Not All Asians are the Same: A Disaggregated Approach to Identify- ing Anti-Asian Racism in Social Media

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.097329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.232894Z digest=sha256:8540ad5168ca76be6b24d39645a5c34aad645768d8cf395db7fbf24d3c68c428

Observation cc1edeb2-f072-4d71-9e79-b69876745678 · outbound

This paper cites Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.058739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.239336Z digest=sha256:efd5a58369e9367e56aa17e3d4ddeb6fd5043c341bb4ef22edd4b6d92e9b2cfb

Observation afb446ee-7248-4722-8ea3-e0267bce9bd5 · outbound

This paper cites Fight Fire with Fire: Fine-tuning Hate Detectors using Large Samples of Generated Hate Speech.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Fight Fire with Fire: Fine-tuning Hate Detectors using Large Samples of Generated Hate Speech

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.928514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.244624Z digest=sha256:fd98c54084195fed4a040e00a4771433c9030138cec5c6c0c8911ed88463adbe

Observation 7d101714-7083-4b0b-9768-8b2465c0bf0d · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Baichuan 2: Open Large-scale Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.254537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.254537Z digest=sha256:6683008b919b438b420558f9d8a37e9c244d6c3a46db03e60fa6cc6419e0f04f

Observation 5e69bce8-4c74-4ab1-b41a-2832fbc0011e · outbound

This paper cites GPT-4chan: This is the worst AI ever.https: //tinyurl.com/2s4jh5p4, 2022.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns GPT-4chan: This is the worst AI ever.https: //tinyurl.com/2s4jh5p4, 2022

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.916299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.263941Z digest=sha256:3f6318b2dfdf70a6f57511aa20f87d852f986bae22a1895c1454af3bdd8ea196

Observation 854195b2-ce6d-410d-84fb-d210bef4997a · outbound

This paper cites OpenAttack: An Open-source Textual Adversarial At- tack Toolkit.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns OpenAttack: An Open-source Textual Adversarial At- tack Toolkit

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.900311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.276493Z digest=sha256:ec2c47c2960c61f3a10db10f03a094bf0550feb7295e5f2659480bdffb012af0

Observation 19b25f2e-afc1-4a23-8ff5-27052e604aa4 · outbound

This paper cites SecurityNet: Assess- ing Machine Learning Vulnerabilities on Public Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns SecurityNet: Assess- ing Machine Learning Vulnerabilities on Public Models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.883125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.293814Z digest=sha256:ea96f526f5a62fcdff0c5af5287874a5976bb474b7f00520df075b66e1df1acd

Observation c9b65d46-ebc3-4c3d-990a-fc43901af425 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns OPT: Open Pre-trained Transformer Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.306584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.306584Z digest=sha256:33c33476384fa1b9a20ea068195aa30527d38382b220671d499463e6e497c964

Observation cb337bfa-4dfa-43ab-a3e2-8bd60141f990 · outbound

This paper cites Generating Natural Adversarial Examples.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Generating Natural Adversarial Examples

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.870122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.319248Z digest=sha256:fdad12fad36c0cedc3b52db9b25f003e901944d6199003a775a6b98ea76a7489

Observation 2d3426a3-2e04-4116-adc9-8c80586222b1 · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T11:03:29.331421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:03:29.331421Z digest=sha256:96652294136410e9ce47d8a57e6693df4695f4887a7c01f331f5e1566e8c7946

Observation 04791f2d-d5d0-434a-8917-1add3bcf9edc · outbound

This paper cites 2, 3, 5, 17, 18.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns 2, 3, 5, 17, 18

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:30.390138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:28.910119Z digest=sha256:929d60f9600173d69b16156cb12ea91e7ef3841480af88ba75fcf736d093e1ff

Observation 9cc12a26-8c36-4a77-8b09-b58c130cb60a · outbound

This paper cites Racism is a Virus: Anti-Asian Hate and Counterspeech in Social Media during the COVID-19 Crisis.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Racism is a Virus: Anti-Asian Hate and Counterspeech in Social Media during the COVID-19 Crisis

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-08-10T11:03:29.417012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.345163Z digest=sha256:2949ad4462f90ae90c56e441980a45897f21e3285a486eea4bdba8ec07eaa224

Observation f44ee80e-a821-4bb2-b532-13a3b18b682a · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:29.845892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.359653Z digest=sha256:ca36f777bb7a06542c57ff89ba16c1159ab1178d23f0171d151a52f0cf898589

Observation e8b5c314-e0f9-4f8d-847e-cd06e98960e6 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:29.832288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.363030Z digest=sha256:c12491f7359dcd9a24ca40f2df9670fcc62c4b7bc1e70879a3b6edf187fbeab9

Observation 988831f8-7a0a-4902-a5af-3abb1bb49485 · outbound

This paper cites an unresolved cited work.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:03:29.788575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.367165Z digest=sha256:6162451a4f3a5fc45a93f070bc69610d79dc188260ef637e034c56f564ffdf85

Observation 81c29158-d397-4a44-9abc-f140e901b128 · outbound

This paper cites identity attack.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns identity attack

Reference 100

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T11:03:29.738939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.370944Z digest=sha256:f5aac622b106648a420a828fec3a4fb6f2cf6dcce57a13440d47241daa082abe

Observation c3846f5f-3496-42a5-8729-02c84077f5c3 · outbound

This paper cites toxicity,.

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns toxicity,

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:03:29.858217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T11:03:29.353557Z digest=sha256:7fb7e4537ac0ba262b6e7938f452f3bec0825a4f9518408bb39a0809749b9f1f

Pith citing papers

Observation 16bb6919-7b9a-4b14-bf1b-9cb3e6c48e0e · inbound

Are Today's LLMs Ready to Explain Well-Being Concepts? cites this paper.

Are Today's LLMs Ready to Explain Well-Being Concepts? HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T01:02:54.096747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T01:02:51.934157Z digest=sha256:ab58e9f50e01dc9c10edf60fb1f7d31405e721bfc44d8cebd46a76a954f0b1e9