Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:40:31.250299Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2509.12672.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:40:31.250299Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 618cc52e-9b6b-4065-b6ed-9e5c3d55ef56 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content , " * write output.state after.block = add.period write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46cdf308-c29a-430e-aa6a-821fff51889b · outbound
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3614bcb-5916-431c-86c4-dcebb732c7e2 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bc19f82-69af-4133-aee9-b35823587b2e · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b74c5b7a-dd7d-4c18-81a7-a2f8c38e83cc · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Mechanistic Interpretability for AI Safety -- A Review
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072c8864-8fc0-4f1c-94cf-4981912c04ea · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Towards Building a Robust Toxicity Predictor
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 247f0cce-2568-4a35-a002-e7f5d21532c6 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38b51520-2d7a-4a6f-9c57-3c6830720a52 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1390aade-ade4-4013-b84e-7deb1b97cd2f · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e91bfd-b394-45b8-af01-0d6b5052dbca · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e2e582d-84cf-4b6f-af44-3536a0ba7fa2 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Towards Automated Circuit Discovery for Mechanistic Interpretability
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1133c7a5-f396-47a1-9fb7-8b5b823790ba · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8163cfac-a653-45b1-8e2d-aaa7b94838d6 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a59922d2-4c9d-4813-8857-b2829dc5399b · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aaad299-99cd-4808-9824-4eb97b536c99 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 140d4b5f-71ba-4710-b87a-debafd5c72a4 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af396301-d35d-4ef6-a04d-098c5fb597c9 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94337268-ddd3-4ed5-bd4a-6ba7e201cbd3 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26eea721-1a69-4a69-a7b5-ac85284641c2 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50d3bb49-f4db-49aa-92d6-7022dbbacd81 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc9cd95e-1d06-44e1-9fe2-c423bae08b88 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf8c2438-d63d-4082-803e-d1245e896c3f · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a04ecd64-ce82-4bcb-b222-78843b17a1a3 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content A New Generation of Perspective API: Efficient Multilingual Character-level Transformers
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63aee8c0-e72c-4077-a558-d254dc22d77c · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a99733c-38c0-447e-aaf8-0988630fe711 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8773139-94a8-4f6e-a960-35935bf51f31 · outbound
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b20839b2-68a0-4c99-9fbe-5a70fe119bb5 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content R.; Li, G.; and Crespi, N
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ff28abf-a7b1-45b6-8beb-025ff87e9613 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b40be0b-e29d-4760-b373-2f63355c226f · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad8a4e2f-5d17-4a7f-9ef6-42f3e805f4a1 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1b7ddb6-f539-44e8-a6dd-12d2629bb605 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 293bf6f3-d479-4531-a742-f2abe67f3e14 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0ec7491-8aa0-4f1d-8f17-166cf60effcd · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cce6712f-334d-4d41-9a63-1f8030cf7501 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aaaba9b-adf5-48c8-a93f-43ac7d596243 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content o ck, F.; and Wagner, C. 2021. “Call me sexist, but
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2ae48b7-6b5d-4803-9419-1809e0dcd2f0 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content HowkGPT: Investigating the Detection of ChatGPT-generated University Student Homework through Context-Aware Perplexity Analysis
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86ca1c4e-0812-4f9a-83dd-4a60ca6afa21 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b27728d-2133-4603-9263-342399895f24 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e72fdbae-b950-4752-92da-431093f8e751 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b88280b-d6f3-4674-a4ee-071c97d11a47 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25799df4-b1cf-4a0a-8220-d889315bb4c5 · outbound
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9849187-ef1a-42d2-ac4e-7c51ab90eeae · outbound
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fd0ebd5-059c-4277-aa79-70989e890b6a · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f925dcdf-a08b-4a7c-8316-fda551ff1c15 · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc5dfd60-e7b4-4b2a-9700-f45348ff631c · outbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.