Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:13:04.217122Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 251 outbound references and 6 inbound Pith citation observations for arXiv:2508.05775.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:13:04.217122Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:06:20.077337Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-25T07:05:26.724164Z
100 of 251 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 261d18ad-fab4-4f43-9dd3-ecb088b7fbac · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df7a80e7-73ef-4c35-b707-6f4833e9d6b6 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Certifying LLM Safety against Adversarial Prompting
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a90b26-9d41-46ec-b784-b65d7cb3fdac · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM and Arora, Simran and Mazeika, Manias and Hendrycks, Dan and Lin, Zinan and Cheng, Yu and Koyejo, Sanmi and Song, Dawn and Li, Bo , title =
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b639068-14e2-43aa-bcb9-7271a1414879 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caaf4c2e-3351-4090-8fc0-c064cb0b8f5f · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Discourse & Society , volume=
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56cbbbe5-c55e-4f2c-a79d-55d70c7ac078 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Advances in Neural Information Processing Systems , volume=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ebaa58-fc08-479b-8ccd-d8f0970e2863 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses , pages=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3024722d-7cfa-4993-86b6-52c8b1367780 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Findings of the association for computational linguistics: EMNLP 2023 , pages=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b62dcf92-388f-4ed0-812f-e6cd008e0d32 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM IEEE Transactions on Cognitive and Developmental Systems , year=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1852a291-a39b-4b13-9bd1-39bd7f229e87 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security , pages=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bce24a2-e0a2-49ac-ba41-f6ce07a77dcd · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 588c10bd-6657-46d9-8169-db93e98a0fa8 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cdb97cf-27b1-4e68-9b34-49645946c50e · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f7c2101-7c04-434e-bff4-bf7a40f3b1d4 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Security and Communication Networks , volume=
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af481b02-ce75-483f-923f-36f1c96e498c · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the First Workshop on Social Influence in Conversations (SICon 2023) , pages=
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc88f121-4999-4919-9bfb-05db88f4f74a · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Alignment is not sufficient to prevent large language models from generating harmful information: A psychoanalytic perspective
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be65711-c14d-4c93-828e-947d429cc5fc · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Systematic Rectification of Language Models via Dead-end Analysis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5f12ff1-141f-460e-8f75-c76ff8b69b78 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Successor Features for Efficient Multisubject Controlled Text Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 699178a1-2ce7-42cb-8f5a-055746c3e91c · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Expert-Guided Extinction of Toxic Tokens for Debiased Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfde5f27-fd0a-4007-bba6-ae62fc3df428 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78d438c3-a207-46c7-ae91-31d8267607d5 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fdd39d6-cfdd-4101-a7fb-caaac4a64647 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM 2024 IEEE Security and Privacy Workshops (SPW) , pages=
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32390486-8e5d-4ff5-916e-cbf37a9521c9 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM I’m fully who I am
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fed7984b-8f98-4228-b648-4a15ac91245a · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM From Text to MITRE Techniques: Exploring the Malicious Use of Large Language Models for Generating Cyber Attack Payloads
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988ccafc-7a5e-4aaa-8ef2-6504b7d2a28e · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Scientific Reports , volume=
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae41fc50-c3db-4742-887b-f06e3f739fa2 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Otolaryngology--Head and Neck Surgery , year=
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8bb058e-c906-4d8a-880f-626b48f65a86 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d56f67aa-6057-4433-8d0b-4a4e33158069 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c0ad005-a921-4062-870c-d34e4c8d7ed4 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Inclusivity in Large Language Models: Personality Traits and Gender Bias in Scientific Abstracts
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe70d7f0-6775-42a6-bfbb-7b348e29d36b · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The African Woman is Rhythmic and Soulful: An Investigation of Implicit Biases in LLM Open-ended Text Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3af60680-828a-4fa1-978f-01de151e4a68 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TATuP-Zeitschrift f
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d570c05-8e04-4d25-a096-1c8596171a84 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66a13d2b-461c-4acf-bf25-6351a60fde69 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Advances in Neural Information Processing Systems , volume=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbef038d-dd7b-4ab2-abae-e276cb072e90 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Attack Prompt Generation for Red Teaming and Defending Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58ed619-db02-496a-ac0b-a68b93782ad9 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Exploring the Adversarial Capabilities of Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4f56dd4-36fb-4e56-8077-3db47afe5ee1 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 485469a0-bdc5-4ccf-80e5-84675f2a08ea · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 042ae570-0dfa-4358-bcfb-f09ecfaf014c · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79103127-ed0f-494b-82bf-a8c035d439d7 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce3e6df4-cdd3-46e3-8ef4-7ed728a81656 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , year=
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40bdc7da-2a16-4037-9803-63721fd03371 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 173e19ea-556e-41c7-b0b1-fbe63074ad43 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Fifty Shades of Bias
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78ea9e60-7d1d-44ab-baff-8594fc267619 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM K o C o S a: K orean Context-aware Sarcasm Detection Dataset
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0203ae01-2eae-47bf-a360-bd268aa7305e · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM ParaFusion: A Large-Scale LLM-Driven English Paraphrase Dataset Infused with High-Quality Lexical and Syntactic Diversity
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb00353-2d98-4488-94c3-cc9b9f61f1b9 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7673df10-b083-4358-b07d-637e107f30ab · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14afb384-feac-4544-b26d-75ad2f90ccbd · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Text generation for dataset augmentation in security classification tasks
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa9f4d3-d632-40d9-a50d-d4ed6826d947 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Explore, Establish, Exploit: Red Teaming Language Models from Scratch
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b0c8281-d9ad-4be3-88b5-307123e3bc52 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Machine Learning , pages=
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8705ad8d-9cd4-466b-b034-18c9e58deb4f · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Adversarial Fine-Tuning of Language Models: An Iterative Optimisation Approach for the Generation and Detection of Problematic Content
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79165f25-f3b1-4ee8-a196-890c4414871e · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TroubleLLM: Align to Red Team Expert
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3620943-fba8-4cdf-8c73-9b0b3674bd14 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Learning diverse attacks on large language models for robust red-teaming and safety tuning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79871108-0d5c-49d3-bc79-dd7c6aee161c · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Outcome-Constrained Large Language Models for Countering Hate Speech
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f166ca7-d232-427c-9d3c-aa44f147d535 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the CHI Conference on Human Factors in Computing Systems , pages=
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8545ce46-1cb4-42ac-8900-8c4ca76e4704 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e9cd479-162b-4609-8a59-504cfcf47563 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9765cd3b-ca29-4b0b-8c55-9a5c99e9e018 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Seeing Seeds Beyond Weeds: Green Teaming Generative AI for Beneficial Uses
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb8d0e4-52c2-4072-94a7-aaa4be383cf5 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3aa813f-3e0d-4da8-a522-c6652f44ff70 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb41db75-9280-4d7b-b079-445d4b108a1e · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Creativity Has Left the Chat: The Price of Debiasing Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fb36029-f869-4632-84e2-4798f5ba304c · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Controlled Text Generation for Large Language Model with Dynamic Attribute Graphs
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e68a8cf-756a-4b8e-8909-781950102777 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fadb5382-8c1b-4e3f-88e9-d283ed392b10 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Yale JL & Tech
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f71c83e-dfd6-4dcf-b82e-b45d171c540b · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Engineering, Technology & Applied Science Research , volume=
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49673c17-e9d5-4249-9414-3d001d1ff6f3 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Machine-Generated Tweets , author=
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c6e79f-2a1f-468c-8481-f059a24e2c7b · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Natural Language Processing , pages=
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa620c31-7410-4c82-ad9c-610d6e605801 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Computational Science and Its Applications , pages=
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec361d9f-ec3f-4332-88fd-83b657ee20e4 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4613330a-9ecb-48f5-bb4f-e743afd3f148 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baf565a0-4225-4baa-9631-8a37500fc0ce · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Regulating Hate Speech Created by Generative AI , pages=
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 005b5e48-1803-4b5c-a922-4f6a68568577 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM A Study of Slang Representation Methods
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f761ee3-f867-44ab-ad42-3032a44c39ae · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Perplexed by Quality: A Perplexity-based Method for Adult and Harmful Content Detection in Multilingual Heterogeneous Web Data
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52f71db4-b053-4998-a6a1-ac5e0bbd3db4 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Human-Guided Fair Classification for Natural Language Processing
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65b692fc-8d19-4aab-9192-b7358227d056 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM arXiv preprint arXiv:2301.12534 , year=
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3042f772-ca5c-4b91-8fcb-c770d6a45a37 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Advances in Social Networks Analysis and Mining , pages=
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f7ed78f-2433-4640-90f3-a95f2603ed0a · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Explicit Toxicity Detection Models with Interactive Visualization , year=
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 164a9473-9137-4419-a9a1-adad34d60acc · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Which Argumentative Aspects of Hate Speech in Social Media can be reliably identified?
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25099d55-a1e2-4d1b-a80b-f0077fbc4f34 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI) , pages=
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a458cb7d-4282-49cc-bc73-972d9425e3d6 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803845f0-f56b-479d-9657-43c8d43a3d27 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Topological Data Mapping of Online Hate Speech, Misinformation, and General Mental Health: A Large Language Model Based Study
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53d99bd2-5e9f-4264-a72f-8f076dcc6532 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Demonstrations Are All You Need: Advancing Offensive Content Paraphrasing using In-Context Learning
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6661e22-6d2c-4b57-b6d2-934321f5527a · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Beyond plain toxic: building datasets for detection of flammable topics and inappropriate statements
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9798d37-cc31-433b-a0b5-a14eb29082c5 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Evaluation of ChatGPT and BERT-based models for Turkish hate speech detection
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ec65427-89d8-4191-b8ae-ac3903d98754 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM HateRephrase: Zero- and Few-Shot Reduction of Hate Intensity in Online Posts using Large Language Models
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10435c6c-a17d-4409-9d63-6408f817ac91 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 302948ef-6193-4c33-911b-47b1b120795a · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5365e4f-0f2d-47a2-8d76-5b266eafd26e · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78766c66-76e8-4a0a-9996-8cb3380fc3b7 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM 2024 , isbn =
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd78b58d-881b-4fd8-9d82-aad9612a5cf3 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Eagle: Ethical Dataset Given from Real Interactions
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54e798c2-c413-47ea-bb80-6a08a6c8fc0c · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the first workshop on language technology for equality, diversity and inclusion , pages=
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54f73beb-ca2e-4f98-b9b6-d76066cfccfe · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34bdbfab-2e6a-4ead-bfb0-5ad5db868f85 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM M isgender M ender: A Community-Informed Approach to Interventions for Misgendering
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a2da99-3a2f-442a-9f47-98af30d9b87c · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d422327e-7c3c-459e-9c4c-774677d84c23 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Toxicity Classification in Ukrainian
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7910204-72e1-4448-a7f3-d09091fa36b5 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the International AAAI Conference on Web and Social Media , volume=
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fae51d7-23e7-4aef-aae7-40ffa4c11406 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the International AAAI Conference on Web and Social Media , volume=
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea2dfb9d-e2ee-4ff1-b6a5-f35c656b1f1a · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop , pages=
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e70d85c-1aa2-4384-84c5-2ffbfd12e9b8 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc7888dc-7bf9-4d8b-8bba-3b06a2bb1d91 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM and Saha, Sriparna and Pasupa, Kitsuchart , title =
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35c4bcef-1500-4a11-a17b-48205080ce62 · outbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM On Calibration of LLM-based Guard Models for Reliable Content Moderation
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56e7e935-1af9-4a4d-aa2f-86ed48c0b5b6 · inbound
BarrierSteer: LLM Safety via Learning Barrier Steering Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b2ca48d-322a-4c1d-9ec2-44a2779397db · inbound
Why Do Large Language Models Generate Harmful Content? Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee3c235e-d736-4b70-b9b0-507b8df37bea · inbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c3d6022-748a-4c66-b3d8-d7f6c01daa7a · inbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6083394-f2d9-440b-aa9c-11737631a9ca · inbound
Do Coding Agents Understand Least-Privilege Authorization? Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62ac8f37-4f52-40ad-b1cc-44d7042493d1 · inbound
Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.