Pith. sign in

Paper Citation Record · LEDGER

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

As of 7 August 2026, this Paper Citation Record lists 100 of 251 outbound references and 6 inbound Pith citation observations for arXiv:2508.05775.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.05775 v3

Coverage vector

measured 100 of 251 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:13:04.217122Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:06:20.077337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T07:05:26.724164Z

Reference resolution

100 of 251 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved90
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 261d18ad-fab4-4f43-9dd3-ecb088b7fbac · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:12:59.700739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:12:59.700739Z digest=sha256:eec39b52b5d40bbefe4230fbf81382025d41830a0325499015ba51135cf84952

Observation df7a80e7-73ef-4c35-b707-6f4833e9d6b6 · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Certifying LLM Safety against Adversarial Prompting

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:12:59.802317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:12:59.802317Z digest=sha256:26786ee0a8f550069fe7d5e666f02b63c2e20abdaaf394e4d4b38b4e684ed304

Observation 47a90b26-9d41-46ec-b784-b65d7cb3fdac · outbound

This paper cites and Arora, Simran and Mazeika, Manias and Hendrycks, Dan and Lin, Zinan and Cheng, Yu and Koyejo, Sanmi and Song, Dawn and Li, Bo , title =.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM and Arora, Simran and Mazeika, Manias and Hendrycks, Dan and Lin, Zinan and Cheng, Yu and Koyejo, Sanmi and Song, Dawn and Li, Bo , title =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:12:59.919407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:12:59.919407Z digest=sha256:6b36c9eaa93daa9e773b5fbb543f14e9921815174fd9247a1cc10c6e9449ae39

Observation 8b639068-14e2-43aa-bcb9-7271a1414879 · outbound

This paper cites TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.005776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.005776Z digest=sha256:17c613f591dff0957890247860a7c0f6242221ca5827591384841b2a90a2f0ab

Observation caaf4c2e-3351-4090-8fc0-c064cb0b8f5f · outbound

This paper cites Discourse & Society , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Discourse & Society , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.104978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.104978Z digest=sha256:25d0779549613295dc439daa6892228c2f1f7d3a04c60401e7cc6a3fcfa16f35

Observation 56cbbbe5-c55e-4f2c-a79d-55d70c7ac078 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Advances in Neural Information Processing Systems , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.181127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.181127Z digest=sha256:23c6ea927198037322f1bf71ebc5bad39679a1ac69abe01c5b7fd58aa3db6e72

Observation 31ebaa58-fc08-479b-8ccd-d8f0970e2863 · outbound

This paper cites Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.321368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.321368Z digest=sha256:59722dd0e3795f7812792d922284ebf74d1670db09b60defa00089e44c8bcc4d

Observation 3024722d-7cfa-4993-86b6-52c8b1367780 · outbound

This paper cites Findings of the association for computational linguistics: EMNLP 2023 , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Findings of the association for computational linguistics: EMNLP 2023 , pages=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.411654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.411654Z digest=sha256:13a60ab363047ff5838ade0f14c6d8a999a36a82a9a9fa26c046a5629b9dbdb0

Observation b62dcf92-388f-4ed0-812f-e6cd008e0d32 · outbound

This paper cites IEEE Transactions on Cognitive and Developmental Systems , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM IEEE Transactions on Cognitive and Developmental Systems , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.524651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.524651Z digest=sha256:1f340039609478fa0885d0905b75a0abe2d34936a117f425d4c9d54fd64fe087

Observation 1852a291-a39b-4b13-9bd1-39bd7f229e87 · outbound

This paper cites Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.699894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.699894Z digest=sha256:f1c7cca5b53699d6b47a55e6b824ef6bf818ae0ed839aed6fa7f79d2f158f3df

Observation 7bce24a2-e0a2-49ac-ba41-f6ce07a77dcd · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.793432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.793432Z digest=sha256:f31d5653a6b6ab881bcc67f177d6468b64acac9eeee8d817bd51439490fe4fe5

Observation 588c10bd-6657-46d9-8169-db93e98a0fa8 · outbound

This paper cites Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.922194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.922194Z digest=sha256:a8beee4a278ccce1b14dd721e0e679babb4503e362acef5e8904b954e84b8d40

Observation 8cdb97cf-27b1-4e68-9b34-49645946c50e · outbound

This paper cites Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.081890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.081890Z digest=sha256:f596f60b308cab30bad4662aeb661e2d2b1f0290b8ed555b5b43e181273cab09

Observation 0f7c2101-7c04-434e-bff4-bf7a40f3b1d4 · outbound

This paper cites Security and Communication Networks , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Security and Communication Networks , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.164060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.164060Z digest=sha256:3ec7ba24e70a959166d899ba5734d08a767bf175d233605becafc1c31699e65f

Observation af481b02-ce75-483f-923f-36f1c96e498c · outbound

This paper cites Proceedings of the First Workshop on Social Influence in Conversations (SICon 2023) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the First Workshop on Social Influence in Conversations (SICon 2023) , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.264593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.264593Z digest=sha256:a5d36e91d6fad28d3404c42708a9155805a58ac98c17202431166e9fba79e520

Observation bc88f121-4999-4919-9bfb-05db88f4f74a · outbound

This paper cites Alignment is not sufficient to prevent large language models from generating harmful information: A psychoanalytic perspective.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Alignment is not sufficient to prevent large language models from generating harmful information: A psychoanalytic perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.410828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.410828Z digest=sha256:05b5d5b7f6a1d5821fb207319477f5b5afb3c13361abcefb012a3a78e55e9c65

Observation 3be65711-c14d-4c93-828e-947d429cc5fc · outbound

This paper cites Systematic Rectification of Language Models via Dead-end Analysis.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Systematic Rectification of Language Models via Dead-end Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.483650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.483650Z digest=sha256:5534b3618260bb645b101fd03cdfd2bff5c5e23da23e241d86edbcc55577d374

Observation f5f12ff1-141f-460e-8f75-c76ff8b69b78 · outbound

This paper cites Successor Features for Efficient Multisubject Controlled Text Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Successor Features for Efficient Multisubject Controlled Text Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.514336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.514336Z digest=sha256:eb0a02a1f31642897c86cb94082c0693eb066164beec6a6ccc918881cfee37ba

Observation 699178a1-2ce7-42cb-8f5a-055746c3e91c · outbound

This paper cites Expert-Guided Extinction of Toxic Tokens for Debiased Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Expert-Guided Extinction of Toxic Tokens for Debiased Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.555889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.555889Z digest=sha256:6d2f64068254b458e57168961bbe2a15d07c8d43bc482f7f0481bf9939951f1f

Observation cfde5f27-fd0a-4007-bba6-ae62fc3df428 · outbound

This paper cites Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.607883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.607883Z digest=sha256:8ec0c8996085e22cf1ba8e3cc9da456ad4abd3229b52b7888d672436b7c26b81

Observation 78d438c3-a207-46c7-ae91-31d8267607d5 · outbound

This paper cites Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.681830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.681830Z digest=sha256:e69fcb2e66df286be1771063df14189b2ca86e21377b9ffb130444e3bed7cad7

Observation 1fdd39d6-cfdd-4101-a7fb-caaac4a64647 · outbound

This paper cites 2024 IEEE Security and Privacy Workshops (SPW) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM 2024 IEEE Security and Privacy Workshops (SPW) , pages=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.769771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.769771Z digest=sha256:9e2a46845a70bf0b2604fb8fd5069eebbc044d7a1ac1a4f620b3e92c52183cc3

Observation 32390486-8e5d-4ff5-916e-cbf37a9521c9 · outbound

This paper cites I’m fully who I am.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM I’m fully who I am

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.803231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.803231Z digest=sha256:168b62a28bcef5920045439940401e7277a7fbb5070e5050fe83f648f74f818b

Observation fed7984b-8f98-4228-b648-4a15ac91245a · outbound

This paper cites From Text to MITRE Techniques: Exploring the Malicious Use of Large Language Models for Generating Cyber Attack Payloads.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM From Text to MITRE Techniques: Exploring the Malicious Use of Large Language Models for Generating Cyber Attack Payloads

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.893879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.893879Z digest=sha256:1ad2215cc502b2b4194f66603c0b53c86bc99990917b28c975c70e66dae39392

Observation 988ccafc-7a5e-4aaa-8ef2-6504b7d2a28e · outbound

This paper cites Scientific Reports , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Scientific Reports , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.986502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.986502Z digest=sha256:431c4d036a2a5168da31279ff589864bb10e5cb9ef69424389d90afb4dddba82

Observation ae41fc50-c3db-4742-887b-f06e3f739fa2 · outbound

This paper cites Otolaryngology--Head and Neck Surgery , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Otolaryngology--Head and Neck Surgery , year=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.013049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.013049Z digest=sha256:7f9cfe5922461ce761e0cfe9307c23d2d798fe84ddeaa04e0997b1f1ba071872

Observation b8bb058e-c906-4d8a-880f-626b48f65a86 · outbound

This paper cites The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.039730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.039730Z digest=sha256:8177c5cfc1da0b9c839cc2d611986bf8c858e2bce75fceb9a3298b90a0590b15

Observation d56f67aa-6057-4433-8d0b-4a4e33158069 · outbound

This paper cites The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.115290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.115290Z digest=sha256:a22fa636149e93b384b0ec1cc4515144152ebe407dc34682279f8b123ce25ca9

Observation 4c0ad005-a921-4062-870c-d34e4c8d7ed4 · outbound

This paper cites Inclusivity in Large Language Models: Personality Traits and Gender Bias in Scientific Abstracts.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Inclusivity in Large Language Models: Personality Traits and Gender Bias in Scientific Abstracts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.198098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.198098Z digest=sha256:ba71ce40b6c73f57c95530f68c51497769e1d83f4955b7f0a1a40d13ffed5d27

Observation fe70d7f0-6775-42a6-bfbb-7b348e29d36b · outbound

This paper cites The African Woman is Rhythmic and Soulful: An Investigation of Implicit Biases in LLM Open-ended Text Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The African Woman is Rhythmic and Soulful: An Investigation of Implicit Biases in LLM Open-ended Text Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.239622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.239622Z digest=sha256:14173b2055d02d374c4920415c107495d2bb43b5de95cc873bd48f925e6037f2

Observation 3af60680-828a-4fa1-978f-01de151e4a68 · outbound

This paper cites TATuP-Zeitschrift f.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TATuP-Zeitschrift f

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.263456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.263456Z digest=sha256:3b8d737165750e6a8e198d8e266bcb56026521505ae3224ca558dbef879d06e1

Observation 1d570c05-8e04-4d25-a096-1c8596171a84 · outbound

This paper cites Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.311367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.311367Z digest=sha256:57f83b03abe8c1ba87ad5155d920f1034aa7e18a7f1399b8a3962d7ec5cfed3b

Observation 66a13d2b-461c-4acf-bf25-6351a60fde69 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Advances in Neural Information Processing Systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.418496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.418496Z digest=sha256:bb0df2705bc4649d5ac6d23d5322c6d69df5abfba5607566406d5373b10e350a

Observation cbef038d-dd7b-4ab2-abae-e276cb072e90 · outbound

This paper cites Attack Prompt Generation for Red Teaming and Defending Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.478257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.478257Z digest=sha256:3e67c8d1f826f369bf8515ca0dc5ee694d58fdf486ac28580fd7924d46a3f618

Observation b58ed619-db02-496a-ac0b-a68b93782ad9 · outbound

This paper cites Exploring the Adversarial Capabilities of Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Exploring the Adversarial Capabilities of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.527686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.527686Z digest=sha256:d06164aa14d42bdecf18d36f63a5a4cb9e6fcaa0351a6adea1e689e15b624947

Observation f4f56dd4-36fb-4e56-8077-3db47afe5ee1 · outbound

This paper cites Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.596616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.596616Z digest=sha256:db655f0971f9884d82e1cf15587725b1ab6462965fda5eec2c500b70a50f291c

Observation 485469a0-bdc5-4ccf-80e5-84675f2a08ea · outbound

This paper cites The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.685636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.685636Z digest=sha256:fa16a729811442e4214f9582d24769eb04d2fddf4bd8aece8ca8e66087669aa5

Observation 042ae570-0dfa-4358-bcfb-f09ecfaf014c · outbound

This paper cites F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.828938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.828938Z digest=sha256:f05e382aa827f0d982fa0ca12a83f7f949b417e235dbaa77a9ff798db238ed6d

Observation 79103127-ed0f-494b-82bf-a8c035d439d7 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.907233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.907233Z digest=sha256:799e82c43c1ba6ed2436c6b946eb429054ed5307a2df2958b342bcae7d50fc6c

Observation ce3e6df4-cdd3-46e3-8ef4-7ed728a81656 · outbound

This paper cites NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.998152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.998152Z digest=sha256:dc7c16e2e9018ab87aed97c546ce33419af8a54fb853f456268a60272bcf0bdf

Observation 40bdc7da-2a16-4037-9803-63721fd03371 · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.063808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.063808Z digest=sha256:f5ad20401f5cbafb68ad1ab07a477d21375c2bb7326d21d36891facb652c07c1

Observation 173e19ea-556e-41c7-b0b1-fbe63074ad43 · outbound

This paper cites Fifty Shades of Bias.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Fifty Shades of Bias

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.146442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.146442Z digest=sha256:f966c7ab6635ec9e4957bb4b15aceb5d3475815de19a00320156de3e9cedea32

Observation 78ea9e60-7d1d-44ab-baff-8594fc267619 · outbound

This paper cites K o C o S a: K orean Context-aware Sarcasm Detection Dataset.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM K o C o S a: K orean Context-aware Sarcasm Detection Dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.198820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.198820Z digest=sha256:ce2f6f4475982461e1d6ca89c90c4c30971c104f6834aacc9d0c35ecf81537f9

Observation 0203ae01-2eae-47bf-a360-bd268aa7305e · outbound

This paper cites ParaFusion: A Large-Scale LLM-Driven English Paraphrase Dataset Infused with High-Quality Lexical and Syntactic Diversity.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM ParaFusion: A Large-Scale LLM-Driven English Paraphrase Dataset Infused with High-Quality Lexical and Syntactic Diversity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.303646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.303646Z digest=sha256:b3b0e94a14883d8890b343f79a643992e2018a887ce3aebe33ff5aa9582a6687

Observation 2eb00353-2d98-4488-94c3-cc9b9f61f1b9 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.368332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.368332Z digest=sha256:2b699a4c42f70d5de685a6879844b495ade6ecc40255a93b8a9738f7ca6ef99e

Observation 7673df10-b083-4358-b07d-637e107f30ab · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.407488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.407488Z digest=sha256:f912f2445770d085ab535a78a72874c5887214ecd9e88f98268e46b59fe8ec70

Observation 14afb384-feac-4544-b26d-75ad2f90ccbd · outbound

This paper cites Text generation for dataset augmentation in security classification tasks.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Text generation for dataset augmentation in security classification tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.471388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.471388Z digest=sha256:54b3ce03fcb79cffae330e56e5a2a18d126bdc1d2b068ae7e562cff75446b64a

Observation 8fa9f4d3-d632-40d9-a50d-d4ed6826d947 · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.550174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.550174Z digest=sha256:90e9a298271e95c1ee43c1208e7733bc18b8208f6fb69ddb5c30776a7ea28abf

Observation 4b0c8281-d9ad-4be3-88b5-307123e3bc52 · outbound

This paper cites International Conference on Machine Learning , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Machine Learning , pages=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.580538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.580538Z digest=sha256:b9e5d30f70e82b0950cf23b78318881c0b88d8f4910ed1d61dc029f1902fe9a3

Observation 8705ad8d-9cd4-466b-b034-18c9e58deb4f · outbound

This paper cites Adversarial Fine-Tuning of Language Models: An Iterative Optimisation Approach for the Generation and Detection of Problematic Content.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Adversarial Fine-Tuning of Language Models: An Iterative Optimisation Approach for the Generation and Detection of Problematic Content

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.657036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.657036Z digest=sha256:48b08cf72ecc575c24a344b271ba9794951f7f69ef201b86b61ccc35b9c278dc

Observation 79165f25-f3b1-4ee8-a196-890c4414871e · outbound

This paper cites TroubleLLM: Align to Red Team Expert.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TroubleLLM: Align to Red Team Expert

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.721737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.721737Z digest=sha256:c888ee0b7cd909e5e4d763bc6e04b36bf2cd5059067a18d466fc6312156c7592

Observation e3620943-fba8-4cdf-8c73-9b0b3674bd14 · outbound

This paper cites Learning diverse attacks on large language models for robust red-teaming and safety tuning.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.820155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.820155Z digest=sha256:6d90622fab415e7ddb2017d1127dfd89aff602ddd8a01a4609139d419e9db0d4

Observation 79871108-0d5c-49d3-bc79-dd7c6aee161c · outbound

This paper cites Outcome-Constrained Large Language Models for Countering Hate Speech.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Outcome-Constrained Large Language Models for Countering Hate Speech

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.900039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.900039Z digest=sha256:0b3e2a33776c9608715ac7a696b4e6511dcd41b3769b4906ec4a5172ddae5dd1

Observation 7f166ca7-d232-427c-9d3c-aa44f147d535 · outbound

This paper cites Proceedings of the CHI Conference on Human Factors in Computing Systems , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the CHI Conference on Human Factors in Computing Systems , pages=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.984396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.984396Z digest=sha256:8b54f64d7dc0a1da3ee37848b93f592e21e2b843185f29a2dd91bfe8f349442e

Observation 8545ce46-1cb4-42ac-8900-8c4ca76e4704 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.024459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.024459Z digest=sha256:4604193caa6bda5c3fc461e85896364cfc989ecbb69e96ef9391f48397013450

Observation 2e9cd479-162b-4609-8a59-504cfcf47563 · outbound

This paper cites Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.028618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.028618Z digest=sha256:e2fcd4f77b7d3ef821ee9d3b2db3cfb1e5130ebd6c900295e4eb343f1711607c

Observation 9765cd3b-ca29-4b0b-8c55-9a5c99e9e018 · outbound

This paper cites Seeing Seeds Beyond Weeds: Green Teaming Generative AI for Beneficial Uses.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Seeing Seeds Beyond Weeds: Green Teaming Generative AI for Beneficial Uses

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.032739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.032739Z digest=sha256:31969f1c037abdb36aad51b4d9c7be9e693e2887afb39ac5d04fe80e6488a692

Observation ffb8d0e4-52c2-4072-94a7-aaa4be383cf5 · outbound

This paper cites Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.037075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.037075Z digest=sha256:4bd754e50f1930f4e411e636bd22f5cecd28c33b62c57d1a7e2311eb48436512

Observation a3aa813f-3e0d-4da8-a522-c6652f44ff70 · outbound

This paper cites Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.041425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.041425Z digest=sha256:82d72180c7e49dcd2db99300fa6187b8a1dadcb0c7698674ba29ae1be8e34b56

Observation eb41db75-9280-4d7b-b079-445d4b108a1e · outbound

This paper cites Creativity Has Left the Chat: The Price of Debiasing Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Creativity Has Left the Chat: The Price of Debiasing Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.046627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.046627Z digest=sha256:8dd9d816b78e19c00ea0870195424c414380c19dfc8c7b1414bd4afc5f6af56b

Observation 4fb36029-f869-4632-84e2-4798f5ba304c · outbound

This paper cites Controlled Text Generation for Large Language Model with Dynamic Attribute Graphs.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Controlled Text Generation for Large Language Model with Dynamic Attribute Graphs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.050958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.050958Z digest=sha256:9c3a694450a7320144724751a1493462a28e780fecdb9e0a9e41001997435367

Observation 7e68a8cf-756a-4b8e-8909-781950102777 · outbound

This paper cites Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.055130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.055130Z digest=sha256:cd91c2ac9925ec6f0c2f206db79976fadfaaa0ee643dfc848aa4e75105d5244f

Observation fadb5382-8c1b-4e3f-88e9-d283ed392b10 · outbound

This paper cites Yale JL & Tech.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Yale JL & Tech

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.059034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.059034Z digest=sha256:3e31ad75edaa39977c9c90b938a1846d536cbbc235b328ff47b5f15deaf9cd02

Observation 7f71c83e-dfd6-4dcf-b82e-b45d171c540b · outbound

This paper cites Engineering, Technology & Applied Science Research , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Engineering, Technology & Applied Science Research , volume=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.062924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.062924Z digest=sha256:0a12ebc84bb79725a23314ce7abab1500461ce184603ec3ef945812947f50f97

Observation 49673c17-e9d5-4249-9414-3d001d1ff6f3 · outbound

This paper cites Machine-Generated Tweets , author=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Machine-Generated Tweets , author=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.066922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.066922Z digest=sha256:7d6e220b9687f95519d23d0bf99ee4ed711f2e11e217a24848f0c4f58b109757

Observation a8c6e79f-2a1f-468c-8481-f059a24e2c7b · outbound

This paper cites Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Natural Language Processing , pages=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.070937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.070937Z digest=sha256:067786e598e196ac7844e8f6fe0f2048c986789efd2c9f832407784609a3e8cc

Observation fa620c31-7410-4c82-ad9c-610d6e605801 · outbound

This paper cites International Conference on Computational Science and Its Applications , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Computational Science and Its Applications , pages=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.075456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.075456Z digest=sha256:34c0e4cc99a609e8d08c4ec654587a1f92c0d9c5f83cd34481acb622522862fe

Observation ec361d9f-ec3f-4332-88fd-83b657ee20e4 · outbound

This paper cites Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.079715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.079715Z digest=sha256:6fc614f60c81a1cbceed7f1884df33c30fbb8e8f6052f2a5fa75a557c22f6439

Observation 4613330a-9ecb-48f5-bb4f-e743afd3f148 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.084319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.084319Z digest=sha256:1c6e2c5b1330e77596f7fbfd5eacbd0d4f2007caa6218396a524644b8f834f63

Observation baf565a0-4225-4baa-9631-8a37500fc0ce · outbound

This paper cites Regulating Hate Speech Created by Generative AI , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Regulating Hate Speech Created by Generative AI , pages=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.088949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.088949Z digest=sha256:161184211fd5007f1a934fe6addff508db966b04f6fef356085ccb91af645348

Observation 005b5e48-1803-4b5c-a922-4f6a68568577 · outbound

This paper cites A Study of Slang Representation Methods.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM A Study of Slang Representation Methods

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:08.069771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.092860Z digest=sha256:c89690262e9a22df9528a2f762dd95428abd968c0e28d238ce88d7b6c1f99225

Observation 6f761ee3-f867-44ab-ad42-3032a44c39ae · outbound

This paper cites Perplexed by Quality: A Perplexity-based Method for Adult and Harmful Content Detection in Multilingual Heterogeneous Web Data.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Perplexed by Quality: A Perplexity-based Method for Adult and Harmful Content Detection in Multilingual Heterogeneous Web Data

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:08.047629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.096997Z digest=sha256:12a5ee42971f9c51587e6f765d652baa9d4b923769a2cf1a0ee4bcbe2d370022

Observation 52f71db4-b053-4998-a6a1-ac5e0bbd3db4 · outbound

This paper cites Human-Guided Fair Classification for Natural Language Processing.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Human-Guided Fair Classification for Natural Language Processing

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.101234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.101234Z digest=sha256:1ee9cff42418177813d986c0bf6003eeec8f0a4b0532a4b91466b631485772a1

Observation 65b692fc-8d19-4aab-9192-b7358227d056 · outbound

This paper cites arXiv preprint arXiv:2301.12534 , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM arXiv preprint arXiv:2301.12534 , year=

Reference 74

Resolution
verified exact
raw_fallback, observed 2026-08-05T23:13:08.009527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.105543Z digest=sha256:bcfe290a546f5ece372c3cb29da385a16914cb1cad0486925d3ec7fcd2afa15f

Observation 3042f772-ca5c-4b91-8fcb-c770d6a45a37 · outbound

This paper cites International Conference on Advances in Social Networks Analysis and Mining , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Advances in Social Networks Analysis and Mining , pages=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.109738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.109738Z digest=sha256:1f818e99cb9081817f90e6ea7e8b45a82b1145e8371467ce1258a3b550177a3f

Observation 5f7ed78f-2433-4640-90f3-a95f2603ed0a · outbound

This paper cites Explicit Toxicity Detection Models with Interactive Visualization , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Explicit Toxicity Detection Models with Interactive Visualization , year=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.114159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.114159Z digest=sha256:7ed7fe775d406398f890da0c891d8048e2898a50fb9bd9c9ad8c3432950d2d19

Observation 164a9473-9137-4419-a9a1-adad34d60acc · outbound

This paper cites Which Argumentative Aspects of Hate Speech in Social Media can be reliably identified?.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Which Argumentative Aspects of Hate Speech in Social Media can be reliably identified?

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.905145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.118824Z digest=sha256:6211ab518c068613d6771435f8a2d15d7e46934cf1ccda4e4f0e92d44eb3784e

Observation 25099d55-a1e2-4d1b-a80b-f0077fbc4f34 · outbound

This paper cites Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI) , pages=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.123304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.123304Z digest=sha256:6941e2d89b11dbff3765033fcd22d22e5ff4c0ea5e4aab9f1b8977e5e7e05878

Observation a458cb7d-4282-49cc-bc73-972d9425e3d6 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.127327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.127327Z digest=sha256:42c523949b1bc954a797c1b1feca61638a76f89da5b437c86d89906ed22583c7

Observation 803845f0-f56b-479d-9657-43c8d43a3d27 · outbound

This paper cites Topological Data Mapping of Online Hate Speech, Misinformation, and General Mental Health: A Large Language Model Based Study.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Topological Data Mapping of Online Hate Speech, Misinformation, and General Mental Health: A Large Language Model Based Study

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:13:07.883128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.131563Z digest=sha256:f871fb9de76a128e7ab476d18461121a048cc698c2143cac903052152cfbf2fd

Observation 53d99bd2-5e9f-4264-a72f-8f076dcc6532 · outbound

This paper cites Demonstrations Are All You Need: Advancing Offensive Content Paraphrasing using In-Context Learning.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Demonstrations Are All You Need: Advancing Offensive Content Paraphrasing using In-Context Learning

Reference 81

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.861922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.135809Z digest=sha256:0a77ffd06f5509b2eec488e6891b63390eb59e79544e3000bc19b2307143d911

Observation d6661e22-6d2c-4b57-b6d2-934321f5527a · outbound

This paper cites Beyond plain toxic: building datasets for detection of flammable topics and inappropriate statements.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Beyond plain toxic: building datasets for detection of flammable topics and inappropriate statements

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.139949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.139949Z digest=sha256:1c6431bf097d3c94a6d8a90711458f5f4ed474ad43d8709246b4f8a6f2660bb1

Observation d9798d37-cc31-433b-a0b5-a14eb29082c5 · outbound

This paper cites Evaluation of ChatGPT and BERT-based models for Turkish hate speech detection.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Evaluation of ChatGPT and BERT-based models for Turkish hate speech detection

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.144250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.144250Z digest=sha256:9a96330361efff68878eef8f78fa0ee8d2dcf39d4dfc67caf7f45cff90fb77ab

Observation 5ec65427-89d8-4191-b8ae-ac3903d98754 · outbound

This paper cites HateRephrase: Zero- and Few-Shot Reduction of Hate Intensity in Online Posts using Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM HateRephrase: Zero- and Few-Shot Reduction of Hate Intensity in Online Posts using Large Language Models

Reference 84

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.840022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.148362Z digest=sha256:1dade84637eb0a5a2751fc26b3b1183b88d05542e87760fea6566f1265bd4f5d

Observation 10435c6c-a17d-4409-9d63-6408f817ac91 · outbound

This paper cites FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.152649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.152649Z digest=sha256:9567c2050c3859217ec9e6b3e3dc9dff3d343ed7ee353e238949657ada103347

Observation 302948ef-6193-4c33-911b-47b1b120795a · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.157130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.157130Z digest=sha256:5b3d5af1e26346b77224dcab6e52907fe7e238e55077405b50fbcb45afa9c8f5

Observation a5365e4f-0f2d-47a2-8d76-5b266eafd26e · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.161107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.161107Z digest=sha256:e725c0265709c755d4e1325016ed2cce5aaf4daa0ce24f80e35d51bcb8f8473d

Observation 78766c66-76e8-4a0a-9996-8cb3380fc3b7 · outbound

This paper cites 2024 , isbn =.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM 2024 , isbn =

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.165543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.165543Z digest=sha256:277c7a9cd3be81dfdc15f185bac8f32ecfc966289a1e8d6a2b62299b151b6e68

Observation dd78b58d-881b-4fd8-9d82-aad9612a5cf3 · outbound

This paper cites Eagle: Ethical Dataset Given from Real Interactions.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Eagle: Ethical Dataset Given from Real Interactions

Reference 89

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.712366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.169918Z digest=sha256:1027b2bf2d9a066d849426ce7dc3870aaafe7bd72cb8a750a97f480321737387

Observation 54e798c2-c413-47ea-bb80-6a08a6c8fc0c · outbound

This paper cites Proceedings of the first workshop on language technology for equality, diversity and inclusion , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the first workshop on language technology for equality, diversity and inclusion , pages=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.174303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.174303Z digest=sha256:626002abc05f10fd5ecbe4e0b2cdd390190cb657044e909ea2ef20dbdf0dfc00

Observation 54f73beb-ca2e-4f98-b9b6-d76066cfccfe · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.178745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.178745Z digest=sha256:c5b1eed0231ea76044f9abf3eef9811a3eee83ee43c18af1aa8a861774ef41f3

Observation 34bdbfab-2e6a-4ead-bfb0-5ad5db868f85 · outbound

This paper cites M isgender M ender: A Community-Informed Approach to Interventions for Misgendering.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM M isgender M ender: A Community-Informed Approach to Interventions for Misgendering

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.183378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.183378Z digest=sha256:04d155a81677733cfe21e8da7e226c80b6ab54220d0ac45270a533e6e5827919

Observation 41a2da99-3a2f-442a-9f47-98af30d9b87c · outbound

This paper cites HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.187341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.187341Z digest=sha256:991e3cdfe23dd889db2917bddb332ed9cf3f99814849a88a95b32654bc2174ff

Observation d422327e-7c3c-459e-9c4c-774677d84c23 · outbound

This paper cites Toxicity Classification in Ukrainian.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Toxicity Classification in Ukrainian

Reference 94

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.675095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.191642Z digest=sha256:ccd95fd0345ae456c6cea85605dd039911bb8a1f56f9055eb9c3590ca997c432

Observation b7910204-72e1-4448-a7f3-d09091fa36b5 · outbound

This paper cites Proceedings of the International AAAI Conference on Web and Social Media , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the International AAAI Conference on Web and Social Media , volume=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.196099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.196099Z digest=sha256:060b71b36a39f5f4763f80d5402783b69f5573b877deb9cd295fad83a00e8d96

Observation 2fae51d7-23e7-4aef-aae7-40ffa4c11406 · outbound

This paper cites Proceedings of the International AAAI Conference on Web and Social Media , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the International AAAI Conference on Web and Social Media , volume=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.200044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.200044Z digest=sha256:c662c768d74be54d6ae0e27025463d7e72d24dce939c6d73f02a4323a5e95b5f

Observation ea2dfb9d-e2ee-4ff1-b6a5-f35c656b1f1a · outbound

This paper cites Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop , pages=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.204022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.204022Z digest=sha256:426150709e4ddbe13010e3eac5829676aaaf8ae9dea1f4d03fc0d8be2d82d67c

Observation 0e70d85c-1aa2-4384-84c5-2ffbfd12e9b8 · outbound

This paper cites A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech

Reference 98

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.653077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.208197Z digest=sha256:576ece17462f8b0b8cdc4759d866187ef2b706ba86e752725f13d55e5c3fe9af

Observation fc7888dc-7bf9-4d8b-8bba-3b06a2bb1d91 · outbound

This paper cites and Saha, Sriparna and Pasupa, Kitsuchart , title =.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM and Saha, Sriparna and Pasupa, Kitsuchart , title =

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.213203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.213203Z digest=sha256:dc79f31f6c39384c3391c80dcd47b1f71e82d22b05ced384816c0876d062ce05

Observation 35c4bcef-1500-4a11-a17b-48205080ce62 · outbound

This paper cites On Calibration of LLM-based Guard Models for Reliable Content Moderation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM On Calibration of LLM-based Guard Models for Reliable Content Moderation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.217122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.217122Z digest=sha256:83f27f1a6c1507b17340e6f6c0ff1524229c2d793f5514e3d6bd8d8b23bff7a0

Pith citing papers

Observation 56e7e935-1af9-4a4d-aa2f-86ed48c0b5b6 · inbound

BarrierSteer: LLM Safety via Learning Barrier Steering cites this paper.

BarrierSteer: LLM Safety via Learning Barrier Steering Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T07:02:03.058731Z digest=sha256:acbe9ee774c09a7dbd02619ab20c4ec97f6ba145976e4884b74bd6c79b8a9dbb

Observation 9b2ca48d-322a-4c1d-9ec2-44a2779397db · inbound

Why Do Large Language Models Generate Harmful Content? cites this paper.

Why Do Large Language Models Generate Harmful Content? Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:31:13.545599Z digest=sha256:0e49c1c956dcdaa8de7931ce36031fca00750f7f867a4562960caca6314838cc

Observation ee3c235e-d736-4b70-b9b0-507b8df37bea · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:17:19.079380Z digest=sha256:d2c1fb521561dca603effececb585b9c04a077f25c01e49d3967888b63b42704

Observation 9c3d6022-748a-4c66-b3d8-d7f6c01daa7a · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:873b73a7b6a612cd32ac496952b505defde2a1b108bd095088301026556a8ba3

Observation a6083394-f2d9-440b-aa9c-11737631a9ca · inbound

Do Coding Agents Understand Least-Privilege Authorization? cites this paper.

Do Coding Agents Understand Least-Privilege Authorization? Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:34:14.379419Z digest=sha256:dceb0ebfe16935f787f6e77f7b55a5e67a1e6094ad753bd138b23e06cd8b52b2

Observation 62ac8f37-4f52-40ad-b1cc-44d7042493d1 · inbound

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows cites this paper.

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T07:06:20.077337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:06:20.077337Z digest=sha256:faa64ca9f73fb9c31ec6277e340ef4506522cd4bae58a1988a64648edaf19921