Pith. sign in

Paper Citation Record · LEDGER

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

As of 14 August 2026, this Paper Citation Record lists 100 of 251 outbound references and 6 inbound Pith citation observations for arXiv:2508.05775.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.05775 v3

Coverage vector

measured 100 of 251 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:13:04.217122Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:06:20.077337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T07:05:26.724164Z

Reference resolution

100 of 251 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved90
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 261d18ad-fab4-4f43-9dd3-ecb088b7fbac · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:12:59.700739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:12:59.700739Z digest=sha256:20effafba0d730de3e86532fab4c3d983a795a0efd19a57c6d3abc1a82dca711

Observation df7a80e7-73ef-4c35-b707-6f4833e9d6b6 · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Certifying LLM Safety against Adversarial Prompting

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:12:59.802317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:12:59.802317Z digest=sha256:20f5e96b7491fc469c3796af8265bdc26be874fbe50b64c1316c293f8dfd93c4

Observation 47a90b26-9d41-46ec-b784-b65d7cb3fdac · outbound

This paper cites and Arora, Simran and Mazeika, Manias and Hendrycks, Dan and Lin, Zinan and Cheng, Yu and Koyejo, Sanmi and Song, Dawn and Li, Bo , title =.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM and Arora, Simran and Mazeika, Manias and Hendrycks, Dan and Lin, Zinan and Cheng, Yu and Koyejo, Sanmi and Song, Dawn and Li, Bo , title =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:12:59.919407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:12:59.919407Z digest=sha256:37e2bd3ab2b1cd2884c2f7fa74ed101d538f1c6467a2f8c9fb0341ba1e5ef5fd

Observation 8b639068-14e2-43aa-bcb9-7271a1414879 · outbound

This paper cites TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.005776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.005776Z digest=sha256:178d399b08567767f07accfd22d9898856a44f3b4b977dbec2c6757ab10b171d

Observation caaf4c2e-3351-4090-8fc0-c064cb0b8f5f · outbound

This paper cites Discourse & Society , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Discourse & Society , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.104978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.104978Z digest=sha256:dc54f055f78946189117fd05ca019f8fdba4f5b0a417bc5cd37919b9f55f70a5

Observation 56cbbbe5-c55e-4f2c-a79d-55d70c7ac078 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Advances in Neural Information Processing Systems , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.181127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.181127Z digest=sha256:1dc6bdb6511caa5c418498411326e1f832bd8be0784664cabe7f9d549c909859

Observation 31ebaa58-fc08-479b-8ccd-d8f0970e2863 · outbound

This paper cites Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.321368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.321368Z digest=sha256:d9af4c15959bf57f52aea24731906fc4760cbd61c322f185278bddf9c8981efc

Observation 3024722d-7cfa-4993-86b6-52c8b1367780 · outbound

This paper cites Findings of the association for computational linguistics: EMNLP 2023 , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Findings of the association for computational linguistics: EMNLP 2023 , pages=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.411654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.411654Z digest=sha256:8f2adc2ee46c0c0c5096d9dd86f8c93cfe273ada32293b7e7e3a2ca69da3a689

Observation b62dcf92-388f-4ed0-812f-e6cd008e0d32 · outbound

This paper cites IEEE Transactions on Cognitive and Developmental Systems , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM IEEE Transactions on Cognitive and Developmental Systems , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.524651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.524651Z digest=sha256:e8c7214901e4c96e651689d7b90aef8fb6b62f84cf7fac2a9a9659f8345d03e1

Observation 1852a291-a39b-4b13-9bd1-39bd7f229e87 · outbound

This paper cites Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.699894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.699894Z digest=sha256:c81557e0b8ff166a55dad7e3dad2221a75bdfdf291b9b80c357656022285e451

Observation 7bce24a2-e0a2-49ac-ba41-f6ce07a77dcd · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.793432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.793432Z digest=sha256:ef7c330c2c8fc968be14dd6fdf659532b045273d3ca1ee83f3ab1b5e2214e60f

Observation 588c10bd-6657-46d9-8169-db93e98a0fa8 · outbound

This paper cites Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:00.922194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:00.922194Z digest=sha256:eeec9b6cb56ddf2814ea8e1dc87bb22b8d5bb03625ce769032a873c79a7dd62b

Observation 8cdb97cf-27b1-4e68-9b34-49645946c50e · outbound

This paper cites Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.081890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.081890Z digest=sha256:5992cece10475d478b095c92c5460e2a8a5c3a61bcff969b3a04901638dbe5ac

Observation 0f7c2101-7c04-434e-bff4-bf7a40f3b1d4 · outbound

This paper cites Security and Communication Networks , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Security and Communication Networks , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.164060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.164060Z digest=sha256:78e3a920a9db3605baaa3bf029377d8729d303d5d1eba75998ba0d2c9c8a881e

Observation af481b02-ce75-483f-923f-36f1c96e498c · outbound

This paper cites Proceedings of the First Workshop on Social Influence in Conversations (SICon 2023) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the First Workshop on Social Influence in Conversations (SICon 2023) , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.264593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.264593Z digest=sha256:f6a0b4a559853e0f43b0fa83fe61cb8e4fb86d462dd4eff086b5d5648f9f1013

Observation bc88f121-4999-4919-9bfb-05db88f4f74a · outbound

This paper cites Alignment is not sufficient to prevent large language models from generating harmful information: A psychoanalytic perspective.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Alignment is not sufficient to prevent large language models from generating harmful information: A psychoanalytic perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.410828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.410828Z digest=sha256:bb49ea15d799f6a1d2a721f96e97645cde4470dcb29510fb3040b50e7e94cc3f

Observation 3be65711-c14d-4c93-828e-947d429cc5fc · outbound

This paper cites Systematic Rectification of Language Models via Dead-end Analysis.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Systematic Rectification of Language Models via Dead-end Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.483650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.483650Z digest=sha256:7e22969747873030261953ee9479c2cddd7baa725b119aa20ffc80e7aeea6da1

Observation f5f12ff1-141f-460e-8f75-c76ff8b69b78 · outbound

This paper cites Successor Features for Efficient Multisubject Controlled Text Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Successor Features for Efficient Multisubject Controlled Text Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.514336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.514336Z digest=sha256:6b132540a13a6ec8826a1d1efd66d02a71375dfd943b97c504c08bac29f333d6

Observation 699178a1-2ce7-42cb-8f5a-055746c3e91c · outbound

This paper cites Expert-Guided Extinction of Toxic Tokens for Debiased Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Expert-Guided Extinction of Toxic Tokens for Debiased Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.555889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.555889Z digest=sha256:5b73ab9b9bf37ed6474935dc6f46ef52e221694aeb57ce6f190365fae136434b

Observation cfde5f27-fd0a-4007-bba6-ae62fc3df428 · outbound

This paper cites Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.607883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.607883Z digest=sha256:7c2f3b7e44dc3add60b4b945a4494d5da60318277aa07886d378b7b70e16f92d

Observation 78d438c3-a207-46c7-ae91-31d8267607d5 · outbound

This paper cites Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.681830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.681830Z digest=sha256:0b90fc5c68b9d3c1f2e0335d47eff273a1b2b1979809ef9e60e3fc4b858dec09

Observation 1fdd39d6-cfdd-4101-a7fb-caaac4a64647 · outbound

This paper cites 2024 IEEE Security and Privacy Workshops (SPW) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM 2024 IEEE Security and Privacy Workshops (SPW) , pages=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.769771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.769771Z digest=sha256:d53cdbd28b146cf5a431397cf81a4b25293d054ba4f555bcd1c8be74cf298120

Observation 32390486-8e5d-4ff5-916e-cbf37a9521c9 · outbound

This paper cites I’m fully who I am.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM I’m fully who I am

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.803231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.803231Z digest=sha256:412f6fca8745c8d18ba8738eb49f51c50799013a70ec579ef43c6663d89552b6

Observation fed7984b-8f98-4228-b648-4a15ac91245a · outbound

This paper cites From Text to MITRE Techniques: Exploring the Malicious Use of Large Language Models for Generating Cyber Attack Payloads.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM From Text to MITRE Techniques: Exploring the Malicious Use of Large Language Models for Generating Cyber Attack Payloads

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.893879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.893879Z digest=sha256:8434f66416897c85999a33234e7bb0699f74c1c21bc67cc04077d4e69a07d347

Observation 988ccafc-7a5e-4aaa-8ef2-6504b7d2a28e · outbound

This paper cites Scientific Reports , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Scientific Reports , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:01.986502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:01.986502Z digest=sha256:3d70706206fb6950742b53a340f1e9dee5ec95958fb5efd76ea0a975fa6af97f

Observation ae41fc50-c3db-4742-887b-f06e3f739fa2 · outbound

This paper cites Otolaryngology--Head and Neck Surgery , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Otolaryngology--Head and Neck Surgery , year=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.013049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.013049Z digest=sha256:977089cb4a04ef21dab9fef618e005f8f5d9419f117b7f7c73d78bbe8dfa51bc

Observation b8bb058e-c906-4d8a-880f-626b48f65a86 · outbound

This paper cites The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.039730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.039730Z digest=sha256:14fe219965546eba89f25ba9c1e555a9e0a6de2ce2a8c088fa3460eec9b13f0d

Observation d56f67aa-6057-4433-8d0b-4a4e33158069 · outbound

This paper cites The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.115290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.115290Z digest=sha256:1c880231c597286b553b79cb344c704922df8daf55b3ff6660a6d397b861bec5

Observation 4c0ad005-a921-4062-870c-d34e4c8d7ed4 · outbound

This paper cites Inclusivity in Large Language Models: Personality Traits and Gender Bias in Scientific Abstracts.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Inclusivity in Large Language Models: Personality Traits and Gender Bias in Scientific Abstracts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.198098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.198098Z digest=sha256:e8c4f3950880063ea5e66e779f43d9a05d00936de5ce64ec44e3302f93319a69

Observation fe70d7f0-6775-42a6-bfbb-7b348e29d36b · outbound

This paper cites The African Woman is Rhythmic and Soulful: An Investigation of Implicit Biases in LLM Open-ended Text Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The African Woman is Rhythmic and Soulful: An Investigation of Implicit Biases in LLM Open-ended Text Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.239622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.239622Z digest=sha256:6e8e66b7d22d71dbfe781513f157502b419441b7aeefe20e0b16a644895b322d

Observation 3af60680-828a-4fa1-978f-01de151e4a68 · outbound

This paper cites TATuP-Zeitschrift f.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TATuP-Zeitschrift f

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.263456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.263456Z digest=sha256:ac3f03f6927a2850a0d18809d5cc084084476daf498c3ccca62fd1b2d4a55b75

Observation 1d570c05-8e04-4d25-a096-1c8596171a84 · outbound

This paper cites Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.311367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.311367Z digest=sha256:9b482e25a6c2a42f5d1de81c5cfc4f391604f9286ed08217d78b30a3cef86f57

Observation 66a13d2b-461c-4acf-bf25-6351a60fde69 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Advances in Neural Information Processing Systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.418496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.418496Z digest=sha256:997c00062c55fab01c3b05688f7b6dc9084bcf368cdcc083dfc072fdaecad90d

Observation cbef038d-dd7b-4ab2-abae-e276cb072e90 · outbound

This paper cites Attack Prompt Generation for Red Teaming and Defending Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.478257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.478257Z digest=sha256:aa1280732b80d1bce5dd2b4dcd98d0256c1a1fd3640a6b79ac796604d1fcf373

Observation b58ed619-db02-496a-ac0b-a68b93782ad9 · outbound

This paper cites Exploring the Adversarial Capabilities of Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Exploring the Adversarial Capabilities of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.527686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.527686Z digest=sha256:cfb07a6db998953d7ed40bbd5f4409a0824eecb188c6d11de37ff0715c13e90e

Observation f4f56dd4-36fb-4e56-8077-3db47afe5ee1 · outbound

This paper cites Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.596616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.596616Z digest=sha256:8d2cc9991d2cb483515388b31bc935c74199c56fe618d441473557d143592d10

Observation 485469a0-bdc5-4ccf-80e5-84675f2a08ea · outbound

This paper cites The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.685636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.685636Z digest=sha256:61437d93ff3b81ea910c05763db0adea327916f45c888f16332a59eb7d50fd8e

Observation 042ae570-0dfa-4358-bcfb-f09ecfaf014c · outbound

This paper cites F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.828938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.828938Z digest=sha256:a7f3dee11935bd9efc2e72cbb5fa9db070255d9788116a7f096bb66ac3b9ae81

Observation 79103127-ed0f-494b-82bf-a8c035d439d7 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.907233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.907233Z digest=sha256:25d70d78d27e9996da1bb189496ce345c04a1c29f37796309021052bca3c9339

Observation ce3e6df4-cdd3-46e3-8ef4-7ed728a81656 · outbound

This paper cites NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.998152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.998152Z digest=sha256:bfef58ed92144677ab2382d029eea964942c91d9bdc556a835d17df7423c935c

Observation 40bdc7da-2a16-4037-9803-63721fd03371 · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.063808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.063808Z digest=sha256:d0bfbb1a294ca5eeb8e5380da27e6df960ef82d47faf133b86d17e4bdb7b815e

Observation 173e19ea-556e-41c7-b0b1-fbe63074ad43 · outbound

This paper cites Fifty Shades of Bias.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Fifty Shades of Bias

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.146442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.146442Z digest=sha256:ee29e4a641504e8bfff9a1a4ade1d95a055742f28306343edca2a2c6ddd2f848

Observation 78ea9e60-7d1d-44ab-baff-8594fc267619 · outbound

This paper cites K o C o S a: K orean Context-aware Sarcasm Detection Dataset.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM K o C o S a: K orean Context-aware Sarcasm Detection Dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.198820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.198820Z digest=sha256:80a94670856924bb56a5fda34aa342adb54cc5c13e479f1b62467386e33f3df0

Observation 0203ae01-2eae-47bf-a360-bd268aa7305e · outbound

This paper cites ParaFusion: A Large-Scale LLM-Driven English Paraphrase Dataset Infused with High-Quality Lexical and Syntactic Diversity.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM ParaFusion: A Large-Scale LLM-Driven English Paraphrase Dataset Infused with High-Quality Lexical and Syntactic Diversity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.303646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.303646Z digest=sha256:96a98814368728106b18091042bf6267321564ef3ce30f178b682aff2b651bf4

Observation 2eb00353-2d98-4488-94c3-cc9b9f61f1b9 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.368332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.368332Z digest=sha256:aff1f8aa16ab8656f56246368593d74853b3b34b5aadc165250565bd14d6642d

Observation 7673df10-b083-4358-b07d-637e107f30ab · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.407488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.407488Z digest=sha256:79b289fa255eb5f44bb156812e5bd2670def546f93e1c349d8bc1e9b8acead28

Observation 14afb384-feac-4544-b26d-75ad2f90ccbd · outbound

This paper cites Text generation for dataset augmentation in security classification tasks.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Text generation for dataset augmentation in security classification tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.471388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.471388Z digest=sha256:863cf887d4d9f38c9004d30c17c338090ab644725485ed953ea3b12b7b7b99c3

Observation 8fa9f4d3-d632-40d9-a50d-d4ed6826d947 · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.550174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.550174Z digest=sha256:ec61fe362c8cc3f95b4e5b38a2337b069c8cf46c41796cbc5a08b1b3e48182a1

Observation 4b0c8281-d9ad-4be3-88b5-307123e3bc52 · outbound

This paper cites International Conference on Machine Learning , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Machine Learning , pages=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.580538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.580538Z digest=sha256:f6eec57f51d1a8688a6a040986aa9eb931ac6e7e2d9dbee66892f5dfd8446ccd

Observation 8705ad8d-9cd4-466b-b034-18c9e58deb4f · outbound

This paper cites Adversarial Fine-Tuning of Language Models: An Iterative Optimisation Approach for the Generation and Detection of Problematic Content.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Adversarial Fine-Tuning of Language Models: An Iterative Optimisation Approach for the Generation and Detection of Problematic Content

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.657036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.657036Z digest=sha256:ff3d86ae271b43f5447bc2bb3db144a1d7e532c82eed34a04e1b9cac00124997

Observation 79165f25-f3b1-4ee8-a196-890c4414871e · outbound

This paper cites TroubleLLM: Align to Red Team Expert.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM TroubleLLM: Align to Red Team Expert

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.721737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.721737Z digest=sha256:ce0735234d569799b0c1163115759d2b17fe3b43f5e3b00f7ae60b4bba80662b

Observation e3620943-fba8-4cdf-8c73-9b0b3674bd14 · outbound

This paper cites Learning diverse attacks on large language models for robust red-teaming and safety tuning.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.820155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.820155Z digest=sha256:1f008228831a1c4907b20c7d3964db2b39a97e146ea819d12abb93b9dac5d0b9

Observation 79871108-0d5c-49d3-bc79-dd7c6aee161c · outbound

This paper cites Outcome-Constrained Large Language Models for Countering Hate Speech.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Outcome-Constrained Large Language Models for Countering Hate Speech

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.900039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.900039Z digest=sha256:d69d6a49b8845f005acc7845098ecc02bf3e8601d6834fbd1fa5c7bf6c28c706

Observation 7f166ca7-d232-427c-9d3c-aa44f147d535 · outbound

This paper cites Proceedings of the CHI Conference on Human Factors in Computing Systems , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the CHI Conference on Human Factors in Computing Systems , pages=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.984396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.984396Z digest=sha256:76983cbfd980d5011172559cb26942b4542cf8da463e8fff1d543fd31eb726d6

Observation 8545ce46-1cb4-42ac-8900-8c4ca76e4704 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.024459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.024459Z digest=sha256:1bb5fea7e7f8a8476bcaab938e709acaadbd417d2a1695fb5942d9b7e7163114

Observation 2e9cd479-162b-4609-8a59-504cfcf47563 · outbound

This paper cites Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.028618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.028618Z digest=sha256:dab442ce57982b3474afa3d04cd1fac966c256384dd12cd8fb19d368e6b26244

Observation 9765cd3b-ca29-4b0b-8c55-9a5c99e9e018 · outbound

This paper cites Seeing Seeds Beyond Weeds: Green Teaming Generative AI for Beneficial Uses.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Seeing Seeds Beyond Weeds: Green Teaming Generative AI for Beneficial Uses

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.032739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.032739Z digest=sha256:d1f554118c8fe772defde9489e8550072c45fa95da34dbfced689a37766f461e

Observation ffb8d0e4-52c2-4072-94a7-aaa4be383cf5 · outbound

This paper cites Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.037075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.037075Z digest=sha256:454f44c16930bab6cbb2a3fcdf54d4d9ba3863c6494c6ec5547b42330a9080e0

Observation a3aa813f-3e0d-4da8-a522-c6652f44ff70 · outbound

This paper cites Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.041425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.041425Z digest=sha256:b09882f37b29c1da85c11327502cdb37f66a1f0f538b38cada9a69b3ecae0ec6

Observation eb41db75-9280-4d7b-b079-445d4b108a1e · outbound

This paper cites Creativity Has Left the Chat: The Price of Debiasing Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Creativity Has Left the Chat: The Price of Debiasing Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.046627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.046627Z digest=sha256:c8c815cb3247cdcb6c68b497b7b0e3cef4728e7a1f9bf11dc37466679bbc59a4

Observation 4fb36029-f869-4632-84e2-4798f5ba304c · outbound

This paper cites Controlled Text Generation for Large Language Model with Dynamic Attribute Graphs.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Controlled Text Generation for Large Language Model with Dynamic Attribute Graphs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.050958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.050958Z digest=sha256:4af1058c91ad165a2c51a8d323401fdc8a3c0360e1ba380ff2ab1ea1221a4864

Observation 7e68a8cf-756a-4b8e-8909-781950102777 · outbound

This paper cites Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.055130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.055130Z digest=sha256:45dc6665762e07caaf1a7c2214c48806961ada2681886ad89469f5e783f13f24

Observation fadb5382-8c1b-4e3f-88e9-d283ed392b10 · outbound

This paper cites Yale JL & Tech.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Yale JL & Tech

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.059034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.059034Z digest=sha256:30cda782fa72864c1f40a57f54b181abe4e3b3d03811c8dc01b5d6914715d9a0

Observation 7f71c83e-dfd6-4dcf-b82e-b45d171c540b · outbound

This paper cites Engineering, Technology & Applied Science Research , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Engineering, Technology & Applied Science Research , volume=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.062924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.062924Z digest=sha256:805cb665fe9d5e61104ae3486b97b056f7281f863c86f15346f16dbaf8ff0e41

Observation 49673c17-e9d5-4249-9414-3d001d1ff6f3 · outbound

This paper cites Machine-Generated Tweets , author=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Machine-Generated Tweets , author=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.066922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.066922Z digest=sha256:e265dd57dac286d156bb315d6781ebd523f9b3785f391736094a41863d205813

Observation a8c6e79f-2a1f-468c-8481-f059a24e2c7b · outbound

This paper cites Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Natural Language Processing , pages=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.070937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.070937Z digest=sha256:0662655f8d0d3dd306e0647b8389509276a4c85b708d721320aa774168cde714

Observation fa620c31-7410-4c82-ad9c-610d6e605801 · outbound

This paper cites International Conference on Computational Science and Its Applications , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Computational Science and Its Applications , pages=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.075456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.075456Z digest=sha256:99fdbbea8c649cd8d12461a0d6df9d796fce5cc0d98efba098488b677a5a4d21

Observation ec361d9f-ec3f-4332-88fd-83b657ee20e4 · outbound

This paper cites Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.079715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.079715Z digest=sha256:d1a6347d94cb07b07d1e407f48436da1e1b374db02375e12b0c2f7d8a1f0e08b

Observation 4613330a-9ecb-48f5-bb4f-e743afd3f148 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.084319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.084319Z digest=sha256:c69647425354e59437cc405edfd36cdf65728a0f21b5861141c7a0c4c03c58a9

Observation baf565a0-4225-4baa-9631-8a37500fc0ce · outbound

This paper cites Regulating Hate Speech Created by Generative AI , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Regulating Hate Speech Created by Generative AI , pages=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.088949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.088949Z digest=sha256:125a7e35af528453100f39ace72274b3df43ad53a3bea9e4da0d3b4f3d3e3477

Observation 005b5e48-1803-4b5c-a922-4f6a68568577 · outbound

This paper cites A Study of Slang Representation Methods.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM A Study of Slang Representation Methods

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:08.069771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.092860Z digest=sha256:a41fb6b1084b7c24a962e44d5b8c8eca96dabc37f5c0c64ecb0face3b15a9f91

Observation 6f761ee3-f867-44ab-ad42-3032a44c39ae · outbound

This paper cites Perplexed by Quality: A Perplexity-based Method for Adult and Harmful Content Detection in Multilingual Heterogeneous Web Data.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Perplexed by Quality: A Perplexity-based Method for Adult and Harmful Content Detection in Multilingual Heterogeneous Web Data

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:08.047629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.096997Z digest=sha256:07ba271e4e1004953850541cf2616a7281744c918dc04a7c16b99f12423b837b

Observation 52f71db4-b053-4998-a6a1-ac5e0bbd3db4 · outbound

This paper cites Human-Guided Fair Classification for Natural Language Processing.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Human-Guided Fair Classification for Natural Language Processing

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.101234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.101234Z digest=sha256:5430e98c304eae3166073746d6175078884113d63dae8d828a06ea3b03436f41

Observation 65b692fc-8d19-4aab-9192-b7358227d056 · outbound

This paper cites arXiv preprint arXiv:2301.12534 , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM arXiv preprint arXiv:2301.12534 , year=

Reference 74

Resolution
verified exact
raw_fallback, observed 2026-08-05T23:13:08.009527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.105543Z digest=sha256:1c09b4a4c86f2cc3d9c7fa9e4d136ba5604ec0bcf2f1bbccee677f764af5829e

Observation 3042f772-ca5c-4b91-8fcb-c770d6a45a37 · outbound

This paper cites International Conference on Advances in Social Networks Analysis and Mining , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM International Conference on Advances in Social Networks Analysis and Mining , pages=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.109738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.109738Z digest=sha256:b14ad1e68559a715b9ab94a7caa80d270f88fc7c771178398ab0c2998502cb54

Observation 5f7ed78f-2433-4640-90f3-a95f2603ed0a · outbound

This paper cites Explicit Toxicity Detection Models with Interactive Visualization , year=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Explicit Toxicity Detection Models with Interactive Visualization , year=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.114159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.114159Z digest=sha256:c021a0c215bddd3efb100e934271a5e01117a87c03aa4c8a897aff4358956260

Observation 164a9473-9137-4419-a9a1-adad34d60acc · outbound

This paper cites Which Argumentative Aspects of Hate Speech in Social Media can be reliably identified?.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Which Argumentative Aspects of Hate Speech in Social Media can be reliably identified?

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.905145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.118824Z digest=sha256:cdea47d5bc52c76daf6e0c02a7e0e6943bf4e4e7c86bbbb20216a245cb67e616

Observation 25099d55-a1e2-4d1b-a80b-f0077fbc4f34 · outbound

This paper cites Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI) , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI) , pages=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.123304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.123304Z digest=sha256:0eaebd5a6199468b44cb6859a1ecd44689e3d34bdd82026340830288353e5a27

Observation a458cb7d-4282-49cc-bc73-972d9425e3d6 · outbound

This paper cites an unresolved cited work.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.127327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.127327Z digest=sha256:7f503b9296a73160cae9e7bc979ed6c8c7cff341c925aadda91d853f905a5a23

Observation 803845f0-f56b-479d-9657-43c8d43a3d27 · outbound

This paper cites Topological Data Mapping of Online Hate Speech, Misinformation, and General Mental Health: A Large Language Model Based Study.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Topological Data Mapping of Online Hate Speech, Misinformation, and General Mental Health: A Large Language Model Based Study

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:13:07.883128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.131563Z digest=sha256:8400455b6d4e855e8bc38ab00566b49ae4a5a6079526423c1cfad829efaae0b8

Observation 53d99bd2-5e9f-4264-a72f-8f076dcc6532 · outbound

This paper cites Demonstrations Are All You Need: Advancing Offensive Content Paraphrasing using In-Context Learning.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Demonstrations Are All You Need: Advancing Offensive Content Paraphrasing using In-Context Learning

Reference 81

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.861922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.135809Z digest=sha256:508abe5edca4c3b7cd225169ea6bf5aff715acff576211031b1f506e25293ce7

Observation d6661e22-6d2c-4b57-b6d2-934321f5527a · outbound

This paper cites Beyond plain toxic: building datasets for detection of flammable topics and inappropriate statements.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Beyond plain toxic: building datasets for detection of flammable topics and inappropriate statements

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.139949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.139949Z digest=sha256:29e63f836b59ef6445d2f3a3ea449043e455bec623637b195e82d9564066425c

Observation d9798d37-cc31-433b-a0b5-a14eb29082c5 · outbound

This paper cites Evaluation of ChatGPT and BERT-based models for Turkish hate speech detection.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Evaluation of ChatGPT and BERT-based models for Turkish hate speech detection

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.144250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.144250Z digest=sha256:55281f64e8317cb788d8d370c74308b211e6da74f3594403c223f3a0a1e82b6c

Observation 5ec65427-89d8-4191-b8ae-ac3903d98754 · outbound

This paper cites HateRephrase: Zero- and Few-Shot Reduction of Hate Intensity in Online Posts using Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM HateRephrase: Zero- and Few-Shot Reduction of Hate Intensity in Online Posts using Large Language Models

Reference 84

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.840022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.148362Z digest=sha256:f0993d0f76ebdf0272ac6bcbb1f9f85214e3aa3259262d319daa0f5f341fe747

Observation 10435c6c-a17d-4409-9d63-6408f817ac91 · outbound

This paper cites FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.152649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.152649Z digest=sha256:fc9f11b4312ff29a5f13846b8e92271c4aa4870732a09c9ee8dafd9bf412c10b

Observation 302948ef-6193-4c33-911b-47b1b120795a · outbound

This paper cites Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.157130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.157130Z digest=sha256:bb5bd67dc1e8f77d99ec5872423f9a3ad0e8e0fc2caf92c2fddb9af55587d8c6

Observation a5365e4f-0f2d-47a2-8d76-5b266eafd26e · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.161107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.161107Z digest=sha256:b1ab7619c3ac4eb51d19089a40c2bd1c747f2137be11a44441be20be107d67a9

Observation 78766c66-76e8-4a0a-9996-8cb3380fc3b7 · outbound

This paper cites 2024 , isbn =.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM 2024 , isbn =

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.165543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.165543Z digest=sha256:80af2ea8c77e7b1bd4425df035034427a59f2c9a711fd04bd289c489a569bc07

Observation dd78b58d-881b-4fd8-9d82-aad9612a5cf3 · outbound

This paper cites Eagle: Ethical Dataset Given from Real Interactions.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Eagle: Ethical Dataset Given from Real Interactions

Reference 89

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.712366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.169918Z digest=sha256:c6302fcf2ef95586f7bbdd97a1f59eb4e72860d1dfb884fc45dae12cb530aedc

Observation 54e798c2-c413-47ea-bb80-6a08a6c8fc0c · outbound

This paper cites Proceedings of the first workshop on language technology for equality, diversity and inclusion , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the first workshop on language technology for equality, diversity and inclusion , pages=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.174303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.174303Z digest=sha256:47a1d010ebb6f2f96102bdcb6718656783e3de281dd889670424030fce00a170

Observation 54f73beb-ca2e-4f98-b9b6-d76066cfccfe · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.178745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.178745Z digest=sha256:1a953c89b51d48d0b2ca8aee3ce5f1e46a15bb85fe92d535df9e5378752c2932

Observation 34bdbfab-2e6a-4ead-bfb0-5ad5db868f85 · outbound

This paper cites M isgender M ender: A Community-Informed Approach to Interventions for Misgendering.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM M isgender M ender: A Community-Informed Approach to Interventions for Misgendering

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.183378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.183378Z digest=sha256:49161096666c9215ad13c1d1145bc42a71164e4af58ac237eace0e0b0421e95d

Observation 41a2da99-3a2f-442a-9f47-98af30d9b87c · outbound

This paper cites HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.187341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.187341Z digest=sha256:d40e094c5defa8997c2f42dcbd809f07462affae0604ff220cb962259567a190

Observation d422327e-7c3c-459e-9c4c-774677d84c23 · outbound

This paper cites Toxicity Classification in Ukrainian.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Toxicity Classification in Ukrainian

Reference 94

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.675095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.191642Z digest=sha256:259c36c70f4e5c35b3058f61f154d0943b6967a700989fa8e047a18b3598ab93

Observation b7910204-72e1-4448-a7f3-d09091fa36b5 · outbound

This paper cites Proceedings of the International AAAI Conference on Web and Social Media , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the International AAAI Conference on Web and Social Media , volume=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.196099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.196099Z digest=sha256:a2900e64fd6de8ed06d498c4417c69498c8f9b6ccb1547073d55587ceef62614

Observation 2fae51d7-23e7-4aef-aae7-40ffa4c11406 · outbound

This paper cites Proceedings of the International AAAI Conference on Web and Social Media , volume=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the International AAAI Conference on Web and Social Media , volume=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.200044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.200044Z digest=sha256:326aa36983e9b6083ba72ee06f8aceafb188aa518e8451f5da5ed69965291df4

Observation ea2dfb9d-e2ee-4ff1-b6a5-f35c656b1f1a · outbound

This paper cites Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop , pages=.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop , pages=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.204022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.204022Z digest=sha256:614ad4e6ac2ec448efb6156a0a01eb7980c6a2cb2b606cb771b03dfeae3cc1ea

Observation 0e70d85c-1aa2-4384-84c5-2ffbfd12e9b8 · outbound

This paper cites A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech

Reference 98

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T23:13:07.653077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T23:13:04.208197Z digest=sha256:411168317e7ab922646888479e33082d29d221e5e004edce9ff9a6125de62667

Observation fc7888dc-7bf9-4d8b-8bba-3b06a2bb1d91 · outbound

This paper cites and Saha, Sriparna and Pasupa, Kitsuchart , title =.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM and Saha, Sriparna and Pasupa, Kitsuchart , title =

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.213203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.213203Z digest=sha256:0e8bfd3dace01c6d4d70cee5141e03797e3f0826304eed8f9b84c4302e2ea450

Observation 35c4bcef-1500-4a11-a17b-48205080ce62 · outbound

This paper cites On Calibration of LLM-based Guard Models for Reliable Content Moderation.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM On Calibration of LLM-based Guard Models for Reliable Content Moderation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.217122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.217122Z digest=sha256:a3f56d79f8f32daa73b1b151c767ea2a7ddc506f477d25c242ccf306ba9248a8

Pith citing papers

Observation 56e7e935-1af9-4a4d-aa2f-86ed48c0b5b6 · inbound

BarrierSteer: LLM Safety via Learning Barrier Steering cites this paper.

BarrierSteer: LLM Safety via Learning Barrier Steering Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T07:02:03.058731Z digest=sha256:a2b6694216908ec104e826ddc69ab75f58a35972e8b587bf36392b68c0e60915

Observation 9b2ca48d-322a-4c1d-9ec2-44a2779397db · inbound

Why Do Large Language Models Generate Harmful Content? cites this paper.

Why Do Large Language Models Generate Harmful Content? Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:31:13.545599Z digest=sha256:1a9bb2c1b7ebf59eec7e8bec08a581e098ca9f049199831e049ecec4d31439a7

Observation ee3c235e-d736-4b70-b9b0-507b8df37bea · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T11:17:19.079380Z digest=sha256:e24697a52ca83b7bc12e8ce0ddc85ad38f060c0787869e02e60afb53e5c8ab17

Observation 9c3d6022-748a-4c66-b3d8-d7f6c01daa7a · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:3193781f3c7a5c25bcb8a31af155d894abf9ed1b1467598f9d3c5bb72fba943d

Observation a6083394-f2d9-440b-aa9c-11737631a9ca · inbound

Do Coding Agents Understand Least-Privilege Authorization? cites this paper.

Do Coding Agents Understand Least-Privilege Authorization? Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T16:34:14.379419Z digest=sha256:d5ec72d4d3c455ecf93c2fea97ab96060b253218efee7da256e0a26504c2cff2

Observation 62ac8f37-4f52-40ad-b1cc-44d7042493d1 · inbound

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows cites this paper.

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T07:06:20.077337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:06:20.077337Z digest=sha256:eb7d3b14980c5e9abd52425b04ae545541f7e10934808d67cbc0c1b60a29a07e