Pith. sign in

Safellm:Unlearningharmfuloutputsfromlargelanguagemodelsagainstjailbreakattacks

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.CR 2 cs.CL 1

years

2026 3

verdicts

UNVERDICTED 3

roles

background 1

polarities

background 1

representative citing papers

EVA: Editing for Versatile Alignment against Jailbreaks

cs.CR · 2026-05-14 · unverdicted · novelty 6.0

EVA applies direct model editing to surgically neutralize jailbreak vulnerabilities in LLMs and VLMs by targeting specific neurons while preserving general capabilities.

Exclusive Unlearning

cs.CL · 2026-04-07 · unverdicted · novelty 6.0

Exclusive Unlearning makes LLMs safe by forgetting all but retained domain knowledge, protecting against jailbreaks while preserving useful responses in areas like medicine and math.

citing papers explorer

Showing 3 of 3 citing papers.

  • EVA: Editing for Versatile Alignment against Jailbreaks cs.CR · 2026-05-14 · unverdicted · none · ref 44

    EVA applies direct model editing to surgically neutralize jailbreak vulnerabilities in LLMs and VLMs by targeting specific neurons while preserving general capabilities.

  • Exclusive Unlearning cs.CL · 2026-04-07 · unverdicted · none · ref 7

    Exclusive Unlearning makes LLMs safe by forgetting all but retained domain knowledge, protecting against jailbreaks while preserving useful responses in areas like medicine and math.

  • From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI cs.CR · 2026-05-15 · unverdicted · none · ref 71

    The paper analyzes evolving security and safety threats in generative AI from content generation to agentic actions, noting that attack surfaces expand faster than defenses and that many safeguards require institutional coordination not yet in place.