Pith. sign in

hub

XSTest: A test suite for identifying exaggerated safety behaviours in large language models

24 Pith papers cite this work, alongside 38 external citations. Polarity classification is still indexing.

24 Pith papers citing it
38 external citations · OpenAlex

hub tools

years

2026 22 2025 2

representative citing papers

Korean Culture into LLM Alignment: Toward Cultural Coherence

cs.CL · 2026-06-05 · unverdicted · novelty 6.0

Presents a Korean harm taxonomy, culturally grounded safe-response guidelines, and DPO fine-tuning that raises cultural safe rates on six open-weight LLMs with little benchmark degradation.

ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack

cs.AI · 2025-09-30 · unverdicted · novelty 6.0

ASGuard identifies tense-vulnerable attention heads via circuit analysis, trains precise scaling vectors on those activations, and applies them in preventative fine-tuning to reduce targeted jailbreaking success across four LLMs with minimal impact on utility or over-refusal.

citing papers explorer

Showing 24 of 24 citing papers.