Pith. sign in

hub Mixed citations

Salad-bench: A hierarchical and com- prehensive safety benchmark for large language models

Mixed citation behavior. Most common role is background (40%).

21 Pith papers citing it
Background 40% of classified citations

hub tools

citation-role summary

background 2 dataset 2 other 1

citation-polarity summary

representative citing papers

Robust and Efficient Guardrails with Latent Reasoning

cs.AI · 2026-05-27 · unverdicted · novelty 7.0

COLAGUARD matches explicit-reasoning guardrail performance on safety benchmarks while delivering 12.9X speedup and 22.4X token reduction by propagating hidden states instead of generating text.

Non-linear Interventions on Large Language Models

cs.CL · 2026-05-14 · unverdicted · novelty 7.0

Presents a non-linear intervention framework for LLMs with a learning procedure for implicit features, validated on refusal bypass steering showing improved precision over linear baselines.

Self-Mined Hardness for Safety Fine-Tuning

cs.LG · 2026-05-04 · unverdicted · novelty 6.0 · 2 refs

Self-mined hardness from model rollouts lowers WildJailbreak attack success to 1-3% on Llama-3 models while raising over-refusal, mitigated by 1:1 interleaving with benign prompts.

ShieldGemma: Generative AI Content Moderation Based on Gemma

cs.CL · 2024-07-31 · unverdicted · novelty 4.0

ShieldGemma delivers a family of Gemma2-based classifiers that outperform Llama Guard and WildCard on public safety benchmarks while introducing a synthetic-data curation pipeline for safety tasks.

citing papers explorer

Showing 21 of 21 citing papers.