Pith. sign in

Llama guard 3-1b-int4: Compact and 9 efficient safeguard for human-ai conversations

9 Pith papers cite this work. Polarity classification is still indexing.

9 Pith papers citing it

citation-role summary

other 1

citation-polarity summary

years

2026 7 2025 2

roles

other 1

polarities

unclear 1

representative citing papers

RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks

cs.CR · 2026-06-06 · unverdicted · novelty 6.0

RecurGuard monitors recurrence rate, volume growth, and query progress in exposed reasoning traces to terminate generation on token-consumption attacks, reporting 99% detection on OverThink and 92% on ExtendAttack with near-zero false positives.

Distilling Safe LLM Systems via Soft Prompts for On Device Settings

cs.LG · 2026-06-08 · unverdicted · novelty 4.0

Soft prompt distillation with total variation and KL divergence transfers safety behaviors from guard models to on-device LLMs and outperforms LoRA adapters, steering vectors, and direct optimization in safety-usefulness trade-offs with minimal inference cost.

citing papers explorer

Showing 9 of 9 citing papers.