Pith. sign in

Systematic benchmarking of guardrails against prompt injection attacks

4 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

4 Pith papers citing it
1 external citations · external index

citation-role summary

background 2 other 1

citation-polarity summary

years

2026 3 2025 1

polarities

background 2 unclear 1

representative citing papers

When AI reviews science: Can we trust the referee?

cs.AI · 2026-04-26 · unverdicted · novelty 6.0

AI peer review systems are vulnerable to prompt injections, prestige biases, assertion strength effects, and contextual poisoning, as demonstrated by a new attack taxonomy and causal experiments on real conference submissions.

SoK: Robustness in Large Language Models against Jailbreak Attacks

cs.CR · 2026-05-06 · accept · novelty 5.0

The paper taxonomizes jailbreak attacks and defenses for LLMs, introduces the Security Cube multi-dimensional evaluation framework, benchmarks 13 attacks and 5 defenses, and identifies open challenges in LLM robustness.

citing papers explorer

Showing 4 of 4 citing papers.