HarmChip is a new benchmark exposing an alignment paradox where LLMs refuse legitimate hardware security queries but comply with semantically disguised malicious requests.
Metacipher: A general and extensible reinforcement learning framework for obfuscation-based jailbreak attacks on black-box llms
3 Pith papers cite this work. Polarity classification is still indexing.
3
Pith papers citing it
fields
cs.CR 3years
2026 3representative citing papers
An LLM-agent framework with RAG generates structured vulnerability analysis reports from source code, achieving 54.21% average quality on 105 NIST-SARD samples evaluated by an LLM judge.
citing papers explorer
-
HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking
HarmChip is a new benchmark exposing an alignment paradox where LLMs refuse legitimate hardware security queries but comply with semantically disguised malicious requests.
-
RAVEN: Retrieval-Augmented Vulnerability Exploration Network for Memory Corruption Analysis in User Code and Binary Programs
An LLM-agent framework with RAG generates structured vulnerability analysis reports from source code, achieving 54.21% average quality on 105 NIST-SARD samples evaluated by an LLM judge.
- Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking