The paper claims that converting harmful queries into AMR/RDF/JSON graphs and prompting code generation jailbreaks leading LLMs with up to 87% success, but the evidence is inconsistent and not reproducible.
Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms
The paper claims that converting harmful queries into AMR/RDF/JSON graphs and prompting code generation jailbreaks leading LLMs with up to 87% success, but the evidence is inconsistent and not reproducible.