An ablation study of an RL-based jailbreaker finds the attack succeeds across tested open-weight models and safeguards, but its headline conclusion about dense rewards and long episodes is contradicted by its own data.
Vullibgen: Identifying vulnerable third- party libraries via generative pre-trained model
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
A Systematic Investigation of RL-Jailbreaking in LLMs
An ablation study of an RL-based jailbreaker finds the attack succeeds across tested open-weight models and safeguards, but its headline conclusion about dense rewards and long episodes is contradicted by its own data.