REVIEW 3 major objections 4 minor 94 references
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that an automated red teaming and adversarial training loop, with no Process Reward Model, can beat PRM-based security alignment while cutting computational cost by 61%.
desk verdict A plausible-sounding framework paper whose headline cost claim is contradicted by its own tables; with no released models or code, nothing load-bearing can be checked. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the closed attack–train loop. The red-teaming side scores each candidate attack prompt with a fitness function $f(x) = \alpha \cdot \mathrm{ASR}(x) + \beta \cdot \mathrm{SIM}(x, x_{\mathrm{orig}}) + \gamma \cdot \mathrm{DIV}(x, P) + \delta \cdot \mathrm{TRANS}(x) + \epsilon \cdot \mathrm{SEVER}(x)$ that balances attack success, semantic similarity, population diversity, cross-model transfer, and severity. The training side minimizes a five-term loss $L_{\mathrm{total}} = \lambda_1 L_{\mathrm{standard}} + \lambda_2 L_{\mathrm{adversarial}} + \lambda_3 L_{\mathrm{regularization}} + \lambda_4 L_{\mathrm{alignment}} + \lambda_5 L_{\mathrm{utility}}$ that combines language modeling, adversarial robustness, regularization, alignment, and utility objectives. Curriculum learning ramps attack difficulty, and adaptive regularization (Elastic Weight Consolidation plus memory replay) is claimed to prevent catastrophic forgetting. This loop replaces PRM step scoring with attack-level and outcome-level scores from evaluator agents, which is what makes the 61% cost reduction possible.
What would settle it
Run the PRM-free pipeline and the PRM baseline against a fixed, externally held-out set of attacks—for instance, a public jailbreak benchmark scored by human raters who never see the framework's own evaluator scores—and compare how often each model produces harmful responses. If the external scores do not show the PRM-free model at least matching the PRM baseline on safety, the paper's central superiority claim fails.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that outcome-level signals from attacks can replace step-level reward signals in security alignment. The paper reports that the PRM-free framework finds vulnerabilities with 68.2% attack success rate versus 56.7% for PRM-Basic, covers 89% of known attack vectors versus 71% for PRM-based methods, and improves robustness scores by 12–18% across five LLMs, while using 480 GPU-hours versus 1240 for the PRM baseline. It also reports utility preservation near 97% on HellaSwag, 95.8% on MMLU, and 94.1% on HumanEval, with only 2.1% benign-task degradation versus 8.7% for standard adversarial training. The claimed reason is that the diversity, severity, and transferability of evolved attacks make a better curriculum than step-by-step reward feedback.
Load-bearing premise
The load-bearing premise is that the framework's own evaluator agents, using automated metrics and simulated human judgment, measure security accurately enough to compare methods; if those self-administered scores are biased, the reported superiority over PRM baselines is an artifact.
Editorial extensions
If this is right
- If the central claim is right, organizations can skip training a step-level reward model and instead run a red teaming plus adversarial training cycle, getting comparable or better safety at 61% lower compute.
- The framework's discovered attack library, with 34% more unique attack patterns than baselines, could be reused as a portable adversarial training or evaluation set for other models.
- The reported 84% cross-model attack transferability means a single red teaming campaign on one architecture would expose weaknesses relevant to related LLMs, making the output act like a shared security benchmark.
- The claimed utility preservation (94–97% on standard benchmarks) means the safety gains would not require giving up ordinary task performance.
Reading between the lines
- Beyond the paper, the same evaluator-agent loop could function as a continuous production monitoring layer, so security alignment becomes an ongoing operational process rather than a one-time training phase.
- A direct test of the claim would freeze the discovered vulnerability library and release it as a static benchmark, letting independent groups check whether the defenses transfer when the attack generator is no longer adaptive.
- The 61% figure covers the aligned model's training and inference; an independent accounting of the red teaming loop's own compute is needed before calling it an end-to-end saving.
- Replacing the framework's simulated human judgment with human raters on a fixed attack sample would separate the value of the red teaming method from any bias in the self-administered metrics.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a PRM-free framework for LLM security alignment that combines automated red teaming (genetic algorithms, multi-agent simulation, prompt mutation) with adversarial training (curriculum learning, adaptive regularization, multi-objective losses), plus reporting and audit components. The abstract and Section 1 claim that this approach achieves superior security alignment compared to PRM-based baselines while reducing computational costs by 61%. Experiments are reported on five anonymized models and seven baselines, with metrics such as Attack Success Rate (ASR), Vulnerability Severity Index (VSI), and Robustness Score (RS). The central quantitative claims are an ASR of 68.2% versus 56.7% for PRM-Basic and 480 versus 1240 GPU-hours for the proposed method versus PRM-Basic.
Significance. If the reported results were reproducible and the metrics externally validated, the proposed pipeline would be a practically significant contribution, lowering the compute barrier for security alignment and offering a scalable red-teaming-plus-training loop. The framework description is detailed, the authors include utility benchmarks alongside security metrics, and Section 7.5 honestly lists several limitations. However, the paper ships no code or data, the evaluated models are anonymized, the core security metrics are computed by the framework's own evaluator agents, and the headline 61% efficiency claim is internally contradicted by the tables. As a result, the significance of the contribution cannot be assessed from the manuscript as submitted.
major comments (3)
- [§5.2.2, Table 2; §5.2.1, Table 1; §4.1] The headline claim of a 61% computational-cost reduction is internally contradicted by the reported numbers. Table 1 lists the proposed method at 9.2 h versus 18.5 h for PRM-Basic (a 50.3% reduction), while Table 2 lists 7.8 h versus 18.2 h (a 57.1% reduction) yet reports 480 versus 1240 GPU-hours (a 61.3% reduction). For fixed hardware, the percentage reduction in GPU-hours must equal the percentage reduction in wall-clock time; the discrepancy implies different GPU counts for the two tables. Moreover, §4.1 and §5.1.3 state that primary experiments used 8×NVIDIA A100 GPUs, so 480 GPU-hours corresponds to 60 h, not 7.8 or 9.2 h. These inconsistencies mean the abstract's quantitative efficiency claim is not supported by the manuscript's own data.
- [§3.2.2, §5.1.2, §5.2.1] The core security metrics (ASR, VSI, RS) are computed by the framework's own evaluator agents using "automated metrics and simulated human judgment" (§3.2.2). The central claim of superior security alignment—e.g., 68.2% ASR versus 56.7% for PRM-Basic in Table 1—rests entirely on these self-administered scores. No external validation is provided correlating these scores with human judgments, independent jailbreak benchmarks, or a fixed public attack suite. As a result, the reported security gains may be artifacts of the evaluator rather than genuine model improvement; this is load-bearing because no other evidence supports the superiority claim.
- [§5.1.1, §5.1.2, §5.1.5] The empirical evaluation is not reproducible as presented. All five "state-of-the-art" models are anonymized (Model A–E), baseline implementations (PRM-Basic, PRM-Advanced, etc.) are described only generically, and no code, data, prompts, or trained outputs are released. Section 5.1.5 lists standard benchmarks such as HellaSwag and MMLU, but the main comparative tables are not tied to any public artifact. For a paper whose contribution is an empirical framework with custom metrics, this makes the central results uncheckable, independent of the internal inconsistencies.
minor comments (4)
- [§5.2.1, Table 1, Table 2] Training time for the proposed method is 9.2 h in Section 5.2.1 and 7.8 h in Table 2; the PRM-Basic, PRM-Advanced, RLHF-Standard, and Constitutional AI entries also differ between the two tables. Please label which configuration each table refers to or unify the numbers.
- [§5.2.1, §6.4] Section 5.2.1 reports an "Extended evaluation over 30 days," while Section 6.4 reports "Extended evaluation over 6 months"; please reconcile the evaluation period.
- [§3.3.2, §3.2.1] Equation (1) and Equation (2) introduce five fitness weights and five loss weights, but no default values, tuning procedure, or sensitivity analysis are reported; please provide the settings used in the experiments.
- [§5.2.4] The claim that all improvements are statistically significant at p < 0.001 is not supported by any reported test statistic, variance, confidence interval, or per-model breakdown; please provide details of the bootstrap procedure and the actual intervals.
Circularity Check
Partial circularity: the vulnerability-discovery superiority is the genetic algorithm's own fitness objective; the efficiency and external-benchmark results remain independent.
-
fitted input called prediction
[Eq. (1) in Section 3.2.1; reported in Section 5.2.1, Table 1]
"Eq. (1): 'f (x) =α · ASR(x) +β · SIM (x, xorig) + γ · DIV (x, P) +δ · T RAN S(x) + ε · SEV ER(x)'. Section 5.2.1: 'Table 1 shows our approach achieving 68.2% ASR versus 56.7% for PRM-Basic and 42.3% for Manual-RT, with superior vulnerability severity (VSI 4.2 vs. 3.1) and diversity (ADM 3.9 vs. 2.4)'."
The genetic algorithm's fitness function in Eq. (1) directly maximizes ASR(x), DIV(x,P), TRANS(x), and SEVER(x). Table 1 then presents Ours as superior on exactly those same quantities (ASR, ADM, Transferability, VSI) relative to baselines that are not optimized with Eq. (1). The reported 'superior vulnerability discovery' is the optimizer reporting its own objective values, not an independent finding; the superiority on those metrics is built into the fitness definition. This partially circularly supports the abstract's 'superior security alignment performance' claim, although external utility/toxicity benchmarks and the efficiency measurements are independent.
full rationale
The paper is an empirical engineering report rather than a derivation, and most of it is not circular: there are no fitted parameters renamed as predictions, no self-citations, no imported uniqueness theorems, and no ansatz smuggled in via prior work. The external benchmarks (ToxiGen, RealToxicityPrompts, HellaSwag, MMLU, HumanEval) provide independent content for the utility and safety claims. The one concrete circularity is localized: the red-teaming GA is instructed in Eq. (1) to maximize ASR, diversity, transferability, and severity, and Table 1 then reports the framework as superior on those same metrics. That is the optimizer celebrating its objective function, so part of the 'superior security alignment' claim is circular by construction. I did not count the framework's self-administered evaluator metrics (Section 3.2.1) as a separate formal circular step because that is a measurement-validity concern rather than an equation-level reduction; similarly, the abstract's 61% compute reduction is internally inconsistent with the wall-clock figures in Tables 1-2 (about 50-57%), but inconsistency is a correctness problem, not circularity. Overall the central PRM-free mechanism and the efficiency story have independent content, so the circularity score is modest.
Assumptions & free parameters
free parameters (4)
- fitness weights alpha, beta, gamma, delta, epsilon
- loss weights lambda_1..lambda_5
- GA population size, tournament size, mutation rate =
100, 5, 0.1
- training iterations, batch size, learning rate =
5000, 32, 1e-5 to 1e-4
assumptions (3)
- domain assumption PRM-based methods are the dominant and appropriate baseline for security alignment.
- ad hoc to paper The framework's self-defined metrics (ASR, VSI, RS) are valid and unbiased measures of security.
- domain assumption Models A-E and the seven baselines are real, distinct, and correctly configured.
Cite this review
Pith. "Pith review of PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training." pith.science (2026). https://pith.science/paper/PPM52GXQ
@misc{pith2026250714202,
author = {Pith},
title = {Pith review of: PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training},
year = {2026},
howpublished = {\url{https://pith.science/paper/PPM52GXQ}},
note = {Machine review of arXiv:2507.14202}
}
read the original abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet they pose significant security risks that threaten their safe deployment in critical domains. Current security alignment methodologies predominantly rely on Process Reward Models (PRMs) to evaluate intermediate reasoning steps, introducing substantial computational overhead and scalability constraints. This paper presents a novel PRM-free security alignment framework that leverages automated red teaming and adversarial training to achieve robust security guarantees while maintaining computational efficiency. Our approach systematically identifies vulnerabilities through sophisticated attack strategies including genetic algorithm optimization, multi-agent simulation, and advanced prompt mutation techniques. The framework enhances model robustness via targeted adversarial training with curriculum learning and adaptive regularization mechanisms. Comprehensive experimental evaluation across five state-of-the-art LLMs demonstrates that our method achieves superior security alignment performance compared to PRM-based approaches while reducing computational costs by 61\%. The framework incorporates transparent reporting and continuous audit mechanisms that enable iterative security improvement and regulatory compliance. Our contributions advance the field of efficient LLM security alignment by democratizing access to robust security measures for resource-constrained organizations and providing a scalable foundation for addressing evolving adversarial threats.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Karan Ahmed, Mahdi Babaei, Ignacio Blanco, Nicolas Gast, Vincent Leroy, and Mathias Lecuyer. 2022. Measuring the carbon intensity of ai in cloud instances. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 1877--1887
2022
-
[4]
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018. Generating natural language adversarial examples. arXiv preprint arXiv:1804.07998
arXiv 2018
-
[5]
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, and 1 others. 2021. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861
arXiv 2021
-
[6]
Eugene Bagdasaryan and Vitaly Shmatikov. 2021. Spinning language models: Risks of propaganda-as-a-service and countermeasures. arXiv preprint arXiv:2112.05224
work page Pith review arXiv 2021
-
[7]
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, and 1 others. 2022 a . Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862
arXiv 2022
-
[9]
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, and 1 others. 2022 c . Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073
arXiv 2022
Show all 94 references
-
[10]
Luke Bailey, Euan Ong, Stuart Russell, and Scott Emmons. 2023. Image hijacks: Adversarial images can control generative models at runtime. arXiv preprint arXiv:2309.00236
2023 arXiv
-
[11]
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610--623
2021
-
[12]
Yoshua Bengio, J \'e r \^o me Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum learning. Proceedings of the 26th annual international conference on machine learning, pages 41--48
2009
-
[13]
Battista Biggio and Fabio Roli. 2018. Wild patterns: Ten years after the rise of adversarial machine learning. Pattern recognition, 84:317--331
2018
-
[14]
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, and 1 others. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258
2021 arXiv
-
[15]
Louise Branch and Dustin Benton. 2022. Evaluating the susceptibility of pre-trained language models via handcrafted adversarial examples. arXiv preprint arXiv:2209.02128
2022 arXiv
-
[16]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and 1 others. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877--1901
2020
-
[17]
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, and 1 others. 2021. Extracting training data from large language models. 30th USENIX Security Symposium, pages 2633--2650
2021
-
[18]
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, J \'e r \'e my Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, and 1 others. 2023 a . Open problems and fundamental limitations of reinforcement learning from human feedback...
2023 arXiv
-
[19]
Stephen Casper, Jason Li, Jiawei Li, Javier Rando, and Gabriel Kreiman. 2023 b . Explore, establish, exploit: Red teaming language models from scratch. arXiv preprint arXiv:2306.09442
2023 arXiv
-
[20]
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2023. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419
2023 arXiv
-
[21]
Jianbo Chen, Michael I Jordan, and Martin J Wainwright. 2020. Hopskipjumpattack: A query-efficient decision-based attack. IEEE Symposium on Security and Privacy
2020
-
[22]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, and 1 others. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374
2021 arXiv
-
[23]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, and 1 others. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311
2022 arXiv
-
[24]
Brian Christian. 2020. The alignment problem: Machine learning and human values. WW Norton & Company
2020
-
[25]
Paul Christiano, Buck Buck, Tom Eccles, Jan Leike, Shane Legg, and Dario Amodei. 2018. Supervising strong learners by amplifying weak experts. arXiv preprint arXiv:1810.08575
2018 arXiv
-
[26]
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30
2017
-
[27]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, and 1 others. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168
2021 arXiv
-
[28]
Yue Deng, Wenxuan Zhang, Sinno Jialin Ong, and Lidong Tan. 2023. Multilingual jailbreak challenges in large language models. arXiv preprint arXiv:2310.06474
2023 arXiv
-
[29]
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021. Bold: Dataset and metrics for measuring biases in open-ended language generation. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transpa...
2021
-
[30]
Emily Dinan, Gavin Abercrombie, A Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser. 2019. Safety for e2e conversational ai: A human-in-the-loop evaluation of conversation models using safety-focused human feedback. arXiv preprint arXiv:1911.07754
2019 arXiv
-
[31]
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2022. Safetykit: First aid for measuring safety in open-domain conversational ai. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, pages 4113--4133
2022
-
[32]
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2017. Hotflip: White-box adversarial examples for text classification. arXiv preprint arXiv:1712.06751
2017 arXiv
-
[33]
Luciano Floridi, Josh Cowls, Monica Beltrametti, Raja Chatila, Patrice Chazerand, Virginia Dignum, Christoph Luetge, Robert Madelin, Ugo Pagallo, Francesca Rossi, and 1 others. 2020. Translating uncertainty about ai into the language of risk. AI & Society, 35(4):947--963
2020
-
[34]
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, and 1 others. 2022. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint ar...
2022 arXiv
-
[35]
Leo Gao, John Schulman, and Jacob Hilton. 2022. Scaling laws for reward model overoptimization. arXiv preprint arXiv:2210.10760
2022 arXiv
-
[36]
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462
2020 arXiv
-
[37]
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572
2014 arXiv
-
[38]
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023. Not what you've signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. arXiv preprint arXiv:2302.12173
2023 arXiv
-
[39]
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. Proceedings of the 60th Annual Meeting of the Association for Computational Li...
2022
-
[40]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300
2020 arXiv
-
[41]
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, and 1 others. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556
2022 arXiv
-
[42]
Geoffrey Irving, Paul Christiano, and Dario Amodei. 2018. Ai safety via debate. arXiv preprint arXiv:1805.00899
2018 arXiv
-
[43]
Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao. 2020. Smart: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization. Proceedings of the 58th Annual Meeting of the Association for ...
2020
-
[44]
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020. Is bert really robust? a strong baseline for natural language attack on text classification and entailment. Proceedings of the AAAI conference on artificial intelligence, 34(05):8018--8025
2020
-
[45]
Anna Jobin, Marcello Ienca, and Effy Vayena. 2019. The global landscape of ai ethics guidelines. Nature Machine Intelligence, 1(9):389--399
2019
-
[46]
Erik Jones, Anca Dragan, Aditi Raghunathan, and Jacob Steinhardt. 2023. Automatically auditing large language models via discrete optimization. arXiv preprint arXiv:2303.04381
2023 arXiv
-
[47]
Hannah Rose Kirk, Haider Iqbal, Elias Benussi, Frederic Volpin, Frederic Dreyer, Yuki M Asano, and Russell Cavendish. 2023. Understanding and mitigating the uncertainty in zero-shot translation. arXiv preprint arXiv:2311.02520
2023 arXiv
-
[48]
Raz Lapid, Ron Langberg, and Moshe Sipper. 2023. Open sesame! universal black box jailbreaking of large language models. arXiv preprint arXiv:2309.01446
2023 arXiv
-
[49]
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. 2018. Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871
2018 arXiv
-
[50]
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020. Bert-attack: Adversarial attack against bert using bert. arXiv preprint arXiv:2004.09984
2020 arXiv
-
[51]
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. Let's verify step by step. arXiv preprint arXiv:2305.20050
2023 arXiv
-
[52]
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023 a . Jailbreaking chatgpt via prompt engineering: An empirical study. arXiv preprint arXiv:2305.13860
2023 arXiv
-
[53]
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023 b . Prompt injection attacks and defenses in llm-integrated applications. arXiv preprint arXiv:2310.12815
2023
-
[54]
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083
2017 arXiv
-
[55]
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM computing surveys, 54(6):1--35
2021
-
[56]
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. 2023. Tree of attacks: Jailbreaking black-box llms automatically. arXiv preprint arXiv:2312.02119
2023 arXiv
-
[57]
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, and 1 others. 2022. Teaching language models to support answers with verified quotes. arXiv preprint arXiv:2203.11147
2022 arXiv
-
[58]
George A Miller. 1995. Wordnet: a lexical database for english. Communications of the ACM, 38(11):39--41
1995
-
[59]
Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. 2018. Virtual adversarial training: a regularization method for supervised and semi-supervised learning. IEEE transactions on pattern analysis and machine intelligence, 41(8):1979--1993
2018
-
[60]
John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yilun Qi. 2020. Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp. arXiv preprint arXiv:2005.05909
2020 arXiv
-
[61]
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, and 1 others. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332
2021 arXiv
-
[62]
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A Feder Cooper, Daphne Ippolito, Christopher A Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. 2023. Scalable extraction of training data from (production) language models. arXiv preprint arXiv:...
2023 arXiv
-
[63]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing sys...
2022
-
[64]
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 a . Red teaming language models with language models. arXiv preprint arXiv:2202.03286
2022 arXiv
-
[65]
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, and 1 others. 2022 b . Discovering language model behaviors with model-written evaluations. arXiv preprint arXiv:2212.09251
2022 arXiv
-
[66]
Fabi \'a n P \'e rez, Ian Ribeiro, Jonatan Malmaud, Rachel Rudick, and Edward Grefenstette. 2022. Ignore previous prompt: Attack techniques for language models. arXiv preprint arXiv:2211.09527
2022 arXiv
-
[67]
Emmanouil Antonios Platanios, Otilia Stretcu, Graham Neubig, Barnabas Poczos, and Tom M Mitchell. 2019. Competence-based curriculum learning for neural machine translation. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Li...
2019
-
[68]
Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John C Duchi, and Percy Liang. 2019. Adversarial training can hurt generalization. arXiv preprint arXiv:1906.06032
2019 arXiv
-
[69]
Stuart Russell. 2019. Human compatible: Artificial intelligence and the problem of control. Viking
2019
-
[70]
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347
2017 arXiv
-
[71]
Rusheb Shah, Stephen Casper, and Javier Rando. 2023. Scalable and transferable black-box jailbreaks for language models via persona modulation. arXiv preprint arXiv:2311.03348
2023 arXiv
-
[72]
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adri \`a Garriga-Alonso, and 1 others. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. ...
2022 arXiv
-
[73]
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008--3021
2020
-
[74]
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and policy considerations for deep learning in nlp. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics
2019
-
[75]
Florian Tramer, Alexey Kurakin, Nicolas Papernot, Ian Goodfellow, Dan Boneh, and Patrick McDaniel. 2017. Ensemble adversarial training: Attacks and defenses. arXiv preprint arXiv:1705.07204
2017 arXiv
-
[76]
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry. 2018. Robustness may be at odds with accuracy. arXiv preprint arXiv:1805.12152
2018 arXiv
-
[77]
Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. 2022. Solving math word problems with process-and outcome-based feedback. arXiv preprint arXiv:2211.14275
2022 arXiv
-
[78]
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019. Universal adversarial triggers for attacking and analyzing nlp. arXiv preprint arXiv:1908.07125
2019 arXiv
-
[79]
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li. 2021 a . Advglue: A multi-task benchmark for adversarial robustness evaluation of language models. arXiv preprint arXiv:2111.02840
2021 arXiv
-
[80]
Tingwu Wang, Renjie Liao, Jimmy Ba, and Sanja Fidler. 2021 b . Learning robust, real-time, reactive robotic movement. The International Journal of Robotics Research, 40(2-3):260--277
2021
-
[81]
Xiaosen Wang, Hao Jin, and Kun He. 2019. Natural language adversarial attacks and defenses in word level. arXiv preprint arXiv:1909.06723
2019 arXiv
-
[82]
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171
2022 arXiv
-
[83]
Yisen Wang, Difan Zou, Jinfeng Yi, James Bailey, Xingjun Ma, and Quanquan Gu. 2020. Improving adversarial robustness requires revisiting misclassified examples. International Conference on Learning Representations
2020
-
[84]
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How does llm safety training fail? arXiv preprint arXiv:2307.02483
2023 arXiv
-
[85]
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, and 1 others. 2021. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359
2021 arXiv
-
[86]
Zhengxuan Xu, Neel Jain, and Tom Goldstein. 2022. Exploring the landscape of distributional robustness for question answering models. arXiv preprint arXiv:2210.12517
2022 arXiv
-
[87]
Huanrui Yang, Jingyang Zhang, Hongliang Dong, Nathan Inkawhich, Andrew Gardner, Andrew Touchet, Wesley Wilkes, Heath Berry, and Hai Li. 2020. Dverge: Diversifying vulnerabilities for enhanced robust generation of ensembles. Advances in Neural Information Processing Systems, 33...
2020
-
[88]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629
2022 arXiv
-
[89]
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. 2023. Low-resource languages jailbreak gpt-4. arXiv preprint arXiv:2310.02446
2023 arXiv
-
[90]
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. 2023. Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts. arXiv preprint arXiv:2309.10253
2023 arXiv
-
[91]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4800
2019
-
[92]
Xiaodong Zhang, Junqi Zhao, and Yann LeCun. 2020. Adversarial training for large neural language models. arXiv preprint arXiv:2004.08994
2020 arXiv
-
[93]
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2019. Freelb: Enhanced adversarial training for natural language understanding. arXiv preprint arXiv:1909.11764
2019 arXiv
-
[94]
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2022. Adversarial training for high-stakes reliability. arXiv preprint arXiv:2205.01663
2022 arXiv
-
[95]
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043
2023 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.