Pith. sign in

REVIEW 95 references

Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.02654 v2 pith:HOOQTRFB submitted 2025-01-05 cs.CL cs.AI

classification cs.CLcs.AI
keywords adversarialdefencelanguagemodelsnaturalworkbenchmarkbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in natural language processing have highlighted the vulnerability of deep learning models to adversarial attacks. While various defence mechanisms have been proposed, there is a lack of comprehensive benchmarks that evaluate these defences across diverse datasets, models, and tasks. In this work, we address this gap by presenting an extensive benchmark for textual adversarial defence that significantly expands upon previous work. Our benchmark incorporates a wide range of datasets, evaluates state-of-the-art defence mechanisms, and extends the assessment to include critical tasks such as single-sentence classification, similarity and paraphrase identification, natural language inference, and commonsense reasoning. This work not only serves as a valuable resource for researchers and practitioners in the field of adversarial robustness but also identifies key areas for future research in textual adversarial defence. By establishing a new standard for benchmarking in this domain, we aim to accelerate progress towards more robust and reliable natural language processing systems.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

95 extracted references · 12 canonical work pages

  1. [1]

    Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219

  2. [2]

    Balanya, Juan Maro\ n as, and Daniel Ramos

    Sergio A. Balanya, Juan Maro\ n as, and Daniel Ramos. 2024. https://doi.org/10.1007/s00521-024-09505-4 Adaptive temperature scaling for robust calibration of deep neural networks . Neural Comput. Appl., 36(14):8073–8095

  3. [3]

    Rongzhou Bao, Jiayi Wang, and Hai Zhao. 2021. Defending pre-trained language models from adversarial word substitutions without performance sacrifice. arXiv preprint arXiv:2105.14553

  4. [4]

    Ruben Branco, Ant \'o nio Branco, Joao Rodrigues, and Joao Silva. 2021. Shortcutted commonsense: Data spuriousness in deep learning of commonsense reasoning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1504--1521

  5. [5]

    Jianlyu Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024. https://doi.org/10.18653/v1/2024.findings-acl.137 M 3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation . In Findings of the Association for Computational Linguistics: ACL 2024, pages 2318--2335, Bangkok,...

  6. [6]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long a...

  7. [7]

    Bill Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In Third International Workshop on Paraphrasing (IWP2005)

  8. [8]

    Xinshuai Dong, Anh Tuan Luu, Rongrong Ji, and Hong Liu. 2021. Towards robustness against natural language word substitutions. arXiv preprint arXiv:2107.13541

Show all 95 references
  1. [9]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

  2. [10]

    o zde G \

    Steffen Eger, G \"o zde G \"u l S ahin, Andreas R \"u ckl \'e , Ji-Ung Lee, Claudia Schulz, Mohsen Mesgar, Krishnkant Swarnkar, Edwin Simpson, and Iryna Gurevych. 2019 a . https://doi.org/10.18653/v1/N19-1165 Text processing like humans do: Visually attacking and shielding NLP...

  3. [11]

    o zde G \

    Steffen Eger, G \"o zde G \"u l S ahin, Andreas R \"u ckl \'e , Ji-Ung Lee, Claudia Schulz, Mohsen Mesgar, Krishnkant Swarnkar, Edwin Simpson, and Iryna Gurevych. 2019 b . Text processing like humans do: Visually attacking and shielding nlp systems. arXiv preprint arXiv:1903.11508

  4. [12]

    Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi. 2018. Black-box generation of adversarial text sequences to evade deep learning classifiers. In 2018 IEEE Security and Privacy Workshops (SPW), pages 50--56. IEEE

  5. [13]

    Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. https://doi.org/10.18653/v1/2021.acl-long.295 Making pre-trained language models better few-shot learners . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint...

  6. [14]

    Siddhant Garg and Goutham Ramakrishnan. 2020. https://api.semanticscholar.org/CorpusID:214802269 Bae: Bert-based adversarial examples for text classification . ArXiv, abs/2004.01970

  7. [16]

    Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 b . Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572

  8. [17]

    Shreya Goyal, Sumanth Doddapaneni, Mitesh M Khapra, and Balaraman Ravindran. 2023. A survey of adversarial defenses and robustness in nlp. ACM Computing Surveys, 55(14s):1--39

  9. [18]

    Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017. On calibration of modern neural networks. In International conference on machine learning, pages 1321--1330. PMLR

  10. [19]

    Xu Han, Ying Zhang, Wei Wang, and Bin Wang. 2022. Text adversarial attacks and defenses: Issues, taxonomy, and perspectives. Security and Communication Networks, 2022(1):6458488

  11. [20]

    Jonathan Hayase, Ema Borevkovic, Nicholas Carlini, Florian Tram \`e r, and Milad Nasr. 2024. Query-based adversarial prompt generation. arXiv preprint arXiv:2402.12329

  12. [21]

    Jia-long he, Xiao-Lin zhang, Yong-Ping wang, Huan-Xiang zhang, Lu gao, and En-Hui xu. 2023. https://doi.org/10.3233/JIFS-230787 Contrastive adversarial learning in text classification tasks . J. Intell. Fuzzy Syst., 45(2):3473–3484

  13. [22]

    Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020. https://api.semanticscholar.org/CorpusID:219531210 Deberta: Decoding-enhanced bert with disentangled attention . ArXiv, abs/2006.03654

  14. [23]

    Xinrong Hu, Ce Xu, Junlong Ma, Zijian Huang, Jie Yang, Yi Guo, and Johan Barthelemy. 2023. https://doi.org/10.18653/v1/2023.findings-eacl.78 [ MASK ] insertion: a robust method for anti-adversarial attacks . In Findings of the Association for Computational Linguistics: EACL 20...

  15. [24]

    Xiaowei Huang, Wenjie Ruan, Wei Huang, Gaojie Jin, Yi Dong, Changshun Wu, Saddek Bensalem, Ronghui Mu, Yi Qi, Xingyu Zhao, et al. 2024. A survey of safety and trustworthiness of large language models through the lens of verification and validation. Artificial Intelligence Revi...

  16. [25]

    Takashi Ishida, Ikko Yamane, Tomoya Sakai, Gang Niu, and Masashi Sugiyama. 2020. Do we need zero training loss after achieving zero training error? In Proceedings of the 37th International Conference on Machine Learning, ICML'20. JMLR.org

  17. [26]

    Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018. https://doi.org/10.18653/v1/N18-1170 Adversarial example generation with syntactically controlled paraphrase networks . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association ...

  18. [27]

    Jiabao Ji, Bairu Hou, Zhen Zhang, Guanhua Zhang, Wenqi Fan, Qing Li, Yang Zhang, Gaowen Liu, Sijia Liu, and Shiyu Chang. 2024. https://doi.org/10.18653/v1/2024.naacl-short.23 Advancing the robustness of large language models through self-denoised smoothing . In Proceedings of ...

  19. [28]

    Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020. Is bert really robust? a strong baseline for natural language attack on text classification and entailment. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 8018--8025

  20. [29]

    Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018. Scitail: A textual entailment dataset from science question answering. In Proceedings of the AAAI conference on artificial intelligence, volume 32

  21. [30]

    Deokjae Lee, Seungyong Moon, Junhyeok Lee, and Hyun Oh Song. 2022. Query-efficient and scalable black-box adversarial attacks on discrete sequential data via bayesian optimization. In International Conference on Machine Learning, pages 12478--12497. PMLR

  22. [31]

    Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, and Bill Dolan. 2020 a . Contextualized perturbation for textual adversarial attack. arXiv preprint arXiv:2009.07502

  23. [32]

    Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, and Bill Dolan. 2021 a . https://doi.org/10.18653/v1/2021.naacl-main.400 Contextualized perturbation for textual adversarial attack . In Proceedings of the 2021 Conference of the North American Chapte...

  24. [33]

    Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. 2018. https://api.semanticscholar.org/CorpusID:54815878 Textbugger: Generating adversarial text against real-world applications . ArXiv, abs/1812.05271

  25. [34]

    Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020 b . https://doi.org/10.18653/v1/2020.emnlp-main.500 BERT - ATTACK : Adversarial attack against BERT using BERT . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (E...

  26. [35]

    Linyang Li and Xipeng Qiu. 2021. Token-aware virtual adversarial training in natural language understanding. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 8410--8418

  27. [36]

    Zongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li, Xiaoqing Zheng, Qi Zhang, Kai-Wei Chang, and Cho-Jui Hsieh. 2021 b . https://doi.org/10.18653/v1/2021.emnlp-main.251 Searching for an effective defender: Benchmarking defense against adversarial word substitution . In Proceeding...

  28. [37]

    Zongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li, Xiaoqing Zheng, Qi Zhang, Kai-Wei Chang, and Cho-Jui Hsieh. 2021 c . Searching for an effective defender: Benchmarking defense against adversarial word substitution. arXiv preprint arXiv:2108.12777

  29. [38]

    Aiwei Liu, Honghai Yu, Xuming Hu, Shu ' ang Li, Li Lin, Fukun Ma, Yawen Yang, and Lijie Wen. 2022 a . https://doi.org/10.18653/v1/2022.emnlp-main.522 Character-level white-box adversarial attacks against transformers via attachable subwords substitution . In Proceedings of the...

  30. [39]

    Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9):1--35

  31. [40]

    Qin Liu, Rui Zheng, Bao Rong, Jingyi Liu, Zhihua Liu, Zhanzhan Cheng, Liang Qiao, Tao Gui, Qi Zhang, and Xuan-Jing Huang. 2022 b . Flooding-x: Improving bert’s resistance to adversarial attacks via loss-restricted fine-tuning. In Proceedings of the 60th Annual Meeting of the A...

  32. [41]

    Qin Liu, Rui Zheng, Bao Rong, Jingyi Liu, ZhiHua Liu, Zhanzhan Cheng, Liang Qiao, Tao Gui, Qi Zhang, and Xuanjing Huang. 2022 c . https://doi.org/10.18653/v1/2022.acl-long.386 Flooding- X : Improving BERT ' s resistance to adversarial attacks via loss-restricted fine-tuning . ...

  33. [42]

    Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019 a . Multi-task deep neural networks for natural language understanding. arXiv preprint arXiv:1901.11504

  34. [43]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 b . https://api.semanticscholar.org/CorpusID:198953378 Roberta: A robustly optimized bert pretraining approach . ArXiv, abs/1907.11692

  35. [44]

    Ning Lu, Shengcai Liu, Zhirui Zhang, Qi Wang, Haifeng Liu, and Ke Tang. 2024. Less is more: Understanding word-level textual adversarial attack via n-gram frequency descend. In 2024 IEEE Conference on Artificial Intelligence (CAI), pages 823--830. IEEE

  36. [45]

    Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. https://openreview.net/forum?id=rJzIBfZAb Towards deep learning models resistant to adversarial attacks . In International Conference on Learning Representations

  37. [46]

    Rishabh Maheshwary, Saket Maheshwary, and Vikram Pudi. 2021. Generating natural language attacks in a hard label black box setting. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 13525--13533

  38. [47]

    Raha Moraffah, Shubh Khandelwal, Amrita Bhattacharjee, and Huan Liu. 2024. https://doi.org/10.1007/978-981-97-2262-4_6 Adversarial text purification: A large language model approach for defense . In Advances in Knowledge Discovery and Data Mining - 28th Pacific-Asia Conference...

  39. [48]

    John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020. Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Sy...

  40. [49]

    OpenAI. 2023. https://chat.openai.com.chat

  41. [50]

    Lin Pan, Chung-Wei Hang, Avirup Sil, and Saloni Potdar. 2022. Improved text classification via contrastive adversarial training. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 11130--11138

  42. [51]

    Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of the ACL

  43. [52]

    Francesco Periti, Haim Dubossarsky, and Nina Tahmasebi. 2024. (chat) gpt v bert: Dawn of justice for semantic change detection. arXiv preprint arXiv:2401.14040

  44. [53]

    Danish Pruthi, Bhuwan Dhingra, and Zachary C. Lipton. 2019. https://doi.org/10.18653/v1/P19-1561 Combating adversarial misspellings with robust word recognition . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5582--5591, Flor...

  45. [54]

    Vyas Raina and Mark Gales. 2023. Sample attackability in natural language adversarial attacks. arXiv preprint arXiv:2306.12043

  46. [55]

    Vyas Raina, Samson Tan, Volkan Cevher, Aditya Rawal, Sheng Zha, and George Karypis. 2024. https://aclanthology.org/2024.acl-long.137 Extreme miscalibration and the illusion of adversarial robustness . In Proceedings of the 62nd Annual Meeting of the Association for Computation...

  47. [56]

    Sudhanshu Ranjan, Chung-En Sun, Linbo Liu, and Tsui-Wei Weng. 2023. https://openreview.net/forum?id=NIeCTX8prp Fooling GPT with adversarial in-context examples for text classification . In R0-FoMo:Robustness of Few-shot and Zero-shot Learning in Large Foundation Models

  48. [57]

    Steve Rathje, Dan-Mircea Mirea, Ilia Sucholutsky, Raja Marjieh, Claire E Robertson, and Jay J Van Bavel. 2024. Gpt is an effective tool for multilingual psychological text analysis. Proceedings of the National Academy of Sciences, 121(34):e2308950121

  49. [58]

    Qibing Ren, Liangliang Shi, Lanjun Wang, and Junchi Yan. 2022. https://openreview.net/forum?id=VdYTmPf6BZ- Adversarial robustness via adaptive label smoothing

  50. [59]

    Shuhuai Ren, Yihe Deng, Kun He, and Wanxiang Che. 2019. Generating natural language adversarial examples through probability weighted word saliency. In Proceedings of the 57th annual meeting of the association for computational linguistics, pages 1085--1097

  51. [60]

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018. Semantically equivalent adversarial rules for debugging nlp models. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (volume 1: long papers), pages 856--865

  52. [61]

    Elias Abad Rocamora, Yongtao Wu, Fanghui Liu, Grigorios G Chrysos, and Volkan Cevher. 2024. Revisiting character-level adversarial attacks. arXiv preprint arXiv:2405.04346

  53. [62]

    Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. 2019. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728

  54. [63]

    Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael Abu-Ghazaleh. 2023. Survey of vulnerabilities in large language models revealed by adversarial attacks. arXiv preprint arXiv:2310.10844

  55. [64]

    Chenglei Si, Zhengyan Zhang, Fanchao Qi, Zhiyuan Liu, Yasheng Wang, Qun Liu, and Maosong Sun. 2021. https://doi.org/10.18653/v1/2021.findings-acl.137 Better robustness by more coverage: Adversarial and mixup data augmentation for robust finetuning . In Findings of the Associat...

  56. [65]

    Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language...

  57. [66]

    Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818--2826

  58. [67]

    Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2015. Rethinking the inception architecture for computer vision. corr abs/1512.00567 (2015)

  59. [68]

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://doi.org/10.18653/v1/N19-1421 C ommonsense QA : A question answering challenge targeting commonsense knowledge . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associa...

  60. [69]

    A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems

  61. [70]

    Hetvi Waghela, Sneha Rakshit, and Jaydip Sen. 2024. A modified word saliency-based adversarial attack on text classification models. arXiv preprint arXiv:2403.11297

  62. [71]

    Boxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan, Ruoxi Jia, Bo Li, and Jingjing Liu. 2021. Infobert: Improving robustness of language models from an information theoretic perspective. In International Conference on Learning Representations

  63. [72]

    Jia Wang, Min Gao, Zongwei Wang, Chenghua Lin, Wei Zhou, and Junhao Wen. 2022 a . Ada: Adversarial learning based data augmentation for malicious users detection. Applied Soft Computing

  64. [73]

    Jiayi Wang, Rongzhou Bao, Zhuosheng Zhang, and Hai Zhao. 2022 b . Rethinking textual adversarial defense for pre-trained language models. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:2526--2540

  65. [74]

    Zhaoyang Wang, Zhiyue Liu, Xiaopeng Zheng, Qinliang Su, and Jiahai Wang. 2023. Rmlm: A flexible defense framework for proactively mitigating word-level adversarial attacks. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Lo...

  66. [75]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R \'e mi Louf, Morgan Funtowicz, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical met...

  67. [76]

    Yue Xu and Wenjie Wang. 2024. Linkprompt: Natural and universal adversarial attacks on prompt-based language models. arXiv preprint arXiv:2403.16432

  68. [77]

    Kexin Yang, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Xiangpeng Wei, Zhengyuan Liu, and Jun Xie. 2023 a . https://doi.org/10.18653/v1/2023.acl-long.28 Fantastic expressions and where to find them: C hinese simile generation with multiple constraints . In Proceedings of the 61s...

  69. [78]

    Yahan Yang, Soham Dan, Dan Roth, and Insup Lee. 2023 b . https://doi.org/10.18653/v1/2023.acl-short.58 In and out-of-domain text adversarial robustness via label smoothing . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: S...

  70. [79]

    Yuting Yang, Pei Huang, Juan Cao, Jintao Li, Yun Lin, and Feifei Ma. 2024. A prompt-based approach to adversarial example generation and robustness enhancement. Frontiers of Computer Science, 18(4):184318

  71. [80]

    Jin Yong Yoo and Yanjun Qi. 2021. https://doi.org/10.18653/v1/2021.findings-emnlp.81 Towards improving adversarial training of NLP models . In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 945--956, Punta Cana, Dominican Republic. Association for...

  72. [81]

    Yuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu, Meng Zhang, Qun Liu, and Maosong Sun. 2020. https://doi.org/10.18653/v1/2020.acl-main.540 Word-level textual adversarial attacking as combinatorial optimization . In Proceedings of the 58th Annual Meeting of the Association fo...

  73. [82]

    Guoyang Zeng, Fanchao Qi, Qianrui Zhou, Tingji Zhang, Zixian Ma, Bairu Hou, Yuan Zang, Zhiyuan Liu, and Maosong Sun. 2021. https://doi.org/10.18653/v1/2021.acl-demo.43 O pen A ttack: An open-source textual adversarial attack toolkit . In Proceedings of the 59th Annual Meeting ...

  74. [83]

    Jiehang Zeng, Jianhan Xu, Xiaoqing Zheng, and Xuanjing Huang. 2023. Certified robustness to text adversarial attacks by randomized [mask]. Computational Linguistics, 49(2):395--427

  75. [84]

    Pengwei Zhan, Jing Yang, He Wang, Chao Zheng, Xiao Huang, and Liming Wang. 2023. https://doi.org/10.18653/v1/2023.findings-acl.500 Similarizing the influence of words with contrastive learning to defend word-level adversarial text attack . In Findings of the Association for Co...

  76. [85]

    Zeliang Zhang, Wei Yao, Susan Liang, and Chenliang Xu. 2024. https://aclanthology.org/2024.findings-eacl.83 Random smooth-based certified defense against text adversarial attack . In Findings of the Association for Computational Linguistics: EACL 2024, pages 1251--1265, St. Ju...

  77. [86]

    Jiahao Zhao, Wenji Mao, and Daniel Dajun Zeng. 2024. Disentangled text representation learning with information-theoretic perspective for adversarial robustness. IEEE/ACM Transactions on Audio, Speech, and Language Processing

  78. [87]

    Zhengli Zhao, Dheeru Dua, and Sameer Singh. 2018. https://openreview.net/forum?id=H1BLjgZCb Generating natural adversarial examples . In International Conference on Learning Representations

  79. [88]

    Rui Zheng, Rong Bao, Qin Liu, Tao Gui, Qi Zhang, Xuanjing Huang, Rui Xie, and Wei Wu. 2022. https://aclanthology.org/2022.coling-1.253 P lug AT : A plug and play module to defend against textual adversarial attack . In Proceedings of the 29th International Conference on Comput...

  80. [90]

    Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, and Dacheng Tao. 2023 b . https://arxiv.org/abs/2302.10198 Can chatgpt understand too? a comparative study on chatgpt and fine-tuned bert . Preprint, arXiv:2302.10198

  81. [91]

    Yi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-Wei Chang, and Xuanjing Huang. 2021 a . https://doi.org/10.18653/v1/2021.acl-long.426 Defense against synonym substitution-based adversarial attacks via D irichlet neighborhood ensemble . In Proceedings of the 59th Annual Meeting of ...

  82. [92]

    Yi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-Wei Chang, and Xuanjing Huang. 2021 b . Defense against synonym substitution-based adversarial attacks via dirichlet neighborhood ensemble. In ACL

  83. [93]

    Bin Zhu and Yanghui Rao. 2023. Exploring robust overfitting for pre-trained language models. In Findings of the Association for Computational Linguistics: ACL 2023, pages 5506--5522

  84. [94]

    Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2020. https://openreview.net/forum?id=BygzbyHFvB Freelb: Enhanced adversarial training for natural language understanding . In International Conference on Learning Representations

  85. [95]

    Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Yue Zhang, Neil Zhenqiang Gong, et al. 2023. Promptbench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv preprint arXiv:2306.04528

  86. [96]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  87. [97]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools