Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A formal definition and three-part taxonomy organize the scattered field of LLM robustness.

desk verdict A useful organizational survey with a real repository, but the formal definition at Eq. (1) is ill-posed and should be fixed or dropped before this is referee-ready. read the letter →

arxiv 2506.11111 v2 pith:QTDSRWLC submitted 2025-06-08 cs.CL cs.AI

classification cs.CLcs.AI
keywords LargeLanguageModelsRobustnessAdversarialOut-of-DistributionHallucinationEvaluationParameter-EfficientFine-TuningSurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey paper argues that robustness for large language models is a distinct research field deserving its own terminology, and it sets out to provide that terminology. It proposes a formal definition of LLM robustness as a min–max objective over perturbations, and it organizes the existing literature into three branches: adversarial robustness (noise, long context, and toxic prompts), out-of-distribution robustness (detection, tuning, hallucination), and robustness evaluation (datasets, metrics, benchmarks). The paper documents its paper-collection protocol and provides a searchable GitHub repository, so the contribution is as much an index of the field as an argument about its shape. The intended payoff is that researchers can locate their work in a shared topology, see how ML robustness and LLM robustness differ, and use the paper's future-directions table as a roadmap.

What carries the argument

The load-bearing object is Eq. (1), a min–max formulation that treats robustness as the worst-case loss over a set $\Delta$ of perturbation operations, with the original loss, the perturbed loss, and a distance term $d(\cdot,\cdot)$ (e.g., KL divergence) balanced by hyperparameters $\alpha$ and $\beta$. The formula is intended to cover three facets at once — performance on clean inputs, reliability on perturbed inputs, and consistency between the two outputs — and the paper's taxonomy in Fig. 2(b) is derived from it: adversarial robustness corresponds to perturbations $\epsilon$ such as $X'=X+\delta$, attack prompts, or long contexts; OOD robustness corresponds to distribution shifts with a bounded distance $\eta$ and an explicit "Not Known" refusal option; robustness evaluation supplies datasets, metrics, and benchmarks to measure the objective. This machinery carries the survey because every section assignment and every collected paper is justified by which part of Eq. (1) it addresses.

What would settle it

If one collected a random sample of robustness papers published in the past year and found a positively reviewed work that cannot be placed in any of the three branches, or that falls in two branches without double counting, the taxonomy's claim to be a comprehensive partition would fail. A simpler check: instantiate Eq. (1) literally, repairing the unbalanced parentheses and specifying the distance function $d(\cdot,\cdot)$, and show whether the 'reliability' term $L(\mathrm{LLM}(X'),Y')$ for $Y'\neq Y$ is even well-defined as a loss.

Watch

Extended reading notes

Core claim

The paper's central claim is that LLM robustness should be understood as a model's ability to maintain performance, consistency, and reliability across prompt variations, and that this ability can be captured by a single formal objective: $$\mathrm{Eval}(\$\theta$)=\arg\min_\$\theta$ \max_{\epsilon\in\$\Delta$} \big[ L(\mathrm{LLM}(X),Y) + \$\alpha$ L(\mathrm{LLM}(X'),Y') + \$\beta$\, d(L(\mathrm{LLM}(X))\parallel L(\mathrm{LLM}(X'))) \big],$$ where $X',Y'$ are perturbed data, $\Delta$ is a set of perturbation operations, and hyperparameters $\alpha,\beta$ trade off performance, consistency, and reliability. On the basis of this definition and a comparison with ML robustness, the paper proposes a three-part topology — adversarial robustness, OOD robustness, and robustness evaluation — and reviews representative works in each, alongside datasets, benchmarks, and future directions. The paper also asserts it is the first survey devoted specifically to LLM robustness rather than treating it as a subsection of a general LLM survey.

Load-bearing premise

The survey's organization rests on the assumption that robustness can be captured by a single min–max formula with two balancing weights, and that the three-way split of the literature follows naturally from that formula.

Editorial extensions

If this is right

  • Researchers gain a shared vocabulary: 'noise prompt,' 'noise decoding,' 'OOD detection,' 'PEFT methods,' and 'hallucination' become named branches of a single robustness topology rather than separate subfields.
  • The comparison with ML robustness (input, tuning, output, application) gives a checklist for where LLM robustness research is needed, such as prompt quality, parameter-efficient tuning, and knowledge updating.
  • The formal definition implies that an LLM that is robust in the paper's sense must simultaneously be good on clean inputs, stable under perturbation, and consistent across paraphrases — three properties that existing benchmarks often measure separately.
  • The human-in-the-loop framework (annotator, expert, red team, evaluator) provides a concrete process for continuously finding and patching robustness failures, and the future-directions table offers a chronological roadmap to 2029.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the taxonomy is a claim about how the literature clusters, not a theorem; a different perturbation typology (for example, one organized by attack surface rather than by input stage) would produce a different survey with equal plausibility.
  • Editorial inference: Eq. (1) could be turned into a practical evaluation recipe, because fixing $\alpha$ and $\beta$ and sampling perturbations from $\Delta$ defines a family of robustness scores that the field does not yet standardize.
  • Editorial inference: the same three-branch structure could be applied to multimodal and agentic LLMs, where the 'prompt' becomes a trajectory of observations and actions; whether the topology survives that extension is an open question.
  • Editorial inference: the paper's own comparison with ML robustness implies that 'Not Known' refusal is a form of robustness, which points toward a testable extension — systems that learn to abstain under distribution shift may score higher on the paper's objective than systems that always answer.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper is a survey of robustness in large language models. It claims to be the first comprehensive survey focused on LLM robustness, proposes a formal definition of LLM robustness in Eq. (1), and organizes the literature into three parts: adversarial robustness, OOD robustness, and robustness evaluation. It also provides a companion GitHub repository and discusses future directions including human-in-the-loop evaluation.

Significance. If the survey's indexing and taxonomy prove reliable, it would give the community a useful entry point: the three-part division by perturbation type is a reasonable organization, and the companion repository is a practical asset. The discussion of human roles and causal-inference connections is a forward-looking addition. However, the formal foundation is currently not rigorous, and several cited summaries contain errors; because the paper's value as a survey depends on accurate indexing, these issues need to be fixed.

major comments (4)
  1. [Section 1, Eq. (1)] The equation cannot serve as a formal definition of LLM robustness. The expression Eval(theta) is an argmin-max training objective, not a property or metric that characterizes whether a given model is robust. The perturbation set Delta is never linked to a distribution over (X',Y'), so the maximization is not well defined. The distance d(.,.) is left unspecified, the KL term contains unbalanced parentheses and is written with one argument, and the 'reliability' condition (Y' != Y) does not define a loss. Because the paper claims that the three-branch taxonomy is based on this definition, this is a load-bearing issue. The authors should either replace Eq. (1) with a well-typed definition of robustness as a model property, or explicitly present the taxonomy as an organizational choice based on perturbation types, as the abstract already does.
  2. [Section 1, contributions bullet] The claim 'we are the first to concentrate on the LLM Robustness' is contradicted by the paper's own references. Reference [179] (Wang et al., 'On the robustness of ChatGPT: An adversarial and out-of-distribution perspective') is a robustness-focused study of LLMs, and reference [213] (Yuan et al., 'Revisiting out-of-distribution robustness in NLP: Benchmarks, analysis, and LLMs evaluations') explicitly surveys OOD robustness in LLMs. The novelty claim should be qualified or removed, and the related-work comparison should engage with these existing surveys directly.
  3. [Section 5.2.2, reference [37]] The text states that 'Esiobu et al. [37] proposes ROBUST, the first benchmark for evaluating open information extraction models in real-world scenarios,' but reference [37] is titled 'ROBBIE: Robust bias evaluation of large generative language models' and concerns bias evaluation, not open information extraction. The same benchmark is also listed in Table 6 as 'ROBUST [37]' under 'Open domain generalization,' which is inconsistent. This is a substantive misattribution in a survey whose contribution is reliable indexing.
  4. [Section 3.2 and Table 3, reference [167]] The text attributes ALiBi to 'Sun et al. [167]' and says 'Their proposed ALiBi [167]' has been shown to outperform other position embedding methods. Reference [167] is the paper 'A Length-Extrapolatable Transformer' by Sun et al.; ALiBi is by Press et al. [137]. This type of attribution error is material in a survey-as-index. The authors should systematically audit all inline citations and tables against the reference list.
minor comments (5)
  1. [References [104] and [105]] References [104] and [105] are the same paper, 'Lost in the Middle: How Language Models Use Long Contexts,' cited with different volume numbers for the same TACL article; they should be merged, and Section 3.2 should cite a single entry.
  2. [Throughout] There are numerous typos and grammatical errors, including 'wild-range' for 'wide-range,' 'unexpeted' for 'unexpected,' 'perturbated' for 'perturbed,' 'disturber' for 'disturb,' 'finishi' and 'usenormous' in Section 1, and 'fla' for 'flag' in Section 2.2.2. A careful proofreading pass is needed.
  3. [Section 2.2.1, Eq. (2)] Equation (2) writes epsilon in {X' = X + delta, X' = Attack(X), LongContext}, mixing a scalar perturbation symbol with input-transformation conditions; the set-membership notation should be clarified.
  4. [Section 4.4] The sentence 'there also exist other related work [34, 47, 47, 68, 98, 150, 171]' contains a duplicated reference [47] and should be de-duplicated.
  5. [Section 5.2.2] The phrase 'he refined robustness metrics' has an unclear antecedent and appears to be a pronoun error; it should be 'they' or a specific author name should be given.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's definition and taxonomy are stipulated, not derived from the paper's own outputs.

full rationale

This is a survey paper, and its central contributions are a formal definition of LLM robustness (Eq. 1), a three-part taxonomy (Adversarial, OOD, Evaluation), a paper-collection protocol, and a companion GitHub repository. None of these involve fitting a parameter and then predicting a closely related quantity, nor does any claimed result reduce by construction to an input. The formal definition in Eq. (1) is asserted rather than derived, and the taxonomy in Section 2.2 is explicitly justified by 'the comparison between ML robustness and LLM robustness' plus prior surveys [18, 160, 226] and the keyword statistics in Fig. 2(a), not by an inference from Eq. (1). The paper's claim to be 'the first to concentrate on the LLM Robustness' is a novelty assertion, not a load-bearing epistemic derivation, and the companion repository is an organizational asset rather than a circular support. The weaknesses noted by the reader, such as the unbalanced parentheses in Eq. (1), the unspecified distance d(·,·), and the loose mapping from the three bullets to the taxonomy, are correctness and rigor concerns about how well the definition is formalized; they are not instances of a result being equivalent to its own inputs. There is no self-citation chain invoked to forbid alternatives and no renamed empirical pattern presented as a derivation. Accordingly, the circularity score is 0.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The survey introduces no fitted parameters or invented entities. Its framework rests on an unproven formal definition (Eq. 1) and on an asserted taxonomy of the literature, plus a keyword-based collection protocol that is not fully reproducible.

free parameters (2)
  • alpha (Eq. 1)
    Hyperparameter weighting the consistency term in the proposed robustness objective; no value is specified, so the definition is not operationalized.
  • beta (Eq. 1)
    Hyperparameter weighting the reliability term in Eq. (1); no value is specified.
assumptions (3)
  • ad hoc to paper The min-max objective in Eq. (1) defines LLM robustness.
    No derivation or empirical validation is provided; it is a variation of standard adversarial training objectives.
  • domain assumption LLM robustness can be partitioned into adversarial robustness, OOD robustness, and evaluation.
    Section 2.2 asserts this topology without a systematic justification or comparison to alternative taxonomies.
  • domain assumption The keyword-based search over selected venues yields a representative picture of the field.
    Section 2.1 describes the protocol, but exact queries, dates, and inclusion criteria are not fully specified, and the filtering relies on manual title/abstract review.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions." pith.science (2026). https://pith.science/paper/QTDSRWLC

@misc{pith2026250611111,
  author       = {Pith},
  title        = {Pith review of: Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QTDSRWLC}},
  note         = {Machine review of arXiv:2506.11111}
}
read the original abstract

Large Language Models (LLMs) have gained enormous attention in recent years due to their capability of understanding and generating natural languages. With the rapid development and wild-range applications (e.g., Agents, Embodied Intelligence), the robustness of LLMs has received increased attention. As the core brain of many AI applications, the robustness of LLMs requires that models should not only generate consistent contents, but also ensure the correctness and stability of generated content when dealing with unexpeted application scenarios (e.g., toxic prompts, limited noise domain data, outof-distribution (OOD) applications, etc). In this survey paper, we conduct a thorough review of the robustness of LLMs, aiming to provide a comprehensive terminology of concepts and methods around this field and facilitate the community. Specifically, we first give a formal definition of LLM robustness and present the collection protocol of this survey paper. Then, based on the types of perturbated inputs, we organize this survey from the following perspectives: 1) Adversarial Robustness: tackling the problem that prompts are manipulated intentionally, such as noise prompts, long context, data attack, etc; 2) OOD Robustness: dealing with the unexpected real-world application scenarios, such as OOD detection, zero-shot transferring, hallucinations, etc; 3) Evaluation of Robustness: summarizing the new evaluation datasets, metrics, and tools for verifying the robustness of LLMs. After reviewing the representative work from each perspective, we discuss and highlight future opportunities and research directions in this field. Meanwhile, we also organize related works and provide an easy-to-search project (https://github.com/zhangkunzk/Awesome-LLM-Robustness-papers) to support the community.

Figures

Figures reproduced from arXiv: 2506.11111 by the authors.

Figure 1
Figure 1. The pipeline of training and applying Large Language Models (LLMs). [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. (A) The statistics of the topic and keyword distributions from the collected papers. (B) The topology of our survey. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. An example of code agents processing various requirements and generate code solution. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: An example of car agents conducting incorrect decisions when dealing with Out-Of-Distribution scenarios. [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: An example of the human-in-the-loop framework for continually improving the robustness of LLMs. [PITH_FULL_IMAGE:figures/full_fig_p020_5.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    CoT probe-time gains arise primarily from lexical activation and short-range token co-occurrence rather than sentence-level logical derivation.

Reference graph

Works this paper leans on

246 extracted references · 37 canonical work pages · cited by 1 Pith paper

  1. [22]

    Xuanting Chen, Junjie Ye, Can Zu, Nuo Xu, Rui Zheng, Minlong Peng, Jie Zhou, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023. How robust is gpt-3.5 to predecessors? a comprehensive study on language understanding tasks. arXiv preprint arXiv:2303.00293 (2023)

  2. [41]

    Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2024. Bias and fairness in large language models: A survey. Computational Linguistics (2024), 1–79

  3. [213]

    Lifan Yuan, Yangyi Chen, Ganqu Cui, Hongcheng Gao, Fangyuan Zou, Xingyi Cheng, Heng Ji, Zhiyuan Liu, and Maosong Sun. 2023. Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and LLMs evaluations.Advances in Neural Information Processing Systems 36 (2023), 58478–58507

  4. [104]

    Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 11 (2024), 157–173

  5. [105]

    Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

    Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 12 (2024), 157–173

  6. [179]

    Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, et al

  7. [37]

    David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, and Eric Smith. 2023. ROBBIE: Robust bias evaluation of large generative language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 3764–3814

  8. [167]

    Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei. 2023. A Length-Extrapolatable Transformer. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 14590–14604

  9. [137]

    Ofir Press, Noah Smith, and Mike Lewis. 2022. Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. In International Conference on Learning Representations

Show all 246 references
  1. [1]

    Jameel Abdul Samadh, Mohammad Hanan Gani, Noor Hussein, Muhammad Uzair Khattak, Muhammad Muzammal Naseer, Fahad Shahbaz Khan, and Salman H Khan. 2024. Align your prompts: Test-time prompting with distribution alignment for zero-shot generalization. Advances in Neural Informati...

  2. [2]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  3. [3]

    Sravanti Addepalli, Ashish Ramayee Asokan, Lakshay Sharma, and R Venkatesh Babu. 2024. Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 23922–23932

  4. [4]

    Dyah Adila, Changho Shin, Linrong Cai, and Frederic Sala. 2024. Zero-Shot Robustification of Zero-Shot Models. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=fCeUoDr9Tq

  5. [5]

    Abhishek Aich, Calvin-Khang Ta, Akash Gupta, Chengyu Song, Srikanth Krishnamurthy, Salman Asif, and Amit Roy-Chowdhury. 2022. Gama: Generative adversarial multi-object scene attacks. Advances in Neural Information Processing Systems 35 (2022), 36914–36930

  6. [6]

    Shengnan An, Zeqi Lin, Qiang Fu, Bei Chen, Nanning Zheng, Jian-Guang Lou, and Dongmei Zhang. 2023. How Do In-Context Examples Affect Compositional Generalization?. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers...

  7. [7]

    Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng, and Jian-Guang Lou. 2024. Make Your LLM Fully Utilize the Context. arXiv preprint arXiv:2404.16811 (2024)

  8. [8]

    Daman Arora, Himanshu Singh, et al. 2023. Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 7527–7543

  9. [9]

    Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023. Self-rag: Learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations

  10. [10]

    Katherine Atwell, Mert Inan, Anthony B Sicilia, and Malihe Alikhani. 2024. Combining Discourse Coherence with Large Language Models for More Inclusive, Equitable, and Robust Task-Oriented Dialogue. In Proceedings of the 2024 Joint International Conference on Computational Ling...

  11. [11]

    Reza Averly and Wei-Lun Chao. 2023. Unified out-of-distribution detection: A model-specific perspective. InProceedings of the IEEE/CVF International Conference on Computer Vision . 1453–1463

  12. [12]

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073 (2022)

  13. [13]

    Sourya Basu, Pulkit Katdare, Prasanna Sattigeri, Vijil Chenthamarakshan, Katherine Driggs-Campbell, Payel Das, and Lav R Varshney

  14. [14]

    Francisco Bellas, Sara Guerreiro-Santalla, Martin Naya, and Richard J Duro. 2023. AI curriculum for European high schools: An embedded intelligence approach. International Journal of Artificial Intelligence in Education 33, 2 (2023), 399–426

  15. [15]

    Adam Bouyamourn. 2023. Why LLMs Hallucinate, and How to Get (Evidential) Closure: Perceptual, Intensional, and Extensional Learning for Faithful Natural Language Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 3181–3193

  16. [16]

    Shreyas Bhat Brahmavar, Ashwin Srinivasan, Tirtharaj Dash, Sowmya Ramaswamy Krishnan, Lovekesh Vig, Arijit Roy, and Raviprasad Aduri. 2024. Generating Novel Leads for Drug Discovery using LLMs with Logical Feedback. In Proceedings of the AAAI Conference on Artificial Intellige...

  17. [17]

    Houssem Ben Braiek and Foutse Khomh. 2024. Machine Learning Robustness: A Primer. arXiv preprint arXiv:2404.00897 (2024)

  18. [18]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–45

  19. [19]

    Huimin Chen, Chengyu Wang, Yanhao Wang, Cen Chen, and Yinggui Wang. 2024. TaiChi: Improving the Robustness of NLP Models by Seeking Common Ground While Reserving Differences. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resou...

  20. [20]

    Hao Chen, Jindong Wang, Ankit Shah, Ran Tao, Hongxin Wei, Xing Xie, Masashi Sugiyama, and Bhiksha Raj. 2024. Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks. In The Twelfth International Conference on Learning Representations . https://openrevi...

  21. [21]

    Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. [n. d.]. Benchmarking large language models in retrieval-augmented generation. arXiv preprint arXiv:2309.01431 ([n. d.]). Proc. ACM Meas. Anal. Comput. Syst., Vol. 37, No. 4, Article 111. Publication date: August 2024. 111:24 •...

  22. [23]

    Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. 2024. LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models. In The Twelfth International Conference on Learning Representations

  23. [24]

    Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, and Ajay Divakaran. 2024. DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition ...

  24. [25]

    Zining Chen, Weiqiu Wang, Zhicheng Zhao, Fei Su, Aidong Men, and Hongying Meng. 2024. PracticalDG: Perturbation Distillation on Vision-Language Models for Hybrid Domain Generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)....

  25. [26]

    Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023. Adapting Language Models to Compress Contexts. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 3829–3846

  26. [27]

    Zhixiang Chi, Li Gu, Tao Zhong, Huan Liu, YUANHAO YU, Konstantinos N Plataniotis, and Yang Wang. 2024. Adapting to Distribution Shift by Visual Domain Prompt Generation. In The Twelfth International Conference on Learning Representations . https://openreview. net/forum?id=sSaN4gxuEf

  27. [28]

    Glass, and Pengcheng He

    Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, and Pengcheng He. 2024. DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models. In The Twelfth International Conference on Learning Representations . https: //openreview.net/forum?id...

  28. [29]

    Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. arXiv preprint arXiv:1809.05053 (2018)

  29. [30]

    Siddhartha Datta. 2022. Learn2weight: Parameter adaptation against similar-domain adversarial attacks. arXiv preprint arXiv:2205.07315 (2022)

  30. [31]

    Hillary Dawkins, Isar Nejadgholi, Daniel Gillis, and Judi McCuaig. 2024. Projective Methods for Mitigating Gender Bias in Pre-trained Language Models. arXiv preprint arXiv:2403.18803 (2024)

  31. [32]

    Nicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton, Elsbeth Turcan, and Kathleen Mckeown. 2023. Evaluation of African American Language Bias in Natural Language Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 6805–6824

  32. [33]

    Zican Dong, Tianyi Tang, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen. 2023. Bamboo: A comprehensive benchmark for evaluating long text modeling capacities of large language models. arXiv preprint arXiv:2309.13345 (2023)

  33. [34]

    Yingpeng Du, Di Luo, Rui Yan, Xiaopei Wang, Hongzhi Liu, Hengshu Zhu, Yang Song, and Jie Zhang. 2024. Enhancing job recom- mendation through llm-based generative adversarial networks. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 8363–8371

  34. [35]

    Haonan Duan, Adam Dziedzic, Nicolas Papernot, and Franziska Boenisch. 2023. Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models. InThirty-seventh Conference on Neural Information Processing Systems. https://openreview. net/forum?id=u6Xv3FuF8N

  35. [36]

    Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan. 2022. A survey of embodied ai: From simulators to research tasks. IEEE Transactions on Emerging Topics in Computational Intelligence 6, 2 (2022), 230–244

  36. [38]

    Yu Fei, Yifan Hou, Zeming Chen, and Antoine Bosselut. 2023. Mitigating Label Biases for In-context Learning. InProceedings Of The 61St Annual Meeting Of The Association For Computational Linguistics (Acl 2023): Long Papers, Vol 1 . Assoc Computational Linguistics-Acl, 14014–14031

  37. [39]

    Virginia Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May. 2023. WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long ...

  38. [40]

    Steinunn Rut Friðriksdóttir and Hafsteinn Einarsson. 2024. Gendered Grammar or Ingrained Bias? Exploring Gender Bias in Icelandic Language Models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-CO...

  39. [42]

    Zhibin Gou, Zhihong Shao, Yeyun Gong, yelong shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen. 2024. ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving. InThe Twelfth International Conference on Learning Representations. 1–34. Proc. ACM Meas. Ana...

  40. [43]

    Shreya Goyal, Sumanth Doddapaneni, Mitesh M Khapra, and Balaraman Ravindran. 2023. A survey of adversarial defenses and robustness in nlp. Comput. Surveys 55, 14s (2023), 1–39

  41. [44]

    Sachin Goyal, Ananya Kumar, Sankalp Garg, Zico Kolter, and Aditi Raghunathan. 2023. Finetune like you pretrain: Improved finetuning of zero-shot vision models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 19338–19347

  42. [45]

    Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752 (2023)

  43. [46]

    Xinyan Guan, Yanjiang Liu, Hongyu Lin, Yaojie Lu, Ben He, Xianpei Han, and Le Sun. 2024. Mitigating large language model hallucinations via autonomous knowledge graph-based retrofitting. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 18126–18134

  44. [47]

    Anisha Gunjal, Jihan Yin, and Erhan Bas. 2024. Detecting and preventing hallucinations in large vision language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 18135–18143

  45. [48]

    Prakhar Gupta, Cathy Jiao, Yi-Ting Yeh, Shikib Mehri, Maxine Eskenazi, and Jeffrey P Bigham. 2022. InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pr...

  46. [49]

    Fifty Shades of Bias

    Rishav Hada, Agrima Seth, Harshita Diddee, and Kalika Bali. 2023. “Fifty Shades of Bias”: Normative Ratings of Gender Bias in GPT Generated English Text. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 1862–1876

  47. [50]

    Dongchen Han, Xuran Pan, Yizeng Han, Shiji Song, and Gao Huang. 2023. Flatten transformer: Vision transformer using focused linear attention. In Proceedings of the IEEE/CVF international conference on computer vision . 5961–5971

  48. [51]

    Jinwei Han, Zhiwen Lin, Zhongyisun Sun, Yingguo Gao, Ke Yan, Shouhong Ding, Yuan Gao, and Gui-Song Xia. 2024. Anchor-based Robust Finetuning of Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 26919–26928

  49. [52]

    Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu. 2023. Reasoning with Language Model is Planning with World Model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 8154–8173

  50. [53]

    Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2021. Self-attention attribution: Interpreting information interactions inside transformer. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 12963–12971

  51. [54]

    Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. In Proceedings of the 60th Annual Meeting of the Association for Computational...

  52. [55]

    Zexue He, Yu Wang, An Yan, Yao Liu, Eric Chang, Amilcare Gentili, Julian McAuley, and Chun-nan Hsu. 2023. MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model Evaluation. In Proceedings of the 2023 Conference on Empirical Methods in Natural...

  53. [56]

    Peter Henderson, Eric Mitchell, Christopher Manning, Dan Jurafsky, and Chelsea Finn. 2023. Self-Destructing Models: Increasing the Costs of Harmful Dual Uses of Foundation Models. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society . 287–296

  54. [57]

    Eugenio Herrera-Berg, Tomás Browne, Pablo León-Villagrá, Marc-Lluís Vives, and Cristian Calderon. 2023. Large Language Models are biased to overestimate profoundness. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 9653–9661

  55. [58]

    Cheng-Yu Hsieh, Yung-Sung Chuang, Chun-Liang Li, Zifeng Wang, Long T Le, Abhishek Kumar, James Glass, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, et al. 2024. Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization. arXiv preprint arXiv:...

  56. [59]

    Chengang Hu, Xiao Liu, and Yansong Feng. 2023. DiNeR: A Large Realistic Dataset for Evaluating Compositional Generalization. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 14938–14947

  57. [60]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685 (2021)

  58. [61]

    Zeyi Huang, Andy Zhou, Zijian Ling, Mu Cai, Haohan Wang, and Yong Jae Lee. 2023. A Sentence Speaks a Thousand Images: Domain Generalization through Distilling CLIP with Language Guidance. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 11685–11695

  59. [62]

    Shima Imani, Liang Du, and Harsh Shrivastava. 2023. Mathprompter: Mathematical reasoning using large language models. arXiv preprint arXiv:2303.05398 (2023)

  60. [63]

    Hamish Ivison, Akshita Bhagia, Yizhong Wang, Hannaneh Hajishirzi, and Matthew E Peters. 2023. HINT: Hypernetwork Instruction Tuning for Efficient Zero-and Few-Shot Generalisation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volum...

  61. [64]

    Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang. 2017. Search-based neural structured learning for sequential question answering. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 1821–1831

  62. [65]

    Akshita Jha, Aida Mostafazadeh Davani, Chandan K Reddy, Shachi Dave, Vinodkumar Prabhakaran, and Sunipa Dev. 2023. SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models. In Proceedings of the 61st Annual Meeting Proc. ACM Meas. Anal. Com...

  63. [66]

    Akshita Jha and Chandan K Reddy. 2023. Codeattack: Code-based adversarial attacks for pre-trained programming language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 14892–14900

  64. [67]

    Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088 (2024)

  65. [68]

    Chaoya Jiang, Haiyang Xu, Mengfan Dong, Jiaxing Chen, Wei Ye, Ming Yan, Qinghao Ye, Ji Zhang, Fei Huang, and Shikun Zhang. 2024. Hallucination Augmented Contrastive Learning for Multimodal Large Language Model. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...

  66. [69]

    Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023. Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression. arXiv preprint arXiv:2310.06839 (2023)

  67. [70]

    Shuoran Jiang, Qingcai Chen, Yang Xiang, Youcheng Pan, and Yukang Lin. 2024. Linguistic Rule Induction Improves Adversarial and OOD Robustness in Large Language Models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources a...

  68. [71]

    Xue Jiang, Feng Liu, Zhen Fang, Hong Chen, Tongliang Liu, Feng Zheng, and Bo Han. 2024. Negative Label Guided OOD Detection with Pretrained Vision-Language Models. In The Twelfth International Conference on Learning Representations . https://openreview.net/ forum?id=xUO1HXz4an

  69. [72]

    Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu. 2023. Is ChatGPT a good translator? A preliminary study. arXiv preprint arXiv:2301.08745 1, 10 (2023)

  70. [73]

    Zhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal, LYU Zhiheng, Kevin Blin, Fernando Gonzalez Adauto, Max Kleiman- Weiner, Mrinmaya Sachan, et al. 2023. Cladder: Assessing causal reasoning in language models. In Thirty-seventh conference on neural information proc...

  71. [74]

    Erik Jones, Hamid Palangi, Clarisse Simões Ribeiro, Varun Chandrasekaran, Subhabrata Mukherjee, Arindam Mitra, Ahmed Hassan Awadallah, and Ece Kamar. 2024. Teaching Language Models to Hallucinate Less with Synthetic Tasks. InThe Twelfth International Conference on Learning Rep...

  72. [75]

    Adilbek Karmanov, Dayan Guan, Shijian Lu, Abdulmotaleb El Saddik, and Eric Xing. 2024. Efficient Test-Time Adaptation of Vision- Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14162–14171

  73. [76]

    Aly Kassem, Omar Mahmoud, and Sherif Saad. 2023. Preserving privacy through dememorization: An unlearning technique for mitigating memorization risks in language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 4360–4379

  74. [77]

    Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier. 2023. A survey of reinforcement learning from human feedback. arXiv preprint arXiv:2312.14925 (2023)

  75. [78]

    Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh. 2023. ProPILE: Probing Privacy Leakage in Large Language Models. In Thirty-seventh Conference on Neural Information Processing Systems . https://openreview.net/forum?id= QkLpGxUboF

  76. [79]

    Ching-Yun Ko, Pin-Yu Chen, Payel Das, Yung-Sung Chuang, and Luca Daniel. 2023. On Robustness-Accuracy Characterization of Large Language Models using Synthetic Datasets. (2023)

  77. [80]

    Giorgi Kokaia, Pratyush Sinha, Yutong Jiang, and Nozha Boujemaa. 2023. Writing your own book: A method for going from closed to open book QA to improve robustness and performance of smaller LLMs. arXiv preprint arXiv:2305.11334 (2023)

  78. [81]

    Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. MAWPS: A math word problem repository. In Proceedings of the 2016 conference of the north american chapter of the association for computational linguistics: human language technologies....

  79. [82]

    Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez. 2023. Pretraining language models with human preferences. In International Conference on Machine Learning . PMLR, 17506–17533

  80. [83]

    Mayank Kothyari, Dhruva Dhingra, Sunita Sarawagi, and Soumen Chakrabarti. 2023. CRUSH4SQL: Collective Retrieval Using Schema Hallucination For Text2SQL. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 14054–14066

  81. [84]

    Ananya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma, and Percy Liang. 2022. Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution. In International Conference on Learning Representations . https://openreview.net/ forum?id=UYneFzXSJWh

  82. [85]

    Po-Nien Kung, Fan Yin, Di Wu, Kai-Wei Chang, and Nanyun Peng. 2023. Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 1813–1829

  83. [86]

    Yu-Ju Lan and Nian-Shing Chen. 2024. Teachers’ agency in the era of LLM and generative AI. Educational Technology & Society 27, 1 (2024), I–XVIII. Proc. ACM Meas. Anal. Comput. Syst., Vol. 37, No. 4, Article 111. Publication date: August 2024. Evaluating and Improving Robustne...

  84. [87]

    Yoonho Lee, Annie S Chen, Fahim Tajwar, Ananya Kumar, Huaxiu Yao, Percy Liang, and Chelsea Finn. 2023. Surgical Fine-Tuning Improves Adaptation to Distribution Shifts. In The Eleventh International Conference on Learning Representations . https://openreview. net/forum?id=APuPRxjHvZ

  85. [88]

    Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing. 2024. Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...

  86. [89]

    Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen Mckeown, and William Yang Wang. 2022. SafeText: A Benchmark for Exploring Physical Safety in Language Models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pro...

  87. [90]

    Bingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen, Yuekun Yao, and Najoung Kim. 2023. SLOG: A Structural Generalization Benchmark for Semantic Parsing. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 3213–3232

  88. [91]

    Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2023. HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 6449–6464

  89. [92]

    Juncheng Li, Minghe Gao, Longhui Wei, Siliang Tang, Wenqiao Zhang, Mengze Li, Wei Ji, Qi Tian, Tat-Seng Chua, and Yueting Zhuang. 2023. Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language Models. In Proceedings of the IEEE/CVF International Conference on ...

  90. [93]

    Lin Li, Haoyan Guan, Jianing Qiu, and Michael Spratling. 2024. One prompt word is enough to boost adversarial robustness for pre-trained vision-language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 24408–24419

  91. [94]

    Shanda Li, Chong You, Guru Guruganesh, Joshua Ainslie, Santiago Ontanon, Manzil Zaheer, Sumit Sanghai, Yiming Yang, Sanjiv Kumar, and Srinadh Bhojanapalli. 2024. Functional Interpolation for Relative Positions improves Long Context Transformers. In The Twelfth International Co...

  92. [95]

    Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto. 2022. Large Language Models Can Be Strong Differentially Private Learners. In International Conference on Learning Representations . https://openreview.net/forum?id=bVuP3ltATMz

  93. [96]

    Xiao Li, Wei Zhang, Yining Liu, Zhanhao Hu, Bo Zhang, and Xiaolin Hu. 2024. Language-Driven Anchors for Zero-Shot Adversarial Robustness. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 24686–24695

  94. [97]

    Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023. Contrastive Decoding: Open-ended Text Generation as Optimization. In Proceedings of the 61st Annual Meeting of the Association for Computational L...

  95. [98]

    Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. 2023. Evaluating Object Hallucination in Large Vision-Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 292–305

  96. [99]

    Yufei Li, Zexin Li, Yingfan Gao, and Cong Liu. 2023. White-box multi-objective adversarial attack on dialogue generation. arXiv preprint arXiv:2305.03655 (2023)

  97. [100]

    Jian Liang, Ran He, and Tieniu Tan. 2024. A comprehensive survey on test-time adaptation under distribution shifts. International Journal of Computer Vision (2024), 1–34

  98. [101]

    Victoria Lin, Eli Ben-Michael, and Louis-Philippe Morency. [n. d.]. Optimizing Language Models for Human Preferences is a Causal Inference Problem. In The 40th Conference on Uncertainty in Artificial Intelligence

  99. [102]

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. DeepSeek-V3 Technical Report. arXiv preprint arXiv:2412.19437 (2024)

  100. [103]

    Bo Liu, Li-Ming Zhan, Zexin Lu, Yujie Feng, Lei Xue, and Xiao-Ming Wu. 2024. How Good Are LLMs at Out-of-Distribution Detection?. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). 8211–8222

  101. [106]

    Yiran Liu, Xiaoang Xu, Zhiyi Hou, and Yang Yu. 2024. Causality Based Front-door Defense Against Backdoor Attack on Language Models. In Forty-first International Conference on Machine Learning . https://openreview.net/forum?id=dmHHVcHFdM

  102. [107]

    Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. 2023. Trustworthy LLMs: A survey and guideline for evaluating large language models’ alignment. arXiv preprint arXiv:2308.05374 (2023)

  103. [108]

    Dong Lu, Zhiqiang Wang, Teng Wang, Weili Guan, Hongchang Gao, and Feng Zheng. 2023. Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 102–111. Proc...

  104. [109]

    Enzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du, Tao Jiang, Chao Hong, Shaowei Liu, Weiran He, Enming Yuan, Yuzhi Wang, et al

  105. [110]

    Xinyuan Lu, Liangming Pan, Qian Liu, Preslav Nakov, and Min-Yen Kan. 2023. SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 7787–7813

  106. [111]

    Fan Ma, Xiaojie Jin, Heng Wang, Yuchen Xian, Jiashi Feng, and Yi Yang. 2024. VISTA-LLAMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 13151–13160

  107. [112]

    Rishabh Maheshwary, Saket Maheshwary, and Vikram Pudi. 2021. A context aware approach for generating natural language attacks. In Proceedings of the AAAI conference on artificial intelligence , Vol. 35. 15839–15840

  108. [113]

    Rishabh Maheshwary, Saket Maheshwary, and Vikram Pudi. 2021. Generating natural language attacks in a hard label black box setting. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 13525–13533

  109. [114]

    Potsawee Manakul, Adian Liusie, and Mark Gales. 2023. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 9004–9017

  110. [115]

    Yuren Mao, Yuhang Ge, Yijiang Fan, Wenyi Xu, Yu Mi, Zhonghao Hu, and Yunjun Gao. 2024. A Survey on LoRA of Large Language Models. arXiv preprint arXiv:2407.11046 (2024)

  111. [116]

    Weikang Meng, Yadan Luo, Xin Li, Dongmei Jiang, and Zheng Zhang. 2025. PolaFormer: Polarity-aware Linear Attention for Vision Transformers. In The Thirteenth International Conference on Learning Representations

  112. [117]

    Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2021. A diverse corpus for evaluating and developing English math word problem solvers. arXiv preprint arXiv:2106.15772 (2021)

  113. [118]

    Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large language models: A survey. arXiv preprint arXiv:2402.06196 (2024)

  114. [119]

    Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-Task Generalization via Natural Language Crowdsourcing Instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 3470–3487

  115. [120]

    Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, et al. 2023. Crosslingual Generalization through Multitask Finetuning. In Proceedings of the 61st Annual Meeting of th...

  116. [121]

    Manish Nagireddy, Lamogha Chiazor, Moninder Singh, and Ioana Baldini. 2024. Socialstigmaqa: A benchmark to uncover stigma amplification in generative language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 21454–21462

  117. [122]

    Muzammal Naseer, Ahmad Mahmood, Salman Khan, and Fahad Khan. 2023. Boosting Adversarial Transferability using Dynamic Cues. In The Eleventh International Conference on Learning Representations . https://openreview.net/forum?id=SZynfVLGd5

  118. [123]

    Allen Nie, Yuhui Zhang, Atharva Shailesh Amdekar, Chris Piech, Tatsunori B Hashimoto, and Tobias Gerstenberg. 2024. Moca: Measuring human-language model alignment on causal and moral judgment tasks. Advances in Neural Information Processing Systems 36 (2024)

  119. [124]

    Ardavan Salehi Nobandegani, Kevin da Silva Castanheira, Timothy O’Donnell, and Thomas R Shultz. 2019. On Robustness: An Undervalued Dimension of Human Rationality.. In CogSci. 3327

  120. [125]

    Changdae Oh, Junhyuk So, Hoyoon Byun, YongTaek Lim, Minchul Shin, Jong-June Jeon, and Kyungwoo Song. 2024. Geodesic multi-modal mixup for robust fine-tuning. Advances in Neural Information Processing Systems 36 (2024)

  121. [126]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35...

  122. [127]

    Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th annual meeting of the association for computational linguistics (volume 1: long papers) . 1946–1958

  123. [128]

    Giuseppe Paolo, Jonas Gonzalez-Billandon, and Balázs Kégl. 2024. A call for embodied AI. arXiv preprint arXiv:2402.03824 (2024)

  124. [129]

    Panupong Pasupat and Percy Liang. 2015. Compositional semantic parsing on semi-structured tables. arXiv preprint arXiv:1508.00305 (2015)

  125. [130]

    Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems? arXiv preprint arXiv:2103.07191 (2021)

  126. [131]

    Arkil Patel, Satwik Bhattamishra, Siva Reddy, and Dzmitry Bahdanau. 2023. MAGNIFICo: Evaluating the In-Context Learning Ability of Large Language Models to Generalize to Novel Interpretations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Proce...

  127. [132]

    was it “stated

    Roma Patel and Ellie Pavlick. 2021. “was it “stated” or was it “claimed”?: How linguistic bias affects generative language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 10080–10095

  128. [133]

    Kellin Pelrine, Anne Imouza, Camille Thibault, Meilina Reksoprodjo, Caleb Gupta, Joel Christoph, Jean-François Godbout, and Reihaneh Rabbany. 2023. Towards Reliable Misinformation Mitigation: Generalization, Uncertainty, and GPT-4. In Proceedings of the 2023 Conference on Empi...

  129. [134]

    Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Leon Derczynski, Xingjian Du, Matteo Grella, Kranthi Gv, Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bartłomiej Koptyra, Hay...

  130. [135]

    Luiza Pozzobon, Beyza Ermis, Patrick Lewis, and Sara Hooker. 2023. On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 7595–7609

  131. [136]

    Bodhisattwa Prasad Majumder, Zexue He, and Julian McAuley. 2022. InterFair: Debiasing with Natural Language Feedback for Fair Interpretable Predictions. arXiv e-prints (2022), arXiv–2210

  132. [138]

    Ji Qi, Chuchun Zhang, Xiaozhi Wang, Kaisheng Zeng, Jifan Yu, Jinxin Liu, Lei Hou, Juanzi Li, and Xu Bin. 2023. Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information Extraction. In Proceedings of the 2023 Conference on Empirical Methods in Natura...

  133. [139]

    Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. 2024. Visual adversarial examples jailbreak aligned large language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 21527–21536

  134. [140]

    Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2024. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!. In The Twelfth International Conference on Learning Representations . https://openrev...

  135. [141]

    Yifu Qiu, Yftah Ziser, Anna Korhonen, Edoardo Ponti, and Shay B Cohen. 2023. Detecting and Mitigating Hallucinations in Multilingual Summarisation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 8914–8932

  136. [142]

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023. Direct preference optimization: your language model is secretly a reward model. In Proceedings of the 37th International Conference on Neural Information Processing Sys...

  137. [143]

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems 36 (2024)

  138. [144]

    Vivek Ramanujan, Thao Nguyen, Sewoong Oh, Ali Farhadi, and Ludwig Schmidt. 2023. On the Connection between Pre-training Data Diversity and Fine-tuning Robustness. In Thirty-seventh Conference on Neural Information Processing Systems . https://openreview.net/ forum?id=2SScUiWUbn

  139. [145]

    Michael Ryan, Tarek Naous, and Wei Xu. 2023. Revisiting non-English Text Simplification: A Unified Multilingual Benchmark. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 4898–4927

  140. [146]

    Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyo...

  141. [147]

    Sebastin Santy, Jenny T Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap. 2023. NLPositionality: Characterizing Design Biases of Datasets and Models. In The 61st Annual Meeting Of The Association For Computational Linguistics

  142. [148]

    Soumya Sanyal, Zeyi Liao, and Xiang Ren. 2022. RobustLR: A diagnostic benchmark for evaluating logical robustness of deductive reasoners. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 9614–9631

  143. [149]

    Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024. Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations. htt...

  144. [150]

    Ziyu Shang, Wenjun Ke, Nana Xiu, Peng Wang, Jiajun Liu, Yanhui Li, Zhizhao Luo, and Ke Ji. 2024. Ontofact: Unveiling fantastic fact-skeleton of llms via ontology-driven reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 18934–18943

  145. [151]

    NAN SHAO, Zefan Cai, Hanwei xu, Chonghua Liao, Yanan Zheng, and Zhilin Yang. 2023. Compositional Task Representations for Large Language Models. InThe Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=6axIMJA7ME3 Proc. ACM Meas. Ana...

  146. [152]

    Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh. 2024. Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=plmBsXHxgR

  147. [153]

    Ke Shen. 2024. The Generalization and Robustness of Transformer-Based Language Models on Commonsense Reasoning. InProceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 23419–23420

  148. [154]

    Atsushi Shirafuji, Yutaka Watanobe, Takumi Ito, Makoto Morishita, Yuki Nakamura, Yusuke Oda, and Jun Suzuki. 2023. Exploring the robustness of large language models for solving programming problems. arXiv preprint arXiv:2306.14583 (2023)

  149. [155]

    Manli Shu, Weili Nie, De-An Huang, Zhiding Yu, Tom Goldstein, Anima Anandkumar, and Chaowei Xiao. 2022. Test-time prompt tuning for zero-shot generalization in vision-language models. Advances in Neural Information Processing Systems 35 (2022), 14274–14289

  150. [156]

    Yang Shu, Xingzhuo Guo, Jialong Wu, Ximei Wang, Jianmin Wang, and Mingsheng Long. 2023. Clipood: Generalizing clip to out-of-distributions. In International Conference on Machine Learning . PMLR, 31716–31731

  151. [157]

    Chenglei Si, Dan Friedman, Nitish Joshi, Shi Feng, Danqi Chen, and He He. 2023. Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations. Association for Computational Linguistics

  152. [158]

    Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, and Shauli Ravfogel. 2023. The curious case of hallucinatory (un) answerability: Finding truths in the hidden states of over-confident large language models. In Proceedings of the 2023 Conference on Empirical Methods in N...

  153. [159]

    Joe Stacey, Yonatan Belinkov, and Marek Rei. 2022. Supervising model attention with human explanations for robust natural language inference. In Proceedings of the AAAI conference on artificial intelligence , Vol. 36. 11349–11357

  154. [160]

    Michal Štefánik. 2022. Methods for Estimating and improving robustness of language models. arXiv preprint arXiv:2206.08446 (2022)

  155. [161]

    Asa Cooper Stickland, Sailik Sengupta, Jason Krone, Saab Mansour, and He He. 2022. Robustification of multilingual language models to real-world noise in crosslingual zero-shot settings with robust contrastive pretraining. arXiv preprint arXiv:2210.04782 (2022)

  156. [162]

    Alessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schölkopf, and Mrinmaya Sachan. 2022. A causal framework to quantify the robustness of mathematical reasoning with language models. arXiv preprint arXiv:2210.12023 (2022)

  157. [163]

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing 568 (2024), 127063

  158. [164]

    Hao Sun and Mihaela van der Schaar. 2024. Inverse-RLignment: Inverse Reinforcement Learning from Demonstrations for LLM Alignment. arXiv preprint arXiv:2405.15624 (2024)

  159. [165]

    Jiuding Sun, Chantal Shaib, and Byron C Wallace. 2024. Evaluating the Zero-shot Robustness of Instruction-tuned Language Models. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=g9diuvxN6D

  160. [166]

    Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. 2023. Retentive network: A successor to transformer for large language models. arXiv preprint arXiv:2307.08621 (2023)

  161. [168]

    Yuhong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Hui Zhao. 2024. Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem. arXiv preprint arXiv:2403.03558 (2024)

  162. [169]

    Song Tang, Wenxin Su, Mao Ye, and Xiatian Zhu. 2024. Source-Free Domain Adaptation with Frozen Multimodal Foundation Model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 23711–23720

  163. [170]

    Junjiao Tian, Zecheng He, Xiaoliang Dai, Chih-Yao Ma, Yen-Cheng Liu, and Zsolt Kira. 2023. Trainable projected gradient method for robust fine-tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7836–7845

  164. [171]

    Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D Manning, and Chelsea Finn. 2024. Fine-Tuning Language Models for Factuality. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=WPZ2yPag4K

  165. [172]

    Andrea Tocchetti, Lorenzo Corti, Agathe Balayn, Mireia Yurrita, Philip Lippmann, Marco Brambilla, and Jie Yang. 2022. AI robustness: a human-centered perspective on technological challenges and opportunities. Comput. Surveys (2022)

  166. [173]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)

  167. [174]

    Thiagarajan

    Puja Trivedi, Danai Koutra, and Jayaraman J. Thiagarajan. 2023. A Closer Look at Model Adaptation using Feature Distortion and Simplicity Bias. In The Eleventh International Conference on Learning Representations . https://openreview.net/forum?id=wkg_b4-IwTZ

  168. [175]

    Faizad Ullah, Ali Faheem, Ubaid Azam, Muhammad Sohaib Ayub, Faisal Kamiran, and Asim Karim. 2024. Detecting Cybercrimes in Accordance with Pakistani Law: Dataset and Evaluation Using PLMs. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, ...

  169. [176]

    Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023. Poisoning language models during instruction tuning. In International Conference on Machine Learning . PMLR, 35413–35425

  170. [177]

    Guangya Wan, Yuqi Wu, Mengxuan Hu, Zhixuan Chu, and Sheng Li. 2024. Bridging causal discovery and large language models: A comprehensive survey of integrative approaches and future directions. arXiv preprint arXiv:2402.11068 (2024). Proc. ACM Meas. Anal. Comput. Syst., Vol. 37...

  171. [178]

    Boxin Wang, Wei Ping, Chaowei Xiao, Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Bo Li, Anima Anandkumar, and Bryan Catanzaro. 2022. Exploring the limits of domain-adaptive training for detoxifying large-scale language models. Advances in Neural Information Processing Systems 3...

  172. [180]

    Shiqi Wang, Zheng Li, Haifeng Qian, Chenghao Yang, Zijian Wang, Mingyue Shang, Varun Kumar, Samson Tan, Baishakhi Ray, Parminder Bhatia, et al. 2022. Recode: Robustness evaluation of code generation models. arXiv preprint arXiv:2212.10264 (2022)

  173. [181]

    Sibo Wang, Jie Zhang, Zheng Yuan, and Shiguang Shan. 2024. Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 24502–24511

  174. [182]

    Tianlu Wang, Rohit Sridhar, Diyi Yang, and Xuezhi Wang. 2021. Identifying and mitigating spurious correlations for improving robustness in nlp models. arXiv preprint arXiv:2110.07736 (2021)

  175. [183]

    Xuezhi Wang, Haohan Wang, and Diyi Yang. 2021. Measure and improve robustness in NLP models: A survey. arXiv preprint arXiv:2112.08313 (2021)

  176. [184]

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations. 1–24

  177. [185]

    Yihan Wang, Si Si, Daliang Li, Michal Lukasik, Felix Yu, Cho-Jui Hsieh, Inderjit S Dhillon, and Sanjiv Kumar. 2024. Two-stage LLM Fine-tuning with Less Specialization and More Generalization. In The Twelfth International Conference on Learning Representations . https://openrev...

  178. [186]

    Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How Does LLM Safety Training Fail?. In Thirty-seventh Conference on Neural Information Processing Systems . https://openreview.net/forum?id=jA235JGM09

  179. [187]

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682 (2022)

  180. [188]

    Zhipeng Wei, Jingjing Chen, Micah Goldblum, Zuxuan Wu, Tom Goldstein, and Yu-Gang Jiang. 2022. Towards transferable adversarial attacks on vision transformers. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 2668–2676

  181. [189]

    Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang, Chengfei Li, Jinfeng Bai, and Minlie Huang. 2023. Unveiling the Implicit Toxicity in Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 1322–1338

  182. [190]

    Anpeng Wu, Kun Kuang, Minqin Zhu, Yingrong Wang, Yujia Zheng, Kairong Han, Baohong Li, Guangyi Chen, Fei Wu, and Kun Zhang

  183. [191]

    Cheng-En Wu, Yu Tian, Haichao Yu, Heng Wang, Pedro Morgado, Yu Hen Hu, and Linjie Yang. 2023. Why Is Prompt Tuning for Vision-Language Models Robust to Noisy Labels?. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 15488–15497

  184. [192]

    Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui. 2023. Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 3909–3925

  185. [193]

    Yu Xia, Tong Yu, Zhankui He, Handong Zhao, Julian McAuley, and Shuai Li. 2024. Aligning as Debiasing: Causality-Aware Alignment via Reinforcement Learning with Interventional Feedback. In Proceedings of the 2024 Conference of the North American Chapter of the Association for C...

  186. [194]

    CoRR (2024)

    Causality for Large Language Models. CoRR (2024)

  187. [195]

    Yao Xiao, Ziyi Tang, Pengxu Wei, Cong Liu, and Liang Lin. 2023. Masked images are counterfactual samples for robust fine-tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 20301–20310

  188. [196]

    Zi Xiong, Lizhi Qing, Yangyang Kang, Jiawei Liu, Hongsong Li, Changlong Sun, Xiaozhong Liu, and Wei Lu. 2024. Enhance Robustness of Language Models Against Variation Attack through Graph Integration. arXiv preprint arXiv:2404.12014 (2024)

  189. [197]

    Binfeng Xu, Zhiyuan Peng, Bowen Lei, Subhabrata Mukherjee, Yuchen Liu, and Dongkuan Xu. 2023. Rewoo: Decoupling reasoning from observations for efficient augmented language models. arXiv preprint arXiv:2305.18323 (2023)

  190. [198]

    Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li. 2024. BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models. In The Twelfth International Conference on Learning Representations . https: //openreview.net/forum?id=c...

  191. [199]

    Weijia Xu, Batool Haider, and Saab Mansour. 2020. End-to-end slot alignment and recognition for cross-lingual NLU. arXiv preprint arXiv:2004.14353 (2020)

  192. [200]

    Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen, Yepeng Liu, Ladislau Bölöni, and Qian Lou. 2023. TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models. In Thirty-seventh Conference on Neural Information Processing Systems . https://openreview. net/forum?id=ZejTutd...

  193. [201]

    Kevin Yang, Dan Klein, Asli Celikyilmaz, Nanyun Peng, and Yuandong Tian. 2024. RLCD: Reinforcement learning from contrastive distillation for LM alignment. In The Twelfth International Conference on Learning Representations

  194. [202]

    Silei Xu, Shicheng Liu, Theo Culhane, Elizaveta Pertseva, Meng-Hsi Wu, Sina Semnani, and Monica Lam. 2023. Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over Wikidata. In Proceedings of the 2023 Conference on Empirical Methods ...

  195. [203]

    Yu Yang, Besmira Nushi, Hamid Palangi, and Baharan Mirzasoleiman. 2023. Mitigating spurious correlations in multi-modal models during fine-tuning. In International Conference on Machine Learning . PMLR, 39365–39379

  196. [204]

    Ziqing Yang, Xinlei He, Zheng Li, Michael Backes, Mathias Humbert, Pascal Berrang, and Yang Zhang. 2023. Data poisoning attacks against multimodal encoders. In International Conference on Machine Learning . PMLR, 39299–39313

  197. [205]

    Zonghan Yang and Yang Liu. 2022. On Robust Prefix-Tuning for Text Classification. In International Conference on Learning Representa- tions. https://openreview.net/forum?id=eBCmOocUejf

  198. [206]

    Yuting Yang, Pei Huang, Feifei Ma, Juan Cao, and Jintao Li. 2024. PAD: A Robustness Enhancement Ensemble Method via Promoting Attention Diversity. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-CO...

  199. [207]

    Ziyi Yin, Muchao Ye, Tianrong Zhang, Jiaqi Wang, Han Liu, Jinghui Chen, Ting Wang, and Fenglong Ma. 2024. VQAttack: Transferable Adversarial Attacks on Visual Question Answering via Pre-trained Models. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 6755–6763

  200. [208]

    Da Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi, Huseyin A Inan, Gautam Kamath, Janardhan Kulkarni, Yin Tat Lee, Andre Manoel, Lukas Wutschitz, Sergey Yekhanin, and Huishuai Zhang. 2022. Differentially Private Fine-tuning of Language Models. In International Conference on ...

  201. [209]

    Jifan Yu, Xiaohan Zhang, Yifan Xu, Xuanyu Lei, Zijun Yao, Jing Zhang, Lei Hou, and Juanzi Li. 2024. A Cause-Effect Look at Alleviating Hallucination of Knowledge-grounded Dialogue Generation. arXiv preprint arXiv:2404.03491 (2024)

  202. [210]

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. Advances in neural information processing systems 36 (2023), 11809–11822

  203. [211]

    Zichun Yu, Chenyan Xiong, Shi Yu, and Zhiyuan Liu. 2023. Augmentation-Adapted Retriever Improves Generalization of Language Models as Generic Plug-In. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2421–2436

  204. [212]

    Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, YX Wei, Lean Wang, Zhiping Xiao, et al. 2025. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. arXiv preprint arXiv:2502.11089 (2025)

  205. [214]

    Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, and Tat-Seng Chua. 2024. RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback. In Proceedings of the IEEE/CV...

  206. [215]

    Zhao Yukun, Yan Lingyong, Sun Weiwei, Xing Guoliang, Wang Shuaiqiang, Meng Chong, Cheng Zhicong, Ren Zhaochun, and Yin Dawei. 2024. Improving the robustness of large language models via consistency alignment. arXiv preprint arXiv:2403.14221 (2024)

  207. [216]

    Maxime Zanella and Ismail Ben Ayed. 2024. On the Test-Time Zero-Shot Generalization of Vision-Language Models: Do We Really Need Prompt Learning?. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 23783–23793

  208. [217]

    Susskind, and Chen Huang

    Yuhang Zang, Hanlin Goh, Joshua M. Susskind, and Chen Huang. 2024. Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id= PKICZXVY9M

  209. [218]

    Mert Yuksekgonul, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar, Ranjita Naik, Hamid Palangi, Ece Kamar, and Besmira Nushi

  210. [219]

    In The Twelfth International Conference on Learning Representations

    Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=gfFVATffPd

  211. [220]

    Tianhang Zhang, Lin Qiu, Qipeng Guo, Cheng Deng, Yue Zhang, Zheng Zhang, Chenghu Zhou, Xinbing Wang, and Luoyi Fu. 2023. Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Proc...

  212. [221]

    Yuansen Zhang, Xiao Wang, Zhiheng Xi, Han Xia, Tao Gui, Qi Zhang, and Xuanjing Huang. 2024. RoCoIns: Enhancing Robustness of Large Language Models through Code-Style Instructions. arXiv preprint arXiv:2402.16431 (2024)

  213. [222]

    Yabin Zhang, Wenjie Zhu, Hui Tang, Zhiyuan Ma, Kaiyang Zhou, and Lei Zhang. 2024. Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 28718–28728. Proc. ...

  214. [223]

    Abdelrahman Zayed, Gonçalo Mordido, Samira Shabanian, Ioana Baldini, and Sarath Chandar. 2024. Fairness-aware structured pruning in transformers. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 22484–22492

  215. [224]

    Chongzhi Zhang, Aishan Liu, Xianglong Liu, Yitao Xu, Hang Yu, Yuqing Ma, and Tianlin Li. 2020. Interpreting and improving adversarial robustness of deep neural networks with neuron sensitivity. IEEE Transactions on Image Processing 30 (2020), 1291–1304

  216. [225]

    Shuai Zhao, Xiaohan Wang, Linchao Zhu, and Yi Yang. 2024. Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id= kIP0duasBb

  217. [226]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023)

  218. [227]

    Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai man Cheung, and Min Lin. 2023. On Evaluating Adversarial Robustness of Large Vision-Language Models. InThirty-seventh Conference on Neural Information Processing Systems. https://openreview. net/forum?id=xbbknN9QFs

  219. [228]

    Feng Zhao, Wan Xianlin, Cheng Yan, and Chu Kiong Loo. 2024. Correcting Language Model Bias for Text Classification in True Zero-Shot Learning. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING...

  220. [229]

    Jiaxu Zhao, Meng Fang, Zijing Shi, Yitong Li, Ling Chen, and Mykola Pechenizkiy. 2023. CHBias: Bias Evaluation and Mitigation of Chinese Conversational Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...

  221. [230]

    Rui Zheng, Wei Shen, Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou, Zhiheng Xi, Xiao Wang, Haoran Huang, Tao Gui, Qi Zhang, and Xuanjing Huang. 2024. Improving Generalization of Alignment with Human Preferences through Group Invariant Learning. In The Twelfth International Conf...

  222. [231]

    Li Zhong and Zilong Wang. 2024. Can LLM Replace Stack Overflow? A Study on Robustness and Reliability of Large Language Model Code Generation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 21841–21849

  223. [232]

    Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2sql: Generating structured queries from natural language using reinforcement learning. arXiv preprint arXiv:1709.00103 (2017)

  224. [233]

    Yilun Zhao, Chen Zhao, Linyong Nan, Zhenting Qi, Wenlin Zhang, Xiangru Tang, Boyu Mi, and Dragomir Radev. 2023. RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations. In Proceedings of the 61st Annual Meeting of the Association for ...

  225. [234]

    Rui Zheng, Rong Bao, Qin Liu, Tao Gui, Qi Zhang, Xuan-Jing Huang, Rui Xie, and Wei Wu. 2022. Plugat: A plug and play module to defend against textual adversarial attack. In Proceedings of the 29th International Conference on Computational Linguistics . 2873–2882

  226. [235]

    Xuhui Zhou, Yue Zhang, Leyang Cui, and Dandan Huang. 2020. Evaluating commonsense in pre-trained language models. InProceedings of the AAAI conference on artificial intelligence , Vol. 34. 9733–9740

  227. [236]

    Yuhao Zhou, Wenxiang Chen, Rui Zheng, Zhiheng Xi, Tao Gui, Qi Zhang, and Xuan-Jing Huang. 2024. ORTicket: Let One Robust BERT Ticket Transfer across Different Tasks. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and ...

  228. [237]

    Zixuan Zhou, Xuefei Ning, Ke Hong, Tianyu Fu, Jiaming Xu, Shiyao Li, Yuming Lou, Luning Wang, Zhihang Yuan, Xiuhong Li, et al

  229. [238]

    Fan Zhou, Yuzhou Mao, Liu Yu, Yi Yang, and Ting Zhong. 2023. Causal-debias: Unifying debiasing in pretrained language models and fine-tuning via causal invariant learning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Lon...

  230. [239]

    Wenjie Zhou, Qiang Wang, Mingzhou Xu, Ming Chen, and Xiangyu Duan. 2024. Revisiting the Self-Consistency Challenges in Multi-Choice Question Formats for Large Language Model Evaluation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Lan...

  231. [240]

    Shaolin Zhu, Menglong Cui, and Deyi Xiong. 2024. Towards robust in-context learning for machine translation with large language models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)....

  232. [243]

    arXiv preprint arXiv:2404.14294 (2024)

    A survey on efficient inference for large language models. arXiv preprint arXiv:2404.14294 (2024)

  233. [244]

    Zihao Zhou, Qiufeng Wang, Mingyu Jin, Jie Yao, Jianan Ye, Wei Liu, Wei Wang, Xiaowei Huang, and Kaizhu Huang. 2024. Mathattack: Attacking large language models towards math solving ability. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 19750–19758

  234. [245]

    Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Yue Zhang, Neil Zhenqiang Gong, et al. 2023. Promptbench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv preprint arXiv:2306.04528 (2023)

  235. [2023]

    arXiv preprint arXiv:2302.12095 (2023)

    On the robustness of chatgpt: An adversarial and out-of-distribution perspective. arXiv preprint arXiv:2302.12095 (2023)

  236. [2024]

    Advances in Neural Information Processing Systems 36 (2024)

    Efficient equivariant transfer learning from pretrained models. Advances in Neural Information Processing Systems 36 (2024)

  237. [2025]

    arXiv preprint arXiv:2502.13189 (2025)

    MoBA: Mixture of Block Attention for Long-Context LLMs. arXiv preprint arXiv:2502.13189 (2025)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.