Pith. sign in

REVIEW 3 major objections 4 minor 3 cited by

Challenges in Guardrailing Large Language Models for Science

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read General-purpose LLM guardrails are insufficient for science, and this paper proposes a science-specific framework organized around trustworthiness, ethics & bias, safety, and legal dimensions, with white-box, black-box, and gray-box…

desk verdict A well-organized position paper that maps known guardrail dimensions onto scientific applications, but its central insufficiency claim is asserted rather than shown, and one internal inconsistency undercuts the color taxonomy. read the letter →

arxiv 2411.08181 v2 pith:FB7QA3HU submitted 2024-11-12 cs.AI

classification cs.AI
keywords largelanguagemodelsLLMguardrailsscientificintegritytrustworthinessAIsafetyretrieval-augmentedgenerationknowledgecontextualizationtimesensitivity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the guardrails currently wrapped around general-purpose LLMs—filters for toxicity, hallucination, privacy, and the like—are not enough when the models are used in scientific research. It claims that scientific use introduces demands that general guardrails do not meet: outputs must stay current as knowledge changes, must be adapted to disciplinary and local context, must handle conflicting findings honestly, and must respect the intellectual property norms of research. To close that gap, the authors propose a science-specific guardrail framework with four dimensions—trustworthiness, ethics & bias, safety, and legal—and a color-coded map of which sub-dimensions need minimal, moderate, or substantial new work. They then tie each dimension to implementation strategies drawn from white-box, black-box, and gray-box interventions, ranging from fine-tuning to external fact-checking and retrieval-augmented generation. A sympathetic reader would take the paper's contribution to be a structured agenda for building and evaluating safe LLM systems for science, not a demonstrated solution.

What carries the argument

The load-bearing device is the taxonomic framework itself: a two-dimensional organization that pairs four guardrail purposes with the level of intervention, plus a table that connects specific techniques to specific dimensions. The framework's contribution is a color-coded adaptation scale—blue means an existing guardrail transfers with minimal change, orange means it needs refinement, red means it is underdeveloped or absent in the general literature. This scale carries the paper's main argument by showing where the scientific domain is genuinely underserved, and the table carries the implementation claim by showing that each dimension can be addressed through some combination of white-box, black-box, and gray-box methods.

What would settle it

A controlled comparison would settle it: run a set of time-sensitive and context-dependent scientific queries through a general-purpose guardrailed LLM and through a system using the proposed four-dimension framework, then measure factuality, consistency, and appropriateness of output. If the general guardrails match or beat the specialized framework without modification, the insufficiency claim fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that scientific research is a distinct enough deployment context that general-purpose LLM guardrails cannot simply be reused; they must be reorganized around science-specific goals. The proposed framework replaces a one-size-fits-all list of safety properties with four categories—trustworthiness (verification, uncertainty identification, consistency, factuality, hallucination identification, knowledge contextualization, time sensitivity, attribution, explainability), ethics & bias (fairness, societal impact), safety (robustness, jailbreak prevention, out-of-distribution checks, harmfulness detection), and legal (privacy, compliance, IP & copyright). Within this scheme, the paper identifies time sensitivity (responses must track the latest findings), knowledge contextualization (advice must adapt to field, region, and user), conflict resolution (contrasting studies must be reconciled rather than flattened), and intellectual property protection as the dimensions where existing guardrails are most lacking and where the largest design effort is required. It then maps white-box, black-box, and gray-box strategies onto these dimensions in a two-way table, effectively providing a design menu for guarding scientific LLM applications.

Load-bearing premise

The load-bearing premise is that scientific research poses challenges that are categorically distinct from general-purpose use—temporal currency, contextualized knowledge, conflict resolution, and IP handling—so that existing guardrails cannot be made sufficient with only minor extensions.

Editorial extensions

If this is right

  • If the framework is right, project teams building LLMs for science should budget for layered guardrails rather than relying on a single off-the-shelf filter.
  • Evaluators of scientific LLM systems would need to test temporal currency and contextual fit, not just toxicity and hallucination rates.
  • Retrieval-augmented generation and human-in-the-loop verification emerge as core mechanisms for the time-sensitivity and knowledge-contextualization dimensions.
  • The color-coding implies a roadmap: most existing guardrails need only blue-level adaptation, while time sensitivity and knowledge contextualization need red-level development.
  • The framework gives a shared vocabulary for comparing scientific LLM deployments, which could make audit and regulatory review more systematic.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension would be to turn the four dimensions into a scoring rubric for auditing existing scientific LLM applications, with each dimension weighted by domain risk.
  • The framework suggests a testable prediction: science-specific guardrails will reduce task-specific harms (e.g., outdated medical advice, misplaced agricultural recommendations) more than general guardrails do, even when both appear safe on generic toxicity benchmarks.
  • The compliance dimension, as described, implies a more active role for LLMs in research administration—matching grants and avoiding duplication—that goes beyond typical safety guardrails and may require its own evaluation.
  • Because the paper is a framework proposal, the strongest next step implied by its own conclusion is an empirical benchmark; without one, the insufficiency claim remains an assertion.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper argues that existing general-purpose LLM guardrails are insufficient for scientific applications and proposes a taxonomy of guardrail dimensions organized under four categories: trustworthiness, ethics & bias, safety, and legal. Each dimension is color-coded by the amount of adaptation required (blue, orange, red, uncolored). The paper also maps white-box, black-box, and gray-box implementation strategies to these dimensions in Table II. The contribution is primarily conceptual: a structured framework and a strategy matrix, with empirical validation deferred to future work.

Significance. If the framework were validated, it would offer a useful organizing structure for building science-specific LLM guardrails. The paper's strengths are its broad literature grounding, the explicit taxonomy in Fig. 1, and the strategy matrix in Table II, which collects many relevant techniques and citations. However, the central claim that general-purpose guardrails are insufficient is asserted rather than demonstrated, and the framework contains internal inconsistencies. The paper is honest about deferring empirical testing, but as it stands, the significance is as a position paper or survey rather than a validated guideline.

major comments (3)
  1. [Section II-B and Section III] The central claim that 'existing general-purpose LLM guardrails are insufficient' is not established. The paper's own Table II maps the allegedly science-specific dimensions (time sensitivity, knowledge contextualization, IP & copyright) to generic techniques such as retrieval-augmented generation, external fact-checking, and human-in-the-loop systems, which are already part of general-purpose guardrail toolkits. The conclusion explicitly states that empirical validation is future work. The manuscript either needs a systematic comparison showing where general-purpose guardrails fail, or the claim should be weakened to 'require domain-specific adaptation' rather than 'insufficient.'
  2. [Section III-A and III-A-1] There is a direct contradiction in the color-coding scheme. Blue boxes are defined as 'established best practices that can transition smoothly into the scientific context, requiring only slight adjustments,' yet the Compliance dimension, which is blue in Fig. 1, states that 'No guardrail dimensions exist yet that address this problem.' Moreover, Table I lists 'Legality' as supported by Llama Guard, so the statement is factually inconsistent with the paper's own survey. This contradiction undermines the principled basis of the taxonomy.
  3. [Section III and Table II] The dimensions and implementation strategies are not consistently aligned. Conflict resolution is presented as a key challenge in the abstract and in Section II-B-3, but it is not a distinct dimension in Fig. 1; it is subsumed under Knowledge Contextualization in Section III-C-2. Additionally, Table II contains citation errors: the 'Plagiarism Detection' row cites [35] (Carlini et al., 'Extracting Training Data from Large Language Models'), which concerns privacy attacks and training data extraction, not plagiarism detection. These issues undermine the reliability of the framework as a guideline.
minor comments (4)
  1. [Section II-B-3] The heading 'Conflict Identification and Resolutiuon' contains a typo ('Resolutiuon' should be 'Resolution'). Similar typos appear elsewhere, such as 'T emporal Relevancy' in Section II-B-2.
  2. [Section II-A-5] The subsection title 'Toxicity and Legality' does not match Table I, which lists 'Toxicity' and 'Legality' as separate rows; it would be clearer to treat them as distinct aspects.
  3. [Table II] The table's formatting makes it difficult to determine which dimensions are addressed by each strategy; cells contain lists of citations without clear column denotation. The 'Knowledge Base Integration' and 'External Knowledge Integration' rows both cite [121], which duplicates [109].
  4. [Abstract and Section IV] The abstract claims 'comprehensive guidelines,' but the body provides a high-level taxonomy and strategy enumeration rather than actionable procedures. Consider describing the contribution as a 'framework' or 'taxonomy' to avoid overclaiming.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the paper is a conceptual survey and framework proposal with no fitted predictions, derived equations, or self-citation chains that reduce to its own inputs.

full rationale

This manuscript is a position/survey paper, not a derivation-based or prediction-based work. It proposes a taxonomy of LLM guardrail dimensions for science and maps existing techniques onto white-box, black-box, and gray-box implementation strategies. There are no equations, fitted parameters, or empirical predictions whose outputs could equal their inputs by construction. The central claim that general-purpose guardrails are insufficient for science is asserted with examples and citations, but that is a substantiveness/correctness concern, not a circularity concern. The paper explicitly defers empirical validation to future work, which further confirms that no fitted result is dressed up as a prediction. The cited prior work, including Dong et al. and SciGuard, provides external support and is not used to forbid alternatives or to define the framework's dimensions into existence. One can note an internal inconsistency in the color-coding of Compliance, but inconsistency is not circular reasoning. Therefore, no specific circular step can be quoted or exhibited, and the honest finding is a circularity score of 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The framework is built on the assumption that scientific LLM use has unique guardrail needs and that the proposed categorization is useful. These are not empirically verified. No free parameters or invented entities are introduced because the paper makes no quantitative claims.

assumptions (2)
  • domain assumption The scientific domain requires guardrail dimensions beyond general-purpose ones
    Central premise of the paper, asserted in Section II-B and used to justify the proposed framework.
  • ad hoc to paper The white-box/black-box/gray-box categorization, originally for attacks, is a valid organizing principle for guardrail implementation
    Adopted from Dong et al. and extended to implementation strategies in Section IV; no justification beyond analogy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Challenges in Guardrailing Large Language Models for Science." pith.science (2026). https://pith.science/paper/FB7QA3HU

@misc{pith2026241108181,
  author       = {Pith},
  title        = {Pith review of: Challenges in Guardrailing Large Language Models for Science},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FB7QA3HU}},
  note         = {Machine review of arXiv:2411.08181}
}
read the original abstract

The rapid development in large language models (LLMs) has transformed the landscape of natural language processing and understanding (NLP/NLU), offering significant benefits across various domains. However, when applied to scientific research, these powerful models exhibit critical failure modes related to scientific integrity and trustworthiness. Existing general-purpose LLM guardrails are insufficient to address these unique challenges in the scientific domain. We provide comprehensive guidelines for deploying LLM guardrails in the scientific domain. We identify specific challenges -- including time sensitivity, knowledge contextualization, conflict resolution, and intellectual property concerns -- and propose a guideline framework for the guardrails that can align with scientific needs. These guardrail dimensions include trustworthiness, ethics & bias, safety, and legal aspects. We also outline in detail the implementation strategies that employ white-box, black-box, and gray-box methodologies that can be enforced within scientific contexts.

Figures

Figures reproduced from arXiv: 2411.08181 by the authors.

Figure 1
Figure 1. Key Aspects of LLM Guardrails in Scientific Domains [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use

    cs.CL 2026-03 conditional novelty 6.0 of 10

    A single post-training recipe — explicit safety reasoning plus refusal as a first-class action, optimized by pairwise preference RL — reduces harmful tool-use and injection success on open LLM agents, with model-depen...

  2. SI-FACT: Mitigating Knowledge Conflict via Self-Improving Faithfulness-Aware Contrastive Tuning

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Self-generated contrastive tuning (SI-FACT) raises contextual answer rates on ECARE_KRE from 69.8% (best baseline) to 76.0%, and on COSE_KRE from 52.9% to 54.2%.

  3. LLM Harms: A Taxonomy and Discussion

    cs.CY 2025-12 unverdicted novelty 3.0 of 10

    This paper proposes a taxonomy of LLM harms in five categories and suggests mitigation strategies plus a dynamic auditing system for responsible development.

Reference graph

Works this paper leans on

127 extracted references · 19 canonical work pages · cited by 3 Pith papers

  1. [35]

    Extracting training data from large language models,

    N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herb ert-V oss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsso n et al., “Extracting training data from large language models,” in 30th USENIX Security Symposium (USENIX Security 21) , 2021, pp. 2633–2650

  2. [1]

    Language models are few-shot learners,

    T. B. Brown, “Language models are few-shot learners,” arXiv preprint arXiv:2005.14165 , 2020

  3. [2]

    Improving language understanding by gener ative pre-training,

    A. Radford, “Improving language understanding by gener ative pre-training,” 2018

  4. [3]

    Finetuned language models are zero -shot learners,

    J. Wei, M. Bosma, V . Y . Zhao, K. Guu, A. W. Y u, B. Lester, N. D u, A. M. Dai, and Q. V . Le, “Finetuned language models are zero -shot learners,” arXiv preprint arXiv:2109.01652 , 2021

  5. [4]

    On the dangers of stochastic parrots: Can language mod els be too big?

    E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitch ell, “On the dangers of stochastic parrots: Can language mod els be too big?” in Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, 2021, pp. 610–623

  6. [5]

    Bot-adversarial dialogue for safe conversational agents ,

    J. Xu, D. Ju, M. Li, Y .-L. Boureau, J. Weston, and E. Dinan, “Bot-adversarial dialogue for safe conversational agents ,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologi es, 2021, pp. 2950–2968

  7. [6]

    Safeguarding la rge language models: A survey,

    Y . Dong, R. Mu, Y . Zhang, S. Sun, T. Zhang, C. Wu, G. Jin, Y . Q i, J. Hu, J. Meng, S. Bensalem, and X. Huang, “Safeguarding la rge language models: A survey,” arXiv preprint arXiv: 2406.02622 , 2024

  8. [7]

    The social impact of natural lan guage processing,

    D. Hovy and S. L. Spruit, “The social impact of natural lan guage processing,” in Proceedings of the 54th Annual Meeting of the Association fo r Computational Linguistics (V olume 2: Short Papers) , 2016, pp. 591–598

Show all 127 references
  1. [8]

    Dissecting racial bias in an algorithm used to manage the he alth of populations,

    Z. Obermeyer, B. Powers, C. V ogeli, and S. Mullainathan, “Dissecting racial bias in an algorithm used to manage the he alth of populations,” Science, vol. 366, no. 6464, pp. 447–453, 2019

  2. [9]

    Chatgpt is fun, but not an author,

    H. H. Thorp, “Chatgpt is fun, but not an author,” pp. 313–3 13, 2023

  3. [10]

    Gpt-4 technical report,

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L . Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023

  4. [11]

    The llama 3 herd of models,

    A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A . Letman, A. Mathur, A. Schelten, A. Y ang, A. Fan et al. , “The llama 3 herd of models,” arXiv preprint arXiv:2407.21783 , 2024

  5. [12]

    Introducing claude 3.5 sonnet,

    Anthropic, “Introducing claude 3.5 sonnet,” Jun 2024. [Online]. Available: https://www.anthropic.com/news/claude-3-5-sonnet

  6. [13]

    Mixtral of experts,

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Sava ry, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Br essand et al. , “Mixtral of experts,” arXiv preprint arXiv:2401.04088 , 2024

  7. [14]

    Gemini: a family of highly capable multimodal models,

    G. Team, R. Anil, S. Borgeaud, Y . Wu, J.-B. Alayrac, J. Y u , R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth et al., “Gemini: a family of highly capable multimodal models,” arXiv preprint arXiv:2312.11805 , 2023

  8. [15]

    On the opportunities and risks of foundation models,

    R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora , S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunski ll et al., “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258 , 2021

  9. [16]

    The impact of large l anguage models on scientific discovery: a preliminary study using gpt-4,

    M. R. AI4Science and M. A. Quantum, “The impact of large l anguage models on scientific discovery: a preliminary study using gpt-4,” arXiv preprint arXiv:2311.07361, 2023

  10. [17]

    An interdisciplinary outlook on large language models for sci entific research,

    J. Boyko, J. Cohen, N. Fox, M. H. V eiga, J. I.-H. Li, J. Liu , B. Modenesi, A. H. Rauch, K. N. Reid, S. Tribedi, A. Visherat ina, and X. Xie, “An interdisciplinary outlook on large language models for sci entific research,” arXiv preprint arXiv: 2311.04929 , 2023

  11. [18]

    The ethics of chatg pt in medicine and healthcare: a systematic review on large l anguage models (llms),

    J. Haltaufderheide and R. Ranisch, “The ethics of chatg pt in medicine and healthcare: a systematic review on large l anguage models (llms),” npj Digital Medicine , vol. 7, no. 1, Jul. 2024. [Online]. Available: http://dx.d oi.org/10.1038/s41746-024-01157-x

  12. [19]

    Ethic al considerations and policy implications for large langua ge models: Guiding responsible development and deployment,

    J. Zhang, X. Ji, Z. Zhao, X. Hei, and K.-K. R. Choo, “Ethic al considerations and policy implications for large langua ge models: Guiding responsible development and deployment,” arXiv preprint arXiv: 2308.02678 , 2023

  13. [20]

    Ll ama guard: Llm-based input-output safeguard for human-ai conversations,

    H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y . Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, and M. Khabsa, “Ll ama guard: Llm-based input-output safeguard for human-ai conversations,” arXiv preprint arXiv: 2312.06674 , 2023

  14. [21]

    Nemo guardrails: A toolkit for controllable and safe ll m applications with programmable rails,

    T. Rebedea, R. Dinu, M. Sreedhar, C. Parisien, and J. Coh en, “Nemo guardrails: A toolkit for controllable and safe ll m applications with programmable rails,” arXiv preprint arXiv: 2310.10501 , 2023

  15. [22]

    Prompti ng is programming: A query language for large language model s,

    L. Beurer-Kellner, M. Fischer, and M. V echev, “Prompti ng is programming: A query language for large language model s,” Proceedings of the ACM on Programming Languages , vol. 7, no. PLDI, pp. 1946–1969, 2023

  16. [23]

    Guidance: A language f or controlling large language models,

    S. Slundberg and contributors, “Guidance: A language f or controlling large language models,” https://github.co m/guidance-ai/guidance, 2024, accessed: 2024-09-30

  17. [24]

    Guardrails: Adding guard rails to large language models,

    S. Rajpal and contributors, “Guardrails: Adding guard rails to large language models,” https://github.com/guar drails-ai/guardrails, 2024, accessed: 2024- 09-30

  18. [25]

    Ai-generate d clinical summaries require more than accuracy,

    K. E. Goodman, H. Y . Paul, and D. J. Morgan, “Ai-generate d clinical summaries require more than accuracy,” JAMA, 2024

  19. [26]

    A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questi ons,

    L. Huang, W. Y u, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin et al. , “A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questi ons,” arXiv preprint arXiv:2311.05232 , 2023. 13

  20. [27]

    Building guardrails for large language mod els,

    Y . Dong, R. Mu, G. Jin, Y . Qi, J. Hu, X. Zhao, J. Meng, W. Rua n, and X. Huang, “Building guardrails for large language mod els,” arXiv preprint arXiv: 2402.01822 , 2024

  21. [28]

    Chainpoll: A high efficacy metho d for llm hallucination detection,

    R. Friel and A. Sanyal, “Chainpoll: A high efficacy metho d for llm hallucination detection,” arXiv preprint arXiv:2310.18344 , 2023

  22. [29]

    Selfcheckgpt: Z ero-resource black-box hallucination detection for gener ative large language models,

    P . Manakul, A. Liusie, and M. J. Gales, “Selfcheckgpt: Z ero-resource black-box hallucination detection for gener ative large language models,” arXiv preprint arXiv:2303.08896, 2023

  23. [30]

    Gptscore: Evaluat e as you desire,

    J. Fu, S.-K. Ng, Z. Jiang, and P . Liu, “Gptscore: Evaluat e as you desire,” arXiv preprint arXiv:2302.04166 , 2023

  24. [31]

    G-eval: Nlg evaluation using gpt-4 with better human alignment,

    Y . Liu, D. Iter, Y . Xu, S. Wang, R. Xu, and C. Zhu, “G-eval: Nlg evaluation using gpt-4 with better human alignment,” arXiv preprint arXiv:2303.16634, 2023

  25. [32]

    Evaluation and mitigation of the limitations of large language models in cl inical decision-making,

    P . Hager, F. Jungmann, R. Holland, K. Bhagat, I. Hubrech t, M. Knauer, J. Vielhauer, M. Makowski, R. Braren, G. Kaissi s et al. , “Evaluation and mitigation of the limitations of large language models in cl inical decision-making,” Nature medicine, vol. 30, no. 9, pp. 2613–2622, 2024

  26. [33]

    An empirical su rvey of the effectiveness of debiasing techniques for pre-t rained language models,

    N. Meade, E. Poole-Dayan, and S. Reddy, “An empirical su rvey of the effectiveness of debiasing techniques for pre-t rained language models,” arXiv preprint arXiv:2110.08527, 2021

  27. [34]

    Mitigating un wanted biases with adversarial learning,

    B. H. Zhang, B. Lemoine, and M. Mitchell, “Mitigating un wanted biases with adversarial learning,” in Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society , 2018, pp. 335–340

  28. [36]

    The secret sharer: Evalu ating and testing unintended memorization in neural networ ks,

    N. Carlini, C. Liu, ´U. Erlingsson, J. Kos, and D. Song, “The secret sharer: Evalu ating and testing unintended memorization in neural networ ks,” in 28th USENIX security symposium (USENIX security 19) , 2019, pp. 267–284

  29. [37]

    Privacy issues in large language m odels: A survey,

    S. Neel and P . Chang, “Privacy issues in large language m odels: A survey,” arXiv preprint arXiv:2312.06717 , 2023

  30. [38]

    Beyo nd memorization: Violating privacy via inference with larg e language models,

    R. Staab, M. V ero, M. Balunovi’c, and M. T. V echev, “Beyo nd memorization: Violating privacy via inference with larg e language models,” International Conference on Learning Representations , 2023

  31. [39]

    Pro pile: Probing privacy leakage in large language models,

    S. Kim, S. Y un, H. Lee, M. Gubri, S. Y oon, and S. J. Oh, “Pro pile: Probing privacy leakage in large language models,” Advances in Neural Information Processing Systems, vol. 36, 2024

  32. [40]

    Smiles-prompting: A n ovel approach to llm jailbreak attacks in chemical synthesi s,

    A. Wong, H. Cao, Z. Liu, and Y . Li, “Smiles-prompting: A n ovel approach to llm jailbreak attacks in chemical synthesi s,” arXiv preprint arXiv: 2410.15641, 2024

  33. [41]

    Scientific large language models: A survey on bio logical & chemical domains,

    Q. Zhang, K. Ding, T. Lyv, X. Wang, Q. Yin, Y . Zhang, J. Y u, Y . Wang, X. Li, Z. Xiang, K. Feng, X. Zhuang, Z. Wang, M. Qin, M. Zhang, J. Zhang, J. Cui, T. Huang, P . Y an, R. Xu, H. Chen, X. Li, X. Fan, H. Xing, a nd H. Chen, “Scientific large language models: A survey on bi...

  34. [42]

    Baseline defenses for adversarial attacks against aligned language models,

    N. Jain, A. Schwarzschild, Y . Wen, G. Somepalli, J. Kirc henbauer, P . yeh Chiang, M. Goldblum, A. Saha, J. Geiping, an d T. Goldstein, “Baseline defenses for adversarial attacks against aligned language models,” arXiv preprint arXiv: 2309.00614 , 2023

  35. [43]

    Red teaming languag e models with language models,

    E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides , A. Glaese, N. McAleese, and G. Irving, “Red teaming languag e models with language models,” arXiv preprint arXiv:2202.03286 , 2022

  36. [44]

    On the robustness of chatgpt: An adversarial and out-of-distribution perspective,

    J. Wang, X. Hu, W. Hou, H. Chen, R. Zheng, Y . Wang, L. Y ang, H. Huang, W. Y e, X. Geng et al. , “On the robustness of chatgpt: An adversarial and out-of-distribution perspective,” arXiv preprint arXiv:2302.12095 , 2023

  37. [45]

    Red teaming large language models in medicine: Real-world insights on m odel behavior,

    C. T.-T. Chang, H. Farah, H. Gui, S. J. Rezaei, C. Bou-Kha lil, Y .-J. Park, A. Swaminathan, J. A. Omiye, A. Kolluri, A. C haurasia et al. , “Red teaming large language models in medicine: Real-world insights on m odel behavior,” medRxiv, pp. 2024–04, 2024

  38. [46]

    Right to be forgotten in the era of l arge language models: Implications, challenges, and solutions,

    D. Zhang, P . Finckenberg-Broman, T. Hoang, S. Pan, Z. Xi ng, M. Staples, and X. Xu, “Right to be forgotten in the era of l arge language models: Implications, challenges, and solutions,” AI and Ethics , 2023

  39. [47]

    Realtoxicityprompts: Evaluating neural toxic degenera tion in language models,

    S. Gehman, S. Gururangan, M. Sap, Y . Choi, and N. A. Smith , “Realtoxicityprompts: Evaluating neural toxic degenera tion in language models,” arXiv preprint arXiv:2009.11462, 2020

  40. [48]

    Opt: Open pre-trained transformer language models,

    S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Che n, C. Dewan, M. Diab, X. Li, X. V . Lin et al. , “Opt: Open pre-trained transformer language models,” arXiv preprint arXiv:2205.01068 , 2022

  41. [49]

    Ai transparency in the age of llm s: A human-centered research roadmap,

    Q. Liao and J. V aughan, “Ai transparency in the age of llm s: A human-centered research roadmap,” Special Issue 4: Grappling With the Generative AI Revolution, 2023

  42. [50]

    On the calibr ation of large language models and alignment,

    C. Zhu, B. Xu, Q. Wang, Y . Zhang, and Z. Mao, “On the calibr ation of large language models and alignment,” arXiv preprint arXiv: 2311.13240 , 2023

  43. [51]

    Bayesformer: Transformer with uncertainty estimation,

    K. A. Sankararaman, S. Wang, and H. Fang, “Bayesformer: Transformer with uncertainty estimation,” arXiv preprint arXiv:2206.00826 , 2022

  44. [52]

    Large language models, scientific knowledge and factua lity: A framework to streamline human expert evaluation,

    M. Wysocka, O. Wysocki, M. Delmas, V . Mutel, and A. Freit as, “Large language models, scientific knowledge and factua lity: A framework to streamline human expert evaluation,” Journal of Biomedical Informatics , 2023

  45. [53]

    Large language models in drug discovery and development: From disease mechanisms to clinical trials,

    Y . Zheng, H. Y . Koh, M. Y ang, L. Li, L. T. May, G. I. Webb, S. Pan, and G. Church, “Large language models in drug discovery and development: From disease mechanisms to clinical trials,” arXiv preprint arXiv: 2409.04481 , 2024

  46. [54]

    Assessing large language models on climate information,

    J. Bulian, M. S. Sch¨ afer, A. Amini, H. Lam, M. Ciaramita , B. Gaiarin, M. C. Huebscher, C. Buck, N. G. Mede, M. Leippold et al. , “Assessing large language models on climate information,” arXiv preprint arXiv:2310.02932 , 2023

  47. [55]

    Unlearning climate misinformation in la rge language models,

    M. Fore, S. Singh, C. Lee, A. Pandey, A. Anastasopoulos, and D. Stamoulis, “Unlearning climate misinformation in la rge language models,” in Proceedings of the 1st W orkshop on Natural Language Process ing Meets Climate Change (ClimateNLP 2024) , D. Stammbach, J. Ni, T. Schima...

  48. [56]

    Enha ncing large language models with climate resources,

    M. Kraus, J. A. Bingler, M. Leippold, T. Schimanski, C. C . Senni, D. Stammbach, S. A. V aghefi, and N. Webersinke, “Enha ncing large language models with climate resources,” arXiv preprint arXiv:2304.00116 , 2023

  49. [57]

    Better patching using llm prom pting, via self-consistency,

    T. Ahmed and P . Devanbu, “Better patching using llm prom pting, via self-consistency,” in 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 2023, pp. 1742–1746

  50. [58]

    Chain-of-thought prompting elicits reasoning in large l anguage models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q . V . Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large l anguage models,” Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022

  51. [59]

    How can we kno w when language models know? on the calibration of language m odels for question answering,

    Z. Jiang, J. Araki, H. Ding, and G. Neubig, “How can we kno w when language models know? on the calibration of language m odels for question answering,” Transactions of the Association for Computational Linguis tics, vol. 9, pp. 962–977, 2021

  52. [60]

    ”why should I t rust you?

    M. T. Ribeiro, S. Singh, and C. Guestrin, “”why should I t rust you?”: Explaining the predictions of any classifier,” i n Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery an d Data Mining, San Francisco, CA, USA, August 13-17, 2016 , 2016, pp. ...

  53. [61]

    A unified approach to inter preting model predictions,

    S. M. Lundberg and S.-I. Lee, “A unified approach to inter preting model predictions,” in Advances in Neural Information Processing Systems 30 , I. Guyon, U. V . Luxburg, S. Bengio, H. Wallach, R. Fergus, S. V ishwanathan, and R. Garnett, Eds. Curran Associates, Inc., 2017, pp....

  54. [62]

    Actionable auditing: Inv estigating the impact of publicly naming biased performanc e results of commercial ai products,

    I. D. Raji and J. Buolamwini, “Actionable auditing: Inv estigating the impact of publicly naming biased performanc e results of commercial ai products,” in Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, a nd Society , 2019, pp. 429–435

  55. [63]

    On responsible machine learning d atasets with fairness, privacy, and regulatory norms,

    S. Mittal, K. Thakral, R. Singh, M. V atsa, T. Glaser, C. C . Ferrer, and T. Hassner, “On responsible machine learning d atasets with fairness, privacy, and regulatory norms,” arXiv preprint arXiv: 2310.15848 , 2023

  56. [64]

    Privacy a nd fairness in federated learning: on the perspective of tra deoff,

    H. Chen, T. Zhu, T. Zhang, W. Zhou, and P . S. Y u, “Privacy a nd fairness in federated learning: on the perspective of tra deoff,” ACM Computing Surveys, vol. 56, no. 2, pp. 1–37, 2023

  57. [65]

    Low-cost high-po wer membership inference attacks,

    S. Zarifzadeh, P . Liu, and R. Shokri, “Low-cost high-po wer membership inference attacks,” in F orty-first International Conference on Machine Learning, 2024. [Online]. Available: https://openreview.net/for um?id=sT7UJh5CTc 14

  58. [66]

    Membership inf erence attacks and privacy in topic modeling,

    N. Manzonelli, W. Zhang, and S. V adhan, “Membership inf erence attacks and privacy in topic modeling,” arXiv preprint arXiv:2403.04451 , 2024

  59. [67]

    Membership infer ence attacks against in-context learning,

    R. Wen, Z. Li, M. Backes, and Y . Zhang, “Membership infer ence attacks against in-context learning,” arXiv preprint arXiv:2409.01380 , 2024

  60. [68]

    Pr ioritizing safeguarding over autonomy: Risks of llm agents for science,

    X. Tang, Q. Jin, K. Zhu, T. Y uan, Y . Zhang, W. Zhou, M. Qu, Y . Zhao, J. Tang, Z. Zhang, A. Cohan, Z. Lu, and M. Gerstein, “Pr ioritizing safeguarding over autonomy: Risks of llm agents for science,” arXiv preprint arXiv: 2402.04247 , 2024

  61. [69]

    Don’t stop pretraining : Adapt language models to domains and tasks,

    S. Gururangan, A. Marasovi´ c, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith, “Don’t stop pretraining : Adapt language models to domains and tasks,” arXiv preprint arXiv:2004.10964 , 2020

  62. [70]

    Enhancing llm factual accu racy with rag to counter hallucinations: A case study on doma in-specific queries in private knowledge-bases,

    J. Li, Y . Y uan, and Z. Zhang, “Enhancing llm factual accu racy with rag to counter hallucinations: A case study on doma in-specific queries in private knowledge-bases,” arXiv preprint arXiv: 2403.10446 , 2024

  63. [71]

    Learn to refuse: Making large language models m ore controllable and reliable through knowledge scope limi tation and refusal mechanism,

    L. Cao, “Learn to refuse: Making large language models m ore controllable and reliable through knowledge scope limi tation and refusal mechanism,” arXiv preprint arXiv: 2311.01041 , 2023

  64. [72]

    A comprehensive survey of hallucin ation mitigation techniques in large language models,

    S. M. T. I. Tonmoy, S. M. M. Zaman, V . Jain, A. Rani, V . Rawt e, A. Chadha, and A. Das, “A comprehensive survey of hallucin ation mitigation techniques in large language models,” arXiv preprint arXiv: 2401.01313 , 2024

  65. [73]

    Control ris k for potential misuse of artificial intelligence in science,

    J. He, W. Feng, Y . Min, J. Yi, K. Tang, S. Li, J. Zhang, K. Ch en, W. Zhou, X. Xie, W. Zhang, N. Y u, and S. Zheng, “Control ris k for potential misuse of artificial intelligence in science,” arXiv preprint arXiv: 2312.06632 , 2023

  66. [74]

    Compliance d isengagement in research: Development and validation of a n ew measure,

    J. M. DuBois, J. T. Chibnall, and J. Gibbs, “Compliance d isengagement in research: Development and validation of a n ew measure,” Science and engineering ethics , vol. 22, pp. 965–988, 2016

  67. [75]

    Guidance for researchers and peer-review ers on the ethical use of large language models (llms) in scie ntific research workflows,

    R. Watkins, “Guidance for researchers and peer-review ers on the ethical use of large language models (llms) in scie ntific research workflows,” AI and Ethics, pp. 1–6, 2023

  68. [76]

    The ethics of using artific ial intelligence in scientific research: new guidance neede d for a new tool,

    D. B. Resnik and M. Hosseini, “The ethics of using artific ial intelligence in scientific research: new guidance neede d for a new tool,” AI and Ethics , pp. 1–23, 2024

  69. [77]

    F. L. Macrina, Scientific integrity: Text and cases in responsible conduct of research. John Wiley & Sons, 2014

  70. [78]

    N. A. of Sciences, Policy, G. Affairs, and C. on Responsi ble Science, F ostering integrity in research. National Academies Press, 2017

  71. [79]

    Citation: A key to building respo nsible and accountable large language models,

    J. Huang and K. Chang, “Citation: A key to building respo nsible and accountable large language models,” NAACL-HLT, 2023

  72. [80]

    Dyknow:dyn amically verifying time-sensitive factual knowledge in ll ms,

    S. M. Mousavi, S. Alghisi, and G. Riccardi, “Dyknow:dyn amically verifying time-sensitive factual knowledge in ll ms,” arXiv preprint arXiv: 2404.08700, 2024

  73. [81]

    User-llm: Efficient llm cont extualization with user embeddings,

    L. Ning, L. Liu, J. Wu, N. Wu, D. Berlowitz, S. Prakash, B. Green, S. O’Banion, and J. Xie, “User-llm: Efficient llm cont extualization with user embeddings,” arXiv preprint arXiv:2402.13598 , 2024

  74. [82]

    A novel nih rese arch grant recommender using bert,

    J. Zhu, B. G. Patra, H. Wu, and A. Y aseen, “A novel nih rese arch grant recommender using bert,” PloS one , vol. 18, no. 1, p. e0278636, 2023

  75. [83]

    A survey on large lan guage models for recommendation,

    L. Wu, Z. Zheng, Z. Qiu, H. Wang, H. Gu, T. Shen, C. Qin, C. Z hu, H. Zhu, Q. Liu, H. Xiong, and E. Chen, “A survey on large lan guage models for recommendation,” W orld Wide W eb, vol. 27, 2024. [Online]. Available: https://link.spring er.com/article/10.1007/s11280-024-01291-...

  76. [84]

    A unified framework of five princ iples for ai in society,

    L. Floridi and J. Cowls, “A unified framework of five princ iples for ai in society,” Machine learning and the city: Applications in architectur e and urban design , pp. 535–545, 2022

  77. [85]

    Rarr: Researching and revising what language models say, using language models,

    L. Gao, Z. Dai, P . Pasupat, A. Chen, A. T. Chaganty, Y . Fan , V . Y . Zhao, N. Lao, H. Lee, D.-C. Juan et al. , “Rarr: Researching and revising what language models say, using language models,” arXiv preprint arXiv:2210.08726 , 2022

  78. [86]

    Attributed question answering: Evaluation and modeling for attributed large la nguage models,

    B. Bohnet, V . Q. Tran, P . V erga, R. Aharoni, D. Andor, L. B . Soares, M. Ciaramita, J. Eisenstein, K. Ganchev, J. Herzig et al. , “Attributed question answering: Evaluation and modeling for attributed large la nguage models,” arXiv preprint arXiv:2212.08037 , 2022

  79. [87]

    Is your llm outdated? evaluating llms at temporal generaliz ation,

    C. Zhu, N. Chen, Y . Gao, Y . Zhang, P . Tiwari, and B. Wang, “ Is your llm outdated? evaluating llms at temporal generaliz ation,” arXiv preprint arXiv: 2405.08460, 2024

  80. [88]

    Retrieval-augmented generation for knowledge-intensive nlp tasks,

    P . Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. K¨ uttler, M. Lewis, W.-t. Yih, T. Rockt¨ aschelet al., “Retrieval-augmented generation for knowledge-intensive nlp tasks,” Advances in Neural Information Processing Systems , vol. 33, pp. 9459–9474, 2020

  81. [89]

    Updating knowledge in large language models: an empirica l evaluation,

    M. A. Roberto, C. Antonio et al. , “Updating knowledge in large language models: an empirica l evaluation,” in Conference Proceedings: 2024 IEEE International Conference on Evolving and Adaptive Intelli gent Systems (EAIS) , 2024

  82. [90]

    Disinformation capabilities of large lang uage models,

    I. Vykopal, M. Pikuliak, I. Srba, R. M´ oro, D. Macko, and M. Bielikov´ a, “Disinformation capabilities of large lang uage models,” Annual Meeting of the Association for Computational Linguistics , 2023

  83. [91]

    Chameleon: Plug-and-play compositi onal reasoning with large language models,

    P . Lu, B. Peng, H. Cheng, M. Galley, K.-W. Chang, Y . N. Wu, S.-C. Zhu, and J. Gao, “Chameleon: Plug-and-play compositi onal reasoning with large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024

  84. [92]

    Toward adaptive reasoning in large la nguage models with thought rollback,

    S. Chen and B. Li, “Toward adaptive reasoning in large la nguage models with thought rollback,” in F orty-first International Conference on Machine Learning, 2024

  85. [93]

    Let’s sample ste p by step: Adaptive-consistency for efficient reasoning and coding with llms,

    A. M. P . Aggarwal, Y . Y ang, and Mausam, “Let’s sample ste p by step: Adaptive-consistency for efficient reasoning and coding with llms,” in Proceedings of the 2023 Conference on Empirical Methods in N atural Language Processing, EMNLP 2023, Singapore, Decemb er 6-10, 2023, H...

  86. [94]

    Do llms exhibit human-like response biases? a case study in survey design,

    L. Tjuatja, V . Chen, T. Wu, A. Talwalkwar, and G. Neubig, “Do llms exhibit human-like response biases? a case study in survey design,” Transactions of the Association for Computational Linguistics , vol. 12, pp. 1011–1026, 2024

  87. [95]

    Mea suring implicit bias in explicitly unbiased large language models,

    X. Bai, A. Wang, I. Sucholutsky, and T. L. Griffiths, “Mea suring implicit bias in explicitly unbiased large language models,” arXiv preprint arXiv: 2402.04105, 2024

  88. [96]

    Knowledge conflicts for llms: A survey,

    R. Xu, Z. Qi, Z. Guo, C. Wang, H. Wang, Y . Zhang, and W. Xu, “ Knowledge conflicts for llms: A survey,” arXiv preprint arXiv: 2403.08319 , 2024

  89. [97]

    Chain-of-verification reduces hallucination in large language models,

    S. Dhuliawala, M. Komeili, J. Xu, R. Raileanu, X. Li, A. C elikyilmaz, and J. Weston, “Chain-of-verification reduces hallucination in large language models,” arXiv preprint arXiv:2309.11495 , 2023

  90. [98]

    V erify-an d-edit: A knowledge-enhanced chain-of-thought framework ,

    R. Zhao, X. Li, S. R. Joty, C. Qin, and L. Bing, “V erify-an d-edit: A knowledge-enhanced chain-of-thought framework ,” Annual Meeting of the Association for Computational Linguistics , 2023

  91. [99]

    Untangle the KNOT: Interweaving conflicting knowledge an d reasoning skills in large language models,

    Y . Liu, Z. Y ao, X. Lv, Y . Fan, S. Cao, J. Y u, L. Hou, and J. Li , “Untangle the KNOT: Interweaving conflicting knowledge an d reasoning skills in large language models,” in Proceedings of the 2024 Joint International Conference on C omputational Linguistics, Language Resour...

  92. [100]

    Factool: Factuality detectio n in generative ai - a tool augmented framework for multi-task and multi-domain scena rios,

    I.-C. Chern, S. Chern, S. Chen, W. Y uan, K. Feng, C. Zhou , J. He, G. Neubig, and P . Liu, “Factool: Factuality detectio n in generative ai - a tool augmented framework for multi-task and multi-domain scena rios,” arXiv preprint arXiv: 2307.13528 , 2023

  93. [101]

    Artificial intelligence’s fair use crisi s,

    B. L. Sobel, “Artificial intelligence’s fair use crisi s,” Colum. JL & Arts , vol. 41, p. 45, 2017

  94. [102]

    Good models borrow, great models stea l: intellectual property rights and generative ai,

    S. Chesterman, “Good models borrow, great models stea l: intellectual property rights and generative ai,” Policy and Society , p. puae006, 2024

  95. [103]

    Unraveling the copyright conundru m: Exploring ai-generated content and its implications for intellectual property rights,

    I. Abdikhakimov, “Unraveling the copyright conundru m: Exploring ai-generated content and its implications for intellectual property rights,” in International Conference on Legal Sciences , vol. 1, no. 5, 2023, pp. 18–32

  96. [104]

    Training language models to follow instructions with human feedback,

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P . Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022

  97. [105]

    Direct preference optimization: Y our la nguage model is secretly a reward model,

    R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. E rmon, and C. Finn, “Direct preference optimization: Y our la nguage model is secretly a reward model,” Advances in Neural Information Processing Systems , vol. 36, 2024

  98. [106]

    Fast model editing at scale,

    E. Mitchell, C. Lin, A. Bosselut, C. Finn, and C. D. Mann ing, “Fast model editing at scale,” arXiv preprint arXiv:2110.11309 , 2021

  99. [107]

    Memoriz ing transformers,

    Y . Wu, M. N. Rabe, D. Hutchins, and C. Szegedy, “Memoriz ing transformers,” arXiv preprint arXiv:2203.08913 , 2022. 15

  100. [108]

    Glam: Efficient scaling of language models with mixture-of-experts,

    N. Du, Y . Huang, A. M. Dai, S. Tong, D. Lepikhin, Y . Xu, M. Krikun, Y . Zhou, A. W. Y u, O. Firat et al. , “Glam: Efficient scaling of language models with mixture-of-experts,” in International Conference on Machine Learning . PMLR, 2022, pp. 5547–5569

  101. [109]

    Improving language models by retrieving from trill ions of tokens,

    S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherf ord, K. Millican, G. B. V an Den Driessche, J.-B. Lespiau, B. D amoc, A. Clark, D. De Las Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Hua ng, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini , G. Irving, O...

  102. [110]

    A survey of safety and trustworthiness of large language models through the lens of verification and validat ion,

    X. Huang, W. Ruan, W. Huang, G. Jin, Y . Dong, C. Wu, S. Ben salem, R. Mu, Y . Qi, X. Zhao et al. , “A survey of safety and trustworthiness of large language models through the lens of verification and validat ion,” Artificial Intelligence Review , vol. 57, no. 7, p. 175, 2024

  103. [111]

    Training a helpful and harmless assistant with reinforcement learning from human feedback ,

    Y . Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasS arma, D. Drain, S. Fort, D. Ganguli, T. Henighan et al. , “Training a helpful and harmless assistant with reinforcement learning from human feedback ,” arXiv preprint arXiv:2204.05862 , 2022

  104. [112]

    Recipes for safety in open-domain chatbots,

    J. Xu, D. Ju, M. Li, Y .-L. Boureau, J. Weston, and E. Dina n, “Recipes for safety in open-domain chatbots,” arXiv preprint arXiv:2010.07079 , 2020

  105. [113]

    Google fact check explor er - recent claims,

    Google Fact Check Explorer, “Google fact check explor er - recent claims,” https://toolbox.google.com/factche ck/ , 2024, accessed: 2024-10-01

  106. [114]

    Generative large language models in automated fact -checking: A survey,

    I. Vykopal, M. Pikuliak, S. Ostermann, and M. ˇSimko, “Generative large language models in automated fact -checking: A survey,” 2024. [Online]. Available: https://arxiv.org/abs/2407.02351

  107. [115]

    Human-ai collaboration via conditional delegati on: A case study of content moderation,

    V . Lai, S. Carton, R. Bhatnagar, Q. V . Liao, Y . Zhang, an d C. Tan, “Human-ai collaboration via conditional delegati on: A case study of content moderation,” in Proceedings of the 2022 CHI Conference on Human Factors in Co mputing Systems , 2022, pp. 1–18

  108. [116]

    Watch you r language: Investigating content moderation with large la nguage models,

    D. Kumar, Y . A. AbuHashem, and Z. Durumeric, “Watch you r language: Investigating content moderation with large la nguage models,” in Proceedings of the International AAAI Conference on W eb and Social Media , vol. 18, 2024, pp. 865–878

  109. [117]

    Judging llm-as-a-judge with mt-bench and chatbot arena,

    L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zh uang, Z. Lin, Z. Li, D. Li, E. Xing et al. , “Judging llm-as-a-judge with mt-bench and chatbot arena,” Advances in Neural Information Processing Systems , vol. 36, pp. 46 595–46 623, 2023

  110. [118]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805 , 2018

  111. [119]

    Roberta: A robu stly optimized bert pretraining approach,

    Y . Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy , M. Lewis, L. Zettlemoyer, and V . Stoyanov, “Roberta: A robu stly optimized bert pretraining approach,” arXiv preprint arXiv: 1907.11692 , 2019

  112. [120]

    Black-box prompt optimization: Aligning l arge language models without model training,

    J. Cheng, X. Liu, K. Zheng, P . Ke, H. Wang, Y . Dong, J. Tan g, and M. Huang, “Black-box prompt optimization: Aligning l arge language models without model training,” in Proceedings of the 62nd Annual Meeting of the Association fo r Computational Linguistics (V olume 1: Long ...

  113. [121]

    Improving language models by retrieving from trillions of tokens,

    S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherf ord, K. Millican, G. B. V an Den Driessche, J.-B. Lespiau, B. D amoc, A. Clark et al. , “Improving language models by retrieving from trillions of tokens,” in International conference on machine learning . PMLR, 2022, pp....

  114. [122]

    Si mple and scalable predictive uncertainty estimation using deep ensembles,

    B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Si mple and scalable predictive uncertainty estimation using deep ensembles,” Advances in neural information processing systems , vol. 30, 2017

  115. [123]

    Deep reinforcement learning from human pre ferences,

    P . F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg , and D. Amodei, “Deep reinforcement learning from human pre ferences,” Advances in neural information processing systems , vol. 30, 2017

  116. [124]

    Big self-supervised models are strong semi-super vised learners,

    T. Chen, S. Kornblith, K. Swersky, M. Norouzi, and G. E. Hinton, “Big self-supervised models are strong semi-super vised learners,” Advances in neural information processing systems , vol. 33, pp. 22 243–22 255, 2020

  117. [125]

    Fixmatch: Simplifying semi-supervised learning with consistency and confidence,

    K. Sohn, D. Berthelot, N. Carlini, Z. Zhang, H. Zhang, C . A. Raffel, E. D. Cubuk, A. Kurakin, and C.-L. Li, “Fixmatch: Simplifying semi-supervised learning with consistency and confidence,” Advances in neural information processing systems , vol. 33, pp. 596–608, 2020

  118. [126]

    Llm-blender: Ensembli ng large language models with pairwise ranking and generati ve fusion,

    D. Jiang, X. Ren, and B. Y . Lin, “Llm-blender: Ensembli ng large language models with pairwise ranking and generati ve fusion,” arXiv preprint arXiv:2306.02561, 2023

  119. [127]

    From local to global: A graph rag a pproach to query-focused summarization,

    D. Edge, H. Trinh, N. Cheng, J. Bradley, A. Chao, A. Mody , S. Truitt, and J. Larson, “From local to global: A graph rag a pproach to query-focused summarization,” arXiv preprint arXiv:2404.16130 , 2024

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.