Pith. sign in

REVIEW 3 major objections 8 minor 300 references

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead

T0 review · 3 major / 8 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that LLM development is best understood as a software-engineering problem spanning six lifecycle phases, and that treating it this way reveals where established SE practice transfers and where new methods are needed.

desk verdict A useful six-phase organizing survey whose 'first comprehensive' claim outruns its undocumented methodology; referee it, but require the novelty claim to be fixed. read the letter →

arxiv 2506.23762 v1 pith:LHAE4NGN submitted 2025-06-30 cs.SE cs.AI

classification cs.SEcs.AI
keywords largelanguagemodelssoftwareengineeringLLMlifecyclerequirementsdatasetconstructiontestingandevaluationdeploymentoperationsmaintenanceevolution
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to establish that LLM development is, at bottom, a software-engineering problem, and that the field is ready for a life-cycle-wide treatment. It divides the LLM lifecycle into six phases—requirements engineering, dataset construction, model development and enhancement, testing and evaluation, deployment and operations, and maintenance and evolution—and surveys the research status of each, followed by the key challenges and proposed research directions. The intended payoff is a shared map: researchers and practitioners can see which software-engineering practices transfer directly to LLMs, which need new methods, and where the open problems concentrate. If the framing holds, it would give the community a common structure for comparing results across phases and for steering future research toward the least-engineered parts of the lifecycle.

What carries the argument

The load-bearing object is the six-phase lifecycle taxonomy, which splits LLM engineering into requirements engineering, dataset construction, model development and enhancement, testing and evaluation, deployment and operations, and maintenance and evolution. The taxonomy is what turns a large collection of related papers into a structured argument: every phase receives a research-status review, a challenge list, and a 'road ahead' section, and the phases are compared against traditional software on dimensions such as determinism, testability, maintainability, and deployability to decide where software-engineering practice transfers and where it must be invented.

What would settle it

Re-running the paper's literature collection with a documented search protocol would settle the core claim: if a peer-reviewed survey already covers all six lifecycle phases through a software-engineering lens, or if a phase contains a substantial body of software-engineering-for-LLM work that the paper never mentions, the 'first comprehensive' claim and the roadmap's completeness would be refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that no existing research has systematically analyzed the LLM development lifecycle through the lens of software engineering, and that doing so yields a usable six-phase organization: requirements engineering, dataset construction, model development and enhancement, testing and evaluation, deployment and operations, and maintenance and evolution. For each phase the paper catalogs current research, names the dominant challenges (for example, defining requirements for non-deterministic systems, data poisoning and authorization, catastrophic forgetting, benchmark contamination, deployment latency and energy costs, and technical debt and model drift), and proposes directions it believes are promising. It also argues that LLMs sit between traditional software and classical machine-learning systems—they keep maintainability, reusability, and scalability characteristics but lose determinism and easy testability—so software-engineering methods need selective adaptation rather than wholesale transfer. The paper positions itself as the first comprehensive survey of this intersection.

Load-bearing premise

The value of the survey depends on its claim to be the first comprehensive software-engineering treatment of the whole LLM lifecycle, but it does not describe how the reviewed papers were searched for, selected, or screened, so the reader cannot verify that the coverage is systematic or unbiased.

Editorial extensions

If this is right

  • If the six-phase framing is adopted, requirements engineering for LLMs must move beyond treating LLMs as components of AI systems and start defining measurable functional and non-functional requirements for models themselves, including accuracy, latency, explainability, and acceptable trade-offs.
  • Dataset construction becomes an identifiable engineering discipline with its own open problems—poisoning defense, provenance, membership inference, watermarking, and bias-aware diversity scoring—rather than a data-cleaning side task.
  • Testing and evaluation of LLMs needs dedicated infrastructure: contamination-resistant benchmarks, adaptive benchmark management, and robust LLM-as-a-judge frameworks with explicit bias mitigation.
  • Deployment and operations split into cluster, edge, and hybrid modes, so research should target each mode's specific constraints—scheduling and energy in clusters, compression and platform heterogeneity on edge devices, and communication and data security in edge-cloud collaboration.
  • Maintenance and evolution deserves systematic study in its own right, with tools for technical-debt measurement, drift detection beyond full retraining, and automated ethical and legal compliance during model updates.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the lifecycle taxonomy catches on, it becomes a shared coordinate system: any new LLM technique could be classified by phase, and progress across phases could be compared directly; the paper stops short of proposing such a classification scheme for future work.
  • Because the paper repeatedly stresses LLMs' probabilistic outputs, it points toward a research program it does not name: adapting classical statistical and metamorphic testing techniques to model behavior, which would give the testing and evaluation phase a firmer technical foundation.
  • The same six-phase lens could be applied specifically to LLM agents and multi-agent systems, where tool integration, inter-agent protocols, and routing cut across all six phases; that application is left for readers to develop.
  • A testable extension follows from the weakest part of the evidence: re-running the survey with an explicit, reproducible search protocol would show whether the six phase-challenge lists are complete or whether some phases are sparse only because of how the literature was gathered.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. This is a narrative survey of software-engineering (SE) challenges across the full lifecycle of large language models (LLMs). The lifecycle is partitioned into six phases—requirements engineering, dataset construction, model development and enhancement, testing and evaluation, deployment and operations, and maintenance and evolution—and each phase is treated in three parts: research status, key challenges, and future research directions. The paper draws on roughly 476 references, weighted toward 2023–2025 work, includes recent ecosystem developments such as MCP and A2A protocols, and claims in the Abstract, §1, and §9 that it is the first comprehensive survey to examine the LLM development lifecycle through the lens of SE. No experiments or derivations are presented; the paper's contribution is entirely organizational and interpretive.

Significance. The paper's intended contribution is organizational: a six-phase lifecycle frame with per-phase challenge and direction summaries, and with security concerns integrated into each phase rather than isolated in a single security chapter. If the coverage claim holds, this is a useful roadmap for a fast-moving area, and the paper is unusually current, discussing MCP, A2A, alignment faking, and LLMOps from 2024–2025 literature. The manuscript deserves credit for explicitly attributing its subsection taxonomies to prior surveys (§5.4, §5.5) and for citing the authors' own earlier SLR rather than hiding it. The significance is nevertheless conditional: the headline novelty claim cannot be verified as submitted, because no literature-search methodology is documented and the related-work section does not perform a scope comparison against adjacent surveys.

major comments (3)
  1. [Abstract, §1, §9] The paper's central claim—that no prior work systematically explores LLM development from a software-engineering perspective and that this is 'the first comprehensive survey'—is a global claim about the published literature, yet the manuscript documents no methodology to support it. There is no description of search strategy, bibliographic databases, time window, inclusion/exclusion criteria, screening process, or quality assessment, so a reader cannot verify that the coverage is comprehensive, representative, or unbiased. This is load-bearing because the novelty claim is stated unqualified in the Abstract ('no existing research systematically explores') and is listed as the first contribution in §1. Please add a methodology subsection documenting the search and screening protocol, or substantially soften the Abstract's unqualified global negative to match the qualified 'to the best of our knowledge' wording already used in §9; ideally, do both and support the choice with a scope-comparison table of adjacent surveys.
  2. [§9] The related-work section inventories existing surveys but never compares their scope against the proposed six-phase taxonomy, so the claimed gap is asserted rather than demonstrated. The relationship to the authors' own prior systematic literature review [126] is dismissed in a single sentence, without stating the governing distinction between LLMs-for-SE (the topic of [126]) and SE-for-LLMs (the topic of this paper). Likewise, the LLMOps lifecycle survey [59] and the RE4AI mapping study [8] are cited, but no explicit differentiation shows why their lifecycle or requirements coverage does not overlap with the deployment/operations and requirements phases of this survey. Please add a structured comparison (e.g., a table mapping each adjacent survey to the six phases) and state precisely which phases and SE aspects are new in this paper.
  3. [§5.4, §5.5] The manuscript honestly attributes its subsection taxonomies ('the same as [29,206,470]' for model compression and 'the same as [114,389]' for PEFT), but this transparency exposes an opacity at the level of the paper's central frame: a reader cannot tell whether the six-phase lifecycle decomposition emerged inductively from the reviewed corpus or was imposed a priori, and the manuscript does not state which is the case. Since the six-phase division is the paper's second stated contribution in §1, its provenance is load-bearing. Please state explicitly how the phases were derived (e.g., inductively from the literature, or adapted from prior lifecycle frameworks such as [59]), and mention any candidate phases that were considered and excluded.
minor comments (8)
  1. [§9] 'we presents the first comprehensive survey' is a grammatical error and should read 'we present the first comprehensive survey.'
  2. [§1] 'Micros ft's DeepSpeed' appears to be a broken rendering of 'Microsoft's DeepSpeed'; please fix the formatting at the line break.
  3. [§5.2.3] 'making it easier for LLMser to switch between' contains the typo 'LLMser'; it should be 'LLM users' (or 'users').
  4. [§7.1.1] 'the imperfect of LLM deployment frameworks introduce further security concerns' is ungrammatical; it should read 'the imperfections of LLM deployment frameworks.'
  5. [§5.5.1] The term 'soft prompt' is used inconsistently ('soft prompt avoid additional structural complexity', 'soft prompt are vulnerable to adversarial attacks'); please use the standard plural 'soft prompts' throughout.
  6. [§2.2.2] The closing sentence of the first paragraph ('Additionally, the development of modular and standardized LLM service interfaces (e.g., OpenAI API [268])...') is nearly duplicated at the end of the next paragraph; please delete one occurrence.
  7. [§1] 'extensive research has explored LLM capabilities [126,380,471]' cites the authors' own SLR [126], which is a survey of LLMs for software engineering rather than of LLM capabilities; please correct the characterization or the citation.
  8. [References] Several factual claims about tool and system behavior rest on non-peer-reviewed preprints (e.g., [123], [284]); for the journal version, please prefer peer-reviewed sources or clearly mark preprint status in the bibliography.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: this is a literature survey with no derivation chain, no fitted parameters, and no predictions; its self-citations are ordinary evidence, and the unsupported 'first comprehensive survey' claim is a verifiability issue, not a circular one.

full rationale

This paper is a survey, not a derivation: it contains no equations, fits no parameters, and makes no quantitative prediction whose output could reduce to its input. The six-phase lifecycle taxonomy (requirements engineering, dataset construction, model development and enhancement, testing and evaluation, deployment and operations, maintenance and evolution) is an organizing frame imposed by the authors, not a result derived from their own prior work; the paper never claims to derive the phases from anything, so there is no self-definitional step. Where sub-taxonomies are adopted, they are explicitly borrowed from external surveys ('We categorize it into three main approaches: quantization, KD, and pruning, the same as [29, 206, 470]' in §5.4; 'We categorize PEFT methods into three main types: additive, reparameterized, and selective approaches, the same as [114, 389]' in §5.5), which is ordinary citation of prior work, not circular import. The paper's self-citations ([125] on LLM deployment framework vulnerabilities, [126] on an SLR of LLMs for SE, [127] on MCP security) are used as evidence for specific empirical claims and are not the justification for the survey's conclusions or for its taxonomy; none of them is invoked as a uniqueness theorem or as a forced choice, so the 'uniqueness imported from authors' and 'ansatz smuggled via citation' patterns do not apply. The central novelty claim — 'To the best of our knowledge, we presents the first comprehensive survey in this work that examines the LLM development lifecycle through the lens of SE' (§9) — is a global negative about the literature that the paper does not substantiate with a documented search strategy, inclusion criteria, or screening process; the skeptic's load-bearing attack targets exactly this missing methodology. That is a real verifiability and correctness weakness, but it is not circularity: the claim does not reduce to the paper's own definitions, fitted values, or self-citations, and a false or unverified novelty claim is a factual-evidence problem rather than a construction-level equivalence. No step was found in which an output equals an input by construction, so the honest finding is no significant circularity (score 0).

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper is a survey, so it introduces no fitted parameters and no physical or conceptual entities that require independent evidence. The proposed 'intelligent prompt and secure framework' in Section 5.3.3 is a suggested future direction, not a postulated entity. The main implicit assumptions are the validity of the six-phase taxonomy, the representativeness of the literature, and the transferability of SE principles.

assumptions (3)
  • ad hoc to paper The six-phase lifecycle taxonomy (requirements, dataset, development, testing, deployment, maintenance) is a valid and complete decomposition of LLM development from an SE perspective.
    The paper asserts this division in Section 1 and Figure 1 without deriving it from a corpus analysis or a prior framework. The taxonomy is the paper's own organizational scheme.
  • domain assumption The cited literature is representative of the field despite the absence of a documented search and selection method.
    No methodology section describes search queries, inclusion criteria, or screening. The survey's conclusions about research status rest on this unstated assumption.
  • domain assumption Traditional software engineering principles can be transferred to LLMs despite their probabilistic and opaque nature.
    The entire paper frames LLM development as amenable to SE methods, based on the comparison table in Section 2 showing both similarities and differences. This assumption is asserted rather than proven.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead." pith.science (2026). https://pith.science/paper/LHAE4NGN

@misc{pith2026250623762,
  author       = {Pith},
  title        = {Pith review of: Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LHAE4NGN}},
  note         = {Machine review of arXiv:2506.23762}
}
read the original abstract

The rapid advancement of large language models (LLMs) has redefined artificial intelligence (AI), pushing the boundaries of AI research and enabling unbounded possibilities for both academia and the industry. However, LLM development faces increasingly complex challenges throughout its lifecycle, yet no existing research systematically explores these challenges and solutions from the perspective of software engineering (SE) approaches. To fill the gap, we systematically analyze research status throughout the LLM development lifecycle, divided into six phases: requirements engineering, dataset construction, model development and enhancement, testing and evaluation, deployment and operations, and maintenance and evolution. We then conclude by identifying the key challenges for each phase and presenting potential research directions to address these challenges. In general, we provide valuable insights from an SE perspective to facilitate future advances in LLM development.

Figures

Figures reproduced from arXiv: 2506.23762 by the authors.

Figure 1
Figure 1. Phase Organization Overview for the LLM Development Lifecycle. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Detailed Activity Breakdown for the LLM Development Lifecycle. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Challenges and Road Ahead in §3 Requirements Engineering. 3.1 Research Status To the best of our knowledge, research on RE for LLMs remains relatively limited. Due to the strong natural language processing (NLP) capabilities of LLMs, existing studies primarily focus on leverag￾ing LLMs to support RE tasks, while comparatively fewer efforts investigate RE methodologies for LLM development itself. This imbalance mirro… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Challenges and Road Ahead in §4 Dataset Construction. According to the construction method, we broadly categorize current approaches to improving data quality into two main strategies: (a) manual data labeling and rule-based selection, and (b) LLM￾assisted data constru…
Figure 5
Figure 5. Figure 5: Challenges and Road Ahead in §5 Development and Enhancement. effectiveness of pre-training has been a long-standing research focus, with efforts directed towards optimizing training datasets, refining training methodologies, and improving computational effi￾ciency [197…
Figure 6
Figure 6. Figure 6: Overlap between LLM-based agents, multimodal models, and multi-model collaboration. [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: Challenges and Road Ahead in §6 Testing and Evaluation. , Vol. 1, No. 1, Article . Publication date: August 2025 [PITH_FULL_IMAGE:figures/full_fig_p028_7.png]
Figure 8
Figure 8. Figure 8: Challenges and Road Ahead in §7 Deployment and Operations. 7.1 Cluster Deployment 7.1.1 Research Status. Cluster deployment refers to deploying models in high-performance com￾puting clusters, such as the cloud, to leverage distributed computing for large-scale inferenc…
Figure 9
Figure 9. Figure 9: Challenges and Road Ahead in §8 Maintenance and Evolution. The rapid advancement of LLMs often leads to the accumulation of technical debt, as ad￾hoc solutions (e.g., memory management, model compression, and attention optimization) are implemented to address short-ter…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

300 extracted references · 8 canonical work pages

  1. [126]

    Soka Hisaharo, Yuki Nishimura, and Aoi Takahashi. 2024. Optimizing llm inference clusters for enhanced performance and energy efficiency. Authorea Preprints (2024)

  2. [78]

    Hugging Face. 2023. PEFT: Parameter-Efficient Fine-Tuning. https://github.com/huggingface/peft

  3. [210]

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. DeepSeek-V3 Technical Report. arXiv preprint arXiv:2412.19437 (2024)

  4. [59]

    Josu Diaz-De-Arcaya, Juan López-De-Armentia, Raúl Miñón, Iker Lasa Ojanguren, and Ana I Torre-Bastida. 2024. Large Language Model Operations (LLMOps): Definition, Challenges, and Lifecycle Management. In 2024 9th International Conference on Smart and Sustainable Technologies (SpliTech) . IEEE, 1–4

  5. [8]

    Khlood Ahmad, Mohamed Abdelrazek, Chetan Arora, Muneera Bano, and John Grundy. 2023. Requirements engi- neering for artificial intelligence systems: A systematic mapping study. Information and Software Technology 158 (2023), 107176

  6. [1]

    [n. d.]. MLflow Model Registry. https://mlflow.org/docs/latest/model-registry/. Accessed: 2025-04-09

  7. [2]

    Sahar Abdelnabi, Aideen Fay, Giovanni Cherubin, Ahmed Salem, Mario Fritz, and Andrew Paverd. 2024. Are you still on track!? Catching LLM Task Drift with Activations. arXiv preprint arXiv:2406.00799 (2024)

  8. [3]

    Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. 2024. On-policy distillation of language models: Learning from self-generated mistakes. In The Twelfth International Conference on Learning Representations

Show all 300 references
  1. [4]

    Agent Communication Protocol Project. 2025. Introduction to the Agent Communication Protocol. https:// agentcommunicationprotocol.dev/introduction/welcome Accessed: 2025-05-12

  2. [5]

    Agent Network Protocol Project. 2025. Agent Network Protocol (ANP) Official Website. https://agent-network- protocol.com/ Accessed: 2025-05-12

  3. [6]

    Ahmed Agiza, Marina Neseem, and Sherief Reda. 2024. Mtlora: Low-rank adaptation approach for efficient multi-task learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 16196–16205

  4. [7]

    Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming{Throughput-Latency} tradeoff in{LLM} inference with{Sarathi-Serve}. In 18th USENIX Symposium on Operating Systems Design and Impleme...

  5. [9]

    Meta AI. 2025. Llama: Open-Source AI Models. https://www.llama.com/

  6. [10]

    Daniel Alexander Alber, Zihao Yang, Anton Alyakin, Eunice Yang, Sumedha Rai, Aly A Valliani, Jeff Zhang, Gabriel R Rosenbaum, Ashley K Amend-Thomas, David B Kurland, et al. 2025. Medical large language models are vulnerable to data-poisoning attacks. Nature Medicine (2025), 1–9

  7. [11]

    Shaden Alshammari, Yu-Xiong Wang, Deva Ramanan, and Shu Kong. 2022. Long-tailed recognition via weight balancing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 6897–6907

  8. [12]

    Maryam Amirizaniani, Elias Martin, Maryna Sivachenko, Afra Mashhadi, and Chirag Shah. 2024. Can llms reason like humans? assessing theory of mind reasoning in llms for open-ended questions. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Ma...

  9. [13]

    Anthropic. 2025. Claude 3.7 Sonnet and Claude Code. https://www.anthropic.com/news/claude-3-7-sonnet Accessed: 2025-04-20

  10. [14]

    Anthropic. 2025. MCP: Agents and Tools Overview. https://docs.anthropic.com/en/docs/agents-and-tools/mcp Accessed: 2025-04-26

  11. [15]

    Daiyaan Arfeen, Zhen Zhang, Xinwei Fu, Gregory R Ganger, and Yida Wang. 2024. PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training. arXiv preprint arXiv:2410.07192 (2024)

  12. [16]

    Guangji Bai, Zheng Chai, Chen Ling, Shiyu Wang, Jiaying Lu, Nan Zhang, Tingwei Shi, Ziyang Yu, Mengdan Zhu, Yifei Zhang, et al. 2024. Beyond efficiency: A systematic survey of resource-efficient large language models. arXiv preprint arXiv:2401.00625 (2024)

  13. [17]

    Nathalie Baracaldo, Bryant Chen, Heiko Ludwig, and Jaehoon Amir Safavi. 2017. Mitigating poisoning attacks on machine learning models: A data provenance based approach. In Proceedings of the 10th ACM workshop on artificial intelligence and security. 103–110

  14. [18]

    Peter J Barclay and Ashkan Sami. 2024. Investigating Markers and Drivers of Gender Bias in Machine Translations. arXiv preprint arXiv:2403.11896 (2024)

  15. [19]

    Yan Cai, Linlin Wang, Ye Wang, Gerard de Melo, Ya Zhang, Yanfeng Wang, and Liang He. 2024. Medbench: A large-scale chinese benchmark for evaluating medical large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 17709–17717

  16. [20]

    Jialun Cao, Zhiyong Chen, Jiarong Wu, Shing-Chi Cheung, and Chang Xu. 2024. JavaBench: A Benchmark of Object- Oriented Code Generation for Evaluating Large Language Models. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 870–882...

  17. [21]

    Jialun Cao, Wuqi Zhang, and Shing-Chi Cheung. 2024. Concerned with Data Contamination? Assessing Countermea- sures in Code Language Model. arXiv preprint arXiv:2403.16898 (2024)

  18. [22]

    Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, et al. 2022. Multipl-e: A scalable and extensible approach to benchmarking neural code generation. arXiv preprint...

  19. [23]

    David Cerdeira, Nuno Santos, Pedro Fonseca, and Sandro Pinto. 2020. Sok: Understanding the prevailing security vulnerabilities in trustzone-assisted tee systems. In2020 IEEE Symposium on Security and Privacy (SP). IEEE, 1416–1432

  20. [24]

    Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu. 2024. Art or artifice? large language models and the false promise of creativity. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–34

  21. [25]

    Aofei Chang, Jiaqi Wang, Han Liu, Parminder Bhatia, Cao Xiao, Ting Wang, and Fenglong Ma. 2024. BIPEFT: Budget-Guided Iterative Search for Parameter Efficient Fine-Tuning of Large Pretrained Language Models. arXiv preprint arXiv:2410.09079 (2024)

  22. [26]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–45

  23. [27]

    Zachary Charles, Arun Ganesh, Ryan McKenna, H Brendan McMahan, Nicole Mitchell, Krishna Pillutla, and Keith Rush. 2024. Fine-tuning large language models with user-level differential privacy. arXiv preprint arXiv:2407.07737 (2024)

  24. [28]

    Harrison Chase and contributors. 2022. LangChain: Build context-aware reasoning applications. https://github.com/ langchain-ai/langchain

  25. [29]

    Arnav Chavan, Raghav Magazine, Shubham Kushwaha, Mérouane Debbah, and Deepak Gupta. 2024. Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward. arXiv preprint arXiv:2402.01799 (2024)

  26. [30]

    Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher M De Sa. 2024. Quip: 2-bit quantization of large language models with guarantees. Advances in Neural Information Processing Systems 36 (2024)

  27. [31]

    Viktoriia Chekalina, Anna Rudenko, Gleb Mezentsev, Alexander Mikhalev, Alexander Panchenko, and Ivan Oseledets

  28. [32]

    Chaochao Chen, Xiaohua Feng, Jun Zhou, Jianwei Yin, and Xiaolin Zheng. 2023. Federated large language model: A position paper. arXiv e-prints (2023), arXiv–2307

  29. [33]

    Dake Chen, Hanbin Wang, Yunhao Huo, Yuzhao Li, and Haoyang Zhang. 2023. Gamegpt: Multi-agent collaborative framework for game development. arXiv preprint arXiv:2310.08067 (2023)

  30. [34]

    Jieneng Chen, Luoxin Ye, Ju He, Zhaoyang Wang, Daniel Khashabi, and Alan L Yuille. 2025. Efficient large multi-modal models via visual context compression. Advances in Neural Information Processing Systems 37 (2025), 73986–74007

  31. [35]

    Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2025. Sharegpt4v: Improving large multi-modal models with better captions. In European Conference on Computer Vision . Springer, 370–387

  32. [36]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)

  33. [37]

    Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu. 2024. Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models. arXiv:2401.01335 [cs.LG] https://arxiv.org/abs/2401.01335

  34. [38]

    Zitao Chen and Karthik Pattabiraman. 2024. Catch Me if You Can: Detecting Unauthorized Data Use in Deep Learning Models. arXiv preprint arXiv:2409.06280 (2024)

  35. [39]

    Yuheng Cheng, Ceyao Zhang, Zhengwen Zhang, Xiangrui Meng, Sirui Hong, Wenhao Li, Zihao Wang, Zekai Wang, Feng Yin, Junhua Zhao, et al. 2024. Exploring large language model based intelligent agents: Definitions, methods, and prospects. arXiv preprint arXiv:2401.03428 (2024)

  36. [40]

    Wonje Choi, Woo Kyung Kim, Minjong Yoo, and Honguk Woo. 2024. Embodied CoT Distillation From LLM To Off-the-shelf Agents. arXiv:2412.11499 [cs.AI] https://arxiv.org/abs/2412.11499

  37. [41]

    Zhumin Chu, Qingyao Ai, Yiteng Tu, Haitao Li, and Yiqun Liu. 2024. Automatic Large Language Model Evaluation via Peer Review. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management . 384–393

  38. [42]

    John Joon Young Chung, Ece Kamar, and Saleema Amershi. 2023. Increasing diversity while maintaining accuracy: Text data generation with large language models and human interventions. arXiv preprint arXiv:2306.04140 (2023)

  39. [43]

    Woojin Chung, Jiwoo Hong, Na Min An, James Thorne, and Se-Young Yun. 2024. Stable Language Model Pre-training by Reducing Embedding Variability. arXiv preprint arXiv:2409.07787 (2024)

  40. [44]

    Comfy. 2025. ComfyUI. https://www.comfy.org/zh-cn/. Accessed: 2025-05-15. , Vol. 1, No. 1, Article . Publication date: August 2025. Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead 43

  41. [45]

    Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm. Company Blog of Databricks (2023)

  42. [46]

    Domenico Cotroneo, Cristina Improta, Pietro Liguori, and Roberto Natella. 2024. Vulnerabilities in ai code generators: Exploring targeted data poisoning attacks. In Proceedings of the 32nd IEEE/ACM International Conference on Program Comprehension. 280–292

  43. [47]

    Yingqian Cui, Jie Ren, Yuping Lin, Han Xu, Pengfei He, Yue Xing, Lingjuan Lyu, Wenqi Fan, Hui Liu, and Jiliang Tang

  44. [48]

    Sunhao Dai, Chen Xu, Shicheng Xu, Liang Pang, Zhenhua Dong, and Jun Xu. 2024. Bias and unfairness in information retrieval systems: New challenges in the llm era. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 6437–6447

  45. [49]

    Xiangxiang Dai, Jin Li, Xutong Liu, Anqi Yu, and John Lui. 2024. Cost-Effective Online Multi-LLM Selection with Versatile Reward Models. arXiv preprint arXiv:2405.16587 (2024)

  46. [50]

    Arghavan Moradi Dakhel, Amin Nikanjam, Vahid Majdinasab, Foutse Khomh, and Michel C Desmarais. 2024. Effective test generation using pre-trained large language models and mutation testing. Information and Software Technology 171 (2024), 107468

  47. [51]

    Badhan Chandra Das, M Hadi Amini, and Yanzhao Wu. 2025. Security and privacy challenges of large language models: A survey. Comput. Surveys 57, 6 (2025), 1–39

  48. [52]

    Sarkar Snigdha Sarathi Das, Ranran Haoran Zhang, Peng Shi, Wenpeng Yin, and Rui Zhang. 2023. Unified Low- Resource Sequence Labeling by Sample-Aware Dynamic Sparse Finetuning. In Conference on Empirical Methods in Natural Language Processing. https://api.semanticscholar.org/Co...

  49. [53]

    Bibhu Dash. 2024. Zero-Trust Architecture (ZTA): Designing an AI-Powered Cloud Security Framework for LLMs’ Black Box Problems. A vailable at SSRN 4726625 (2024)

  50. [54]

    Daniel DeAlcala, Aythami Morales, Julian Fierrez, Gonzalo Mancera, Ruben Tolosana, and Javier Ortega-Garcia

  51. [55]

    Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. 2025. Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. Advances in Neural Information Processing Systems 37 (2025), 82895–82920

  52. [56]

    Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale. Advances in Neural Information Processing Systems 35 (2022), 30318–30332

  53. [57]

    arXiv preprint arXiv:2402.09225 (2024)

    Is my data in your ai model? membership inference test with application to face images. arXiv preprint arXiv:2402.09225 (2024)

  54. [58]

    Nobel Dhar, Bobin Deng, Dan Lo, Xiaofeng Wu, Liang Zhao, and Kun Suo. 2024. An empirical analysis and resource footprint study of deploying large language models on edge devices. In Proceedings of the 2024 ACM Southeast Conference. 69–76

  55. [60]

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2024. Qlora: Efficient finetuning of quantized llms. Advances in Neural Information Processing Systems 36 (2024)

  56. [61]

    Yangruibo Ding, Zijian Wang, Wasi Ahmad, Hantian Ding, Ming Tan, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, et al. 2023. Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion. Advances in Neural Information...

  57. [62]

    Yangruibo Ding, Zijian Wang, Wasi Ahmad, Hantian Ding, Ming Tan, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, et al. 2024. Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion. Advances in Neural Information...

  58. [63]

    Bosheng Ding, Chengwei Qin, Ruochen Zhao, Tianze Luo, Xinze Li, Guizhen Chen, Wenhan Xia, Junjie Hu, Luu Anh Tuan, and Shafiq Joty. 2024. Data augmentation using llms: Data perspectives, learning paradigms and challenges. In Findings of the Association for Computational Lingui...

  59. [64]

    Shihan Dou, Enyu Zhou, Yan Liu, Songyang Gao, Wei Shen, Limao Xiong, Yuhao Zhou, Xiao Wang, Zhiheng Xi, Xiaoran Fan, et al. 2024. LoRAMoE: Alleviating world knowledge forgetting in large language models via MoE-style plugin. In Proceedings of the 62nd Annual Meeting of the Ass...

  60. [65]

    Wenyu Du, Tongxu Luo, Zihan Qiu, Zeyu Huang, Yikang Shen, Reynold Cheng, Yike Guo, and Jie Fu. 2025. Stacking your transformers: A closer look at model growth for efficient llm pre-training. Advances in Neural Information , Vol. 1, No. 1, Article . Publication date: August 202...

  61. [66]

    Ming Dong, Kang Xue, Bolong Zheng, and Tingting He. 2024. Data-oriented Dynamic Fine-tuning Parameter Selection Strategy for FISH Mask based Efficient Fine-tuning. arXiv preprint arXiv:2403.08484 (2024)

  62. [67]

    Jiangfei Duan, Shuo Zhang, Zerui Wang, Lijuan Jiang, Wenwen Qu, Qinghao Hu, Guoteng Wang, Qizhen Weng, Hang Yan, Xingcheng Zhang, et al. 2024. Efficient training of large language models on distributed infrastructures: a survey. arXiv preprint arXiv:2407.20018 (2024)

  63. [68]

    Yucong Duan. 2024. The Large Language Model (LLM) Bias Evaluation (Age Bias).DIKWP Research Group International Standard Evaluation. DOI 10 (2024)

  64. [69]

    Xueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang, Junwei Liu, Yixuan Chen, Jiayi Feng, Chaofeng Sha, Xin Peng, and Yiling Lou. 2023. Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation. arXiv preprint arXiv:2308.01861 (2023)

  65. [70]

    Lilian Edwards. 2021. The EU AI Act: a summary of its significance and scope. Artificial Intelligence (the EU AI Act) 1 (2021)

  66. [71]

    Abul Ehtesham, Aditi Singh, Gaurav Kumar Gupta, and Saket Kumar. 2025. A survey of agent interoperability protocols: Model Context Protocol (MCP), Agent Communication Protocol (ACP), Agent-to-Agent Protocol (A2A), and Agent Network Protocol (ANP). arXiv preprint arXiv:2505.022...

  67. [72]

    Kristina Dzeparoska, Ali Tizghadam, and Alberto Leon-Garcia. 2024. Intent Assurance using LLMs guided by Intent Drift. arXiv preprint arXiv:2402.00715 (2024)

  68. [73]

    Jonathan Evertz, Merlin Chlosta, Lea Schönherr, and Thorsten Eisenhofer. 2024. Whispers in the Machine: Confiden- tiality in LLM-integrated Systems. arXiv preprint arXiv:2402.06922 (2024)

  69. [74]

    Hugging Face. 2023. Open LLM Leaderboard: Track, rank, and evaluate open LLMs and chatbots. https://huggingface. co/spaces/open-llm-leaderboard/open_llm_leaderboard

  70. [75]

    EleutherAI. 2021. lm-evaluation-harness: A framework for evaluating language models on a wide range of tasks. https://github.com/EleutherAI/lm-evaluation-harness

  71. [76]

    Hugging Face. 2025. Hugging Face Datasets: A Community Library for Datasets. https://huggingface.co/datasets

  72. [77]

    Hugging Face. 2025. Inference Endpoints: Machine Learning At Your Service. https://endpoints.huggingface.co/

  73. [79]

    Alina Fastowski and Gjergji Kasneci. 2024. Understanding Knowledge Drift in LLMs through Misinformation. arXiv preprint arXiv:2409.07085 (2024)

  74. [80]

    Muhammad Fawi. 2024. Curlora: Stable llm continual fine-tuning and catastrophic forgetting mitigation. arXiv preprint arXiv:2408.14572 (2024)

  75. [81]

    Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M Zhang. 2023. Large language models for software engineering: Survey and open problems. In 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineeri...

  76. [82]

    Shangbin Feng, Wenxuan Ding, Alisa Liu, Zifeng Wang, Weijia Shi, Yike Wang, Zejiang Shen, Xiaochuang Han, Hunter Lang, Chen-Yu Lee, et al . 2025. When One LLM Drools, Multi-LLM Collaboration Rules. arXiv preprint arXiv:2502.04506 (2025)

  77. [83]

    Shangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding, Vidhisha Balachandran, and Yulia Tsvetkov. 2024. Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration. arXiv preprint arXiv:2402.00367 (2024)

  78. [84]

    Jia Feng, Jiachen Liu, Cuiyun Gao, Chun Yong Chong, Chaozheng Wang, Shan Gao, and Xin Xia. 2024. Complex- codeeval: A benchmark for evaluating large code models on more complex code. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering...

  79. [85]

    Tao Feng, Lizhen Qu, Niket Tandon, Zhuang Li, Xiaoxi Kang, and Gholamreza Haffari. 2024. From pre-training corpora to large language models: What factors influence llm performance in causal discovery tasks? arXiv preprint arXiv:2407.19638 (2024)

  80. [86]

    Maximilian T Fischer, Yannick Metz, Lucas Joos, Matthias Miller, and Daniel A Keim. 2024. MULTI-CASE: A Transformer-based Ethics-aware Multimodal Investigative Intelligence Framework. arXiv preprint arXiv:2401.01955 (2024)

  81. [87]

    Shangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher, Chan Young Park, Yejin Choi, and Yulia Tsvetkov. 2024. Modular pluralism: Pluralistic alignment via multi-llm collaboration. arXiv preprint arXiv:2406.15951 (2024)

  82. [88]

    Tingchen Fu, Mrinank Sharma, Philip Torr, Shay B Cohen, David Krueger, and Fazl Barez. 2024. PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning. arXiv preprint arXiv:2410.08811 (2024)

  83. [89]

    Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai. 2024. Serverlessllm: Low-latency serverless inference for large language models. In 18th USENIX Symposium on Operating Systems Design and Implementation . USENIX Association, ...

  84. [90]

    Giorgio Franceschelli and Mirco Musolesi. 2024. On the creativity of large language models. AI & SOCIETY (2024), 1–11

  85. [91]

    Gad Gad, Eyad Gad, Zubair Md Fadlullah, Mostafa M Fouda, and Nei Kato. 2024. Communication-efficient and privacy- preserving federated learning via joint knowledge distillation and differential privacy in bandwidth-constrained networks. IEEE Transactions on Vehicular Technology (2024)

  86. [92]

    Kanishk Gandhi, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah Goodman. 2024. Understanding social reasoning in language models with language models. Advances in Neural Information Processing Systems 36 (2024)

  87. [93]

    Gad Gad, Aya Farrag, Ahmed Aboulfotouh, Khaled Bedda, Zubair Md Fadlullah, and Mostafa M Fouda. 2024. Joint self-organizing maps and knowledge-distillation-based communication-efficient federated learning for resource- constrained UAV-IoT systems.IEEE Internet of Things Journa...

  88. [94]

    Shuzheng Gao, Wenxin Mao, Cuiyun Gao, Li Li, Xing Hu, Xin Xia, and Michael R Lyu. 2024. Learning in the wild: Towards leveraging unlabeled data for effectively tuning pre-trained code models. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13

  89. [95]

    Senay A Gebreab, Khaled Salah, Raja Jayaraman, Muhammad Habib ur Rehman, and Samer Ellaham. 2024. LLM-Based Framework for Administrative Task Automation in Healthcare. In 2024 12th International Symposium on Digital Forensics and Security (ISDFS) . IEEE, 1–7

  90. [96]

    Chujie Gao, Dongping Chen, Qihui Zhang, Yue Huang, Yao Wan, and Lichao Sun. 2024. Llm-as-a-coauthor: The challenges of detecting llm-human mixcase. arXiv preprint arXiv:2401.05952 (2024)

  91. [97]

    Linyuan Gong, Sida Wang, Mostafa Elhoushi, and Alvin Cheung. 2024. Evaluation of llms on syntax-aware code fill-in-the-middle tasks. arXiv preprint arXiv:2403.04814 (2024)

  92. [98]

    Zhuocheng Gong, Jiahao Liu, Jingang Wang, Xunliang Cai, Dongyan Zhao, and Rui Yan. 2024. What makes quantiza- tion for large language model hard? an empirical study from the lens of perturbation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 18082–18089

  93. [99]

    Georgi Gerganov and contributors. 2023. llama.cpp: Inference of LLaMA models in pure C/C++. https://github.com/ ggml-org/llama.cpp

  94. [100]

    Sagar Goyal, Eti Rastogi, Sree Prasanna Rajagopal, Dong Yuan, Fen Zhao, Jai Chintagunta, Gautam Naik, and Jeff Ward. 2024. Healai: A healthcare llm for effective medical documentation. InProceedings of the 17th ACM International Conference on Web Search and Data Mining . 1167–1168

  95. [101]

    Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, et al. 2024. Alignment faking in large language models. arXiv preprint arXiv:2412.14093 (2024)

  96. [102]

    Google. 2025. Gemini: Your Personal AI Assistant. https://gemini.google.com/

  97. [103]

    Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. 2024. A survey on llm-as-a-judge. arXiv preprint arXiv:2411.15594 (2024)

  98. [104]

    Jia Gu, Liang Pang, Huawei Shen, and Xueqi Cheng. 2024. Do LLMs Play Dice? Exploring Probability Distribution Sampling in Large Language Models for Behavioral Simulation. arXiv preprint arXiv:2404.09043 (2024)

  99. [105]

    Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Inte...

  100. [106]

    Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zambrano, et al. 2024. Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models. Advances in ...

  101. [107]

    Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence. arXiv preprint arXiv:2401.14196 (2024)

  102. [108]

    Sylvain Gugger, Thomas Wolf, Lysandre Debut, and contributors. 2020. Transformers: State-of-the-art Natural Language Processing. https://github.com/huggingface/transformers

  103. [109]

    Yiduo Guo, Jie Fu, Huishuai Zhang, Dongyan Zhao, and Yikang Shen. 2024. Efficient continual pre-training by mitigating the stability gap. arXiv preprint arXiv:2406.14833 (2024)

  104. [110]

    Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, Deyi Xiong, et al. 2023. Evaluating large language models: A comprehensive survey. arXiv preprint arXiv:2310.19736 (2023)

  105. [111]

    Junfeng Guo, Yiming Li, Ruibo Chen, Yihan Wu, Heng Huang, et al. 2025. ZeroMark: Towards Dataset Ownership Verification without Disclosing Watermark. Advances in Neural Information Processing Systems 37 (2025), 120468– 120500

  106. [112]

    Muhammad Usman Hadi, Rizwan Qureshi, Abbas Shah, Muhammad Irfan, Anas Zafar, Muhammad Bilal Shaikh, Naveed Akhtar, Jia Wu, Seyedali Mirjalili, et al. 2023. A survey on large language models: Applications, challenges, limitations, and practical usage. Authorea Preprints 3 (2023...

  107. [113]

    Songyue Han, Mingyu Wang, Jialong Zhang, Dongdong Li, and Junhong Duan. 2024. A Review of Large Lan- guage Models: Fundamental Architectures, Key Technological Evolutions, Interdisciplinary Technologies Integration, Optimization and Compression Techniques, Applications, and Ch...

  108. [114]

    Akshat Gupta, Anurag Rao, and Gopala Anumanchipalli. 2024. Model editing at scale leads to gradual and catastrophic forgetting. arXiv preprint arXiv:2401.07453 (2024)

  109. [115]

    Zixu Hao, Huiqiang Jiang, Shiqi Jiang, Ju Ren, and Ting Cao. 2024. Hybrid slm and llm for edge-cloud collaborative inference. In Proceedings of the Workshop on Edge and Mobile Foundation Models . 36–41

  110. [116]

    Helia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme, and Chris Kedzie. 2024. LLM-rubric: A multidi- mensional, calibrated approach to automated evaluation of natural language texts. arXiv preprint arXiv:2501.00274 (2024)

  111. [117]

    Zeyu Han, Chao Gao, Jinyang Liu, Jeff Zhang, and Sai Qian Zhang. 2024. Parameter-efficient fine-tuning for large models: A comprehensive survey. arXiv preprint arXiv:2403.14608 (2024)

  112. [118]

    Soufiane Hayou, Nikhil Ghosh, and Bin Yu. 2024. Lora+: Efficient low rank adaptation of large models. arXiv preprint arXiv:2402.12354 (2024)

  113. [119]

    Chaoyang He, Shen Li, Mahdi Soltanolkotabi, and Salman Avestimehr. 2021. PipeTransformer: Automated elastic pipelining for distributed training of large-scale models. In International Conference on Machine Learning . PMLR, 4150–4159

  114. [120]

    Shabnam Hassani, Mehrdad Sabetzadeh, and Daniel Amyot. 2025. An empirical study on LLM-based classification of requirements-related provisions in food-safety regulations. Empirical Software Engineering 30, 3 (2025), 72

  115. [121]

    Ying He, Jingcheng Fang, F Richard Yu, and Victor C Leung. 2024. Large language models (LLMs) inference offloading and resource allocation in cloud-edge computing: An active inference approach.IEEE Transactions on Mobile Computing (2024)

  116. [122]

    Tharindu B Hewage, Shashikant Ilager, Maria Rodriguez Read, and Rajkumar Buyya. 2025. Aging-aware CPU Core Management for Embodied Carbon Amortization in Cloud LLM Inference. arXiv preprint arXiv:2501.15829 (2025)

  117. [123]

    Jiaao He and Jidong Zhai. 2024. FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines. arXiv preprint arXiv:2403.11421 (2024)

  118. [124]

    Guiyang Hou, Yongliang Shen, and Weiming Lu. 2024. Progressive Tuning: Towards Generic Sentiment Abilities for Large Language Models. In Findings of the Association for Computational Linguistics ACL 2024 . 14392–14402

  119. [125]

    Xinyi Hou, Jiahao Han, Yanjie Zhao, and Haoyu Wang. 2025. Unveiling the Landscape of LLM Deployment in the Wild: An Empirical Study. arXiv preprint arXiv:2505.02502 (2025)

  120. [127]

    Xinyi Hou, Yanjie Zhao, Shenao Wang, and Haoyu Wang. 2025. Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions. arXiv preprint arXiv:2503.23278 (2025)

  121. [128]

    Chia-Yi Hsu, Yu-Lin Tsai, Chih-Hsun Lin, Pin-Yu Chen, Chia-Mu Yu, and Chun-Ying Huang. 2024. Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models. arXiv preprint arXiv:2405.16833 (2024)

  122. [129]

    Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024. Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology 33, 8 (2024), 1–79

  123. [130]

    Zhiqiang Hu, Lei Wang, Yihuai Lan, Wanyu Xu, Ee-Peng Lim, Lidong Bing, Xing Xu, Soujanya Poria, and Roy Ka-Wei Lee. 2023. Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models. arXiv preprint arXiv:2304.01933 (2023)

  124. [131]

    Chun-Yin Huang, Kartik Srinivas, Xin Zhang, and Xiaoxiao Li. 2024. Overcoming data and model heterogeneities in decentralized federated learning via synthetic anchors. arXiv preprint arXiv:2405.11525 (2024)

  125. [132]

    Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay. 2024. MARS: A Benchmark for Multi-LLM Algorithmic Routing System. In ICLR 2024 Workshop: How Far Are We From AGI

  126. [133]

    Dong Huang, Jie M Zhang, Michael Luck, Qingwen Bu, Yuhao Qing, and Heming Cui. 2023. Agentcoder: Multi-agent- based code generation with iterative testing and optimisation. arXiv preprint arXiv:2312.13010 (2023)

  127. [134]

    Jianheng Huang, Leyang Cui, Ante Wang, Chengyi Yang, Xinting Liao, Linfeng Song, Junfeng Yao, and Jinsong Su

  128. [135]

    Dong Huang, Guangtao Zeng, Jianbo Dai, Meng Luo, Han Weng, Yuhao Qing, Heming Cui, Zhijiang Guo, and Jie M Zhang. 2024. Effi-Code: Unleashing Code Efficiency in Language Models. arXiv preprint arXiv:2410.10209 (2024)

  129. [136]

    Wei Huang, Xingyu Zheng, Xudong Ma, Haotong Qin, Chengtao Lv, Hong Chen, Jie Luo, Xiaojuan Qi, Xianglong Liu, and Michele Magno. 2024. An empirical study of llama3 quantization: From llms to mllms. Visual Intelligence 2, 1 (2024), 36

  130. [137]

    Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Keming Lu, et al. 2024. Qwen2. 5-coder technical report. arXiv preprint arXiv:2409.12186 (2024)

  131. [138]

    arXiv preprint arXiv:2403.01244 (2024)

    Mitigating catastrophic forgetting in large language models with self-synthesized rehearsal. arXiv preprint arXiv:2403.01244 (2024)

  132. [139]

    Wei Huang, Xudong Ma, Haotong Qin, Xingyu Zheng, Chengtao Lv, Hong Chen, Jie Luo, Xiaojuan Qi, Xianglong Liu, and Michele Magno. 2024. How good are low-bit quantized llama3 models? an empirical study. arXiv e-prints (2024), arXiv–2404. , Vol. 1, No. 1, Article . Publication da...

  133. [140]

    Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, and Irina Rish. 2024. Simple and scalable strategies to continually pre-train large language models. arXiv preprint arXiv:2403.08763 (2024)

  134. [141]

    Intel. 2025. Intel ® Software Guard Extensions (Intel® SGX). https://www.intel.com/content/www/us/en/products/ docs/accelerator-engines/software-guard-extensions.html

  135. [142]

    Bo Hui, Haolin Yuan, Neil Gong, Philippe Burlina, and Yinzhi Cao. 2024. Pleak: Prompt leaking attacks against large language model applications. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 3600–3614

  136. [143]

    Fushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang, and Song Guo. 2024. C2KD: Bridging the Modality Gap for Cross-Modal Knowledge Distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 16006–16015

  137. [144]

    Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. 2024. Beavertails: Towards improved safety alignment of llm via a human-preference dataset.Advances in Neural Information Processing Systems 36 (2024)

  138. [145]

    Bowen Jiang, Yangxinyu Xie, Xiaomeng Wang, Weijie J Su, Camillo Jose Taylor, and Tanwi Mallick. 2024. Multi-modal and multi-agent systems meet rationality: A survey. In ICML 2024 Workshop on LLMs and Cognition

  139. [146]

    Myeongjun Jang, Antonios Georgiadis, Yiyun Zhao, and Fran Silavong. 2024. DriftWatch: A Tool that Automatically Detects Data Drift and Extracts Representative Examples Affected by Drift. In Proceedings of the 2024 Conference of the North American Chapter of the Association for...

  140. [147]

    Hyesung Jeon, Yulhwa Kim, and Jae-joon Kim. 2024. L4q: Parameter efficient quantization-aware training on large language models via lora-wise lsq. arXiv preprint arXiv:2402.04902 (2024)

  141. [148]

    Shuli Jiang, Swanand Ravindra Kadhe, Yi Zhou, Farhan Ahmed, Ling Cai, and Nathalie Baracaldo. 2024. Turning Generative Models Degenerate: The Power of Data Poisoning Attacks. arXiv preprint arXiv:2407.12281 (2024)

  142. [149]

    Xiaotang Jiang, Huan Wang, Yiliu Chen, Ziqi Wu, Lichuan Wang, Bin Zou, Yafeng Yang, Zongyang Cui, Yu Cai, Tianhang Yu, et al. 2020. MNN: A universal and efficient inference engine. Proceedings of Machine Learning and Systems 2 (2020), 1–13

  143. [150]

    Gangwei Jiang, Caigao Jiang, Zhaoyi Li, Siqiao Xue, Jun Zhou, Linqi Song, Defu Lian, and Ying Wei. 2024. Interpretable catastrophic forgetting of large language model fine-tuning via instruction vector. arXiv preprint arXiv:2406.12227 (2024)

  144. [151]

    Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim. 2024. A survey on large language models for code generation. arXiv preprint arXiv:2406.00515 (2024)

  145. [152]

    Lyudong Jin, Yanning Zhang, Yanhan Li, Shurong Wang, Howard H Yang, Jian Wu, and Meng Zhang. 2025. MoE2: Optimizing Collaborative Inference for Edge Large Language Models. arXiv preprint arXiv:2501.09410 (2025)

  146. [153]

    Renren Jin, Jiangcun Du, Wuwei Huang, Wei Liu, Jian Luan, Bin Wang, and Deyi Xiong. 2024. A comprehensive evaluation of quantization strategies for large language models. In Findings of the Association for Computational Linguistics ACL 2024. 12186–12215

  147. [154]

    Junfeng Jiao, Saleh Afroogh, Yiming Xu, and Connor Phillips. 2024. Navigating llm ethics: Advancements, challenges, and future directions. arXiv preprint arXiv:2406.18841 (2024)

  148. [155]

    Hongpeng Jin and Yanzhao Wu. 2024. CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration. arXiv preprint arXiv:2411.02829 (2024)

  149. [156]

    Junmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang, Jacob Hansen, James Glass, David Cox, Rameswar Panda, Rogerio Feris, and Alan Ritter. 2024. Self-moe: Towards compositional large language models with self-specialized experts. arXiv preprint arXiv:2406.12034 (2024)

  150. [157]

    Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al. 2023. ChatGPT for good? On opportunities and challenges of large language models for education. Learning an...

  151. [158]

    Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. 2023. Challenges and applications of large language models. arXiv preprint arXiv:2307.10169 (2023)

  152. [159]

    Damjan Kalajdzievski. 2024. Scaling laws for forgetting when fine-tuning large language models. arXiv preprint arXiv:2401.05605 (2024)

  153. [160]

    Mohammad Abdullah Matin Khan, M Saiful Bari, Xuan Long Do, Weishi Wang, Md Rizwan Parvez, and Shafiq Joty

  154. [161]

    Sanghyeon Kim, Hyunmo Yang, Yunghyun Kim, Youngjoon Hong, and Eunbyung Park. 2024. Hydra: Multi-head low-rank adaptation for parameter efficient fine-tuning. Neural Networks (2024), 106414

  155. [162]

    Krishnaram Kenthapadi, Mehrnoosh Sameki, and Ankur Taly. 2024. Grounding and evaluation for large language models: Practical challenges and lessons learned (survey). In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 6523–6533

  156. [163]

    Mohammed Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan, Sumanth Doddapaneni, Suriyaprasaad B, Varun G, Sparsh Jain, Anoop Kunchukuttan, Pratyush Kumar, Raj Dabre, and Mitesh Khapra. 2024. IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets...

  157. [164]

    Jing Yu Koh, Daniel Fried, and Russ R Salakhutdinov. 2024. Generating images with multimodal language models. Advances in Neural Information Processing Systems 36 (2024)

  158. [165]

    Patrick Tser Jern Kon, Jiachen Liu, Qiuyi Ding, Yiming Qiu, Zhenning Yang, Yibo Huang, Jayanth Srinivasa, Myungjin Lee, Mosharaf Chowdhury, and Ang Chen. 2025. Curie: Toward rigorous and automated scientific experimentation with ai agents. arXiv preprint arXiv:2502.16069 (2025)

  159. [166]

    Rui Kong, Qiyang Li, Xinyu Fang, Qingtian Feng, Qingfeng He, Yazhu Dong, Weijun Wang, Yuanchun Li, Linghe Kong, and Yunxin Liu. 2024. LoRA-Switch: Boosting the Efficiency of Dynamic LLM Adapters via System-Algorithm Co-design. arXiv preprint arXiv:2405.17741 (2024)

  160. [167]

    Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh. 2024. Propile: Probing privacy leakage in large language models. Advances in Neural Information Processing Systems 36 (2024)

  161. [168]

    Young Jin Kim, Ammar Ahmad Awan, Alexandre Muzio, Andres Felipe Cruz Salinas, Liyang Lu, Amr Hendy, Samyam Rajbhandari, Yuxiong He, and Hany Hassan Awadalla. 2021. Scalable and efficient moe training for multitask multilingual models. arXiv preprint arXiv:2109.10465 (2021)

  162. [169]

    Shachi H Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Beckage, Hsuan Su, Hung-yi Lee, and Lama Nachman. 2024. Decoding biases: Automated methods and llm judges for gender bias detection in language models. arXiv preprint arXiv:2408.03907 (2024)

  163. [170]

    Alexey Kurakin, Natalia Ponomareva, Umar Syed, Liam MacDermed, and Andreas Terzis. 2023. Harnessing large- language models to generate private synthetic text. arXiv preprint arXiv:2306.01684 (2023)

  164. [171]

    Malgorzata Lazuka, Andreea Anghel, and Thomas Parnell. 2024. LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services. In SC24: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–18

  165. [172]

    Weirui Kuang, Bingchen Qian, Zitao Li, Daoyuan Chen, Dawei Gao, Xuchen Pan, Yuexiang Xie, Yaliang Li, Bolin Ding, and Jingren Zhou. 2024. Federatedscope-llm: A comprehensive package for fine-tuning large language models in federated learning. In Proceedings of the 30th ACM SIG...

  166. [173]

    Sonu Kumar, Anubhav Girdhar, Ritesh Patil, and Divyansh Tripathi. 2025. MCP Guardian: A Security-First Layer for Safeguarding MCP-Based AI System. arXiv preprint arXiv:2504.12757 (2025)

  167. [174]

    Bingdong Li, Zixiang Di, Yanting Yang, Hong Qian, Peng Yang, Hao Hao, Ke Tang, and Aimin Zhou. 2024. It’s Morphing Time: Unleashing the Potential of Multiple LLMs via Multi-objective Optimization. arXiv preprint arXiv:2407.00487 (2024)

  168. [175]

    Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. 2024. SEED-Bench: Benchmarking Multimodal Large Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13299–13308

  169. [176]

    Chen-An Li and Hung-Yi Lee. 2024. Examining forgetting in continual pre-training of aligned large language models. arXiv preprint arXiv:2401.03129 (2024)

  170. [177]

    Donghyun Lee and Mo Tiwari. 2024. Prompt infection: Llm-to-llm prompt injection within multi-agent systems. arXiv preprint arXiv:2410.07283 (2024)

  171. [178]

    Weixian Lei, Yixiao Ge, Kun Yi, Jianfeng Zhang, Difei Gao, Dylan Sun, Yuying Ge, Ying Shan, and Mike Zheng Shou

  172. [179]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Vit-lens: Towards omni-modal representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 26647–26657

  173. [180]

    Haitao Li, Qian Dong, Junjie Chen, Huixue Su, Yujia Zhou, Qingyao Ai, Ziyi Ye, and Yiqun Liu. 2024. Llms-as-judges: a comprehensive survey on llm-based evaluation methods. arXiv preprint arXiv:2412.05579 (2024)

  174. [181]

    Haoling Li, Xin Zhang, Xiao Liu, Yeyun Gong, Yifan Wang, Yujiu Yang, Qi Chen, and Peng Cheng. 2024. Gradient-Mask Tuning Elevates the Upper Limits of LLM Performance. arXiv preprint arXiv:2406.15330 (2024)

  175. [182]

    Jia Li, Ge Li, Xuanming Zhang, Yihong Dong, and Zhi Jin. 2024. EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories. arXiv preprint arXiv:2404.00599 (2024)

  176. [183]

    Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, et al. 2024. From generation to judgment: Opportunities and challenges of llm-as-a-judge. arXiv preprint arXiv:2411.16594 (2024). , ...

  177. [184]

    Dong Li, Meng Yan, Yaosheng Zhang, Zhongxin Liu, Chao Liu, Xiaohong Zhang, Ting Chen, and David Lo. 2024. CoSec: On-the-Fly Security Hardening of Code LLMs via Supervised Co-decoding. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analys...

  178. [185]

    Hongyu Li, Liang Ding, Meng Fang, and Dacheng Tao. 2024. Revisiting Catastrophic Forgetting in Large Language Model Tuning. arXiv preprint arXiv:2406.04836 (2024)

  179. [186]

    Ming Li, Jiuhai Chen, Lichang Chen, and Tianyi Zhou. 2024. Can llms speak for diverse people? tuning llms via debate to generate controllable controversial statements. arXiv preprint arXiv:2402.10614 (2024)

  180. [187]

    Miaoge Li, Jingcai Guo, Richard Yi Da Xu, Dongsheng Wang, Xiaofeng Cao, and Song Guo. 2024. TsCA: On the Semantic Consistency Alignment via Conditional Transport for Compositional Zero-Shot Learning. arXiv preprint arXiv:2408.08703 (2024)

  181. [188]

    Qinfeng Li, Zhiqiang Shen, Zhenghan Qin, Yangfan Xie, Xuhong Zhang, Tianyu Du, Sheng Cheng, Xun Wang, and Jianwei Yin. 2024. TransLinkGuard: Safeguarding Transformer Models Against Model Stealing in Edge Deployment. In Proceedings of the 32nd ACM International Conference on Mu...

  182. [189]

    Jing Li, Zhijie Sun, Xuan He, Li Zeng, Yi Lin, Entong Li, Binfan Zheng, Rongqian Zhao, and Xin Chen. 2024. Locmoe: A low-overhead moe for large language model training. arXiv preprint arXiv:2401.13920 (2024)

  183. [190]

    Luchang Li, Sheng Qian, Jie Lu, Lunxi Yuan, Rui Wang, and Qin Xie. 2024. Transformer-lite: High-efficiency deployment of large language models on mobile phone gpus. arXiv preprint arXiv:2403.20041 (2024)

  184. [191]

    Miaomiao Li, Hao Chen, Yang Wang, Tingyuan Zhu, Weijia Zhang, Kaijie Zhu, Kam-Fai Wong, and Jindong Wang

  185. [192]

    arXiv preprint arXiv:2502.04419 (2025)

    Understanding and Mitigating the Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks. arXiv preprint arXiv:2502.04419 (2025)

  186. [193]

    Yannan Li, Yong Yu, Willy Susilo, Zhiyong Hong, and Mohsen Guizani. 2021. Security and privacy for edge intelligence in 5G and beyond networks: Challenges and solutions. IEEE Wireless Communications 28, 2 (2021), 63–69

  187. [194]

    Zhaowei Li, Wei Wang, YiQing Cai, Xu Qi, Pengyu Wang, Dong Zhang, Hang Song, Botian Jiang, Zhida Huang, and Tao Wang. 2024. Unifiedmllm: Enabling unified representation for multi-modal multi-tasks with large language model. arXiv preprint arXiv:2408.02503 (2024)

  188. [195]

    Zhaowei Li, Qi Xu, Dong Zhang, Hang Song, Yiqing Cai, Qi Qi, Ran Zhou, Junting Pan, Zefeng Li, Vu Tu, et al. 2024. Groundinggpt: Language enhanced multi-modal grounding model. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1:...

  189. [196]

    Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. 2023. Starcoder: may the source be with you!arXiv preprint arXiv:2305.06161 (2023)

  190. [197]

    Tianhao Li, Shangjie Li, Binbin Xie, Deyi Xiong, and Baosong Yang. 2024. MoE-CT: a novel approach for large language models training with resistance to catastrophic forgetting. arXiv preprint arXiv:2407.00875 (2024)

  191. [198]

    Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190 (2021)

  192. [199]

    Yichen Li, Yun Peng, Yintong Huo, and Michael R Lyu. 2024. Enhancing llm-based coding tools through native integration of ide-derived static context. In Proceedings of the 1st International Workshop on Large Language Models for Code. 70–74

  193. [200]

    Yang Lin, Xinyu Ma, Xu Chu, Yujie Jin, Zhibang Yang, Yasha Wang, and Hong Mei. 2024. Lora dropout as a sparsity regularizer for overfitting control. arXiv preprint arXiv:2404.09610 (2024)

  194. [201]

    Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, and Song Han. 2024. Qserve: W4a8kv4 quantization and system co-design for efficient llm serving. arXiv preprint arXiv:2405.04532 (2024)

  195. [202]

    Yen-Ting Lin and Yun-Nung Chen. 2023. Llm-eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models. arXiv preprint arXiv:2305.13711 (2023)

  196. [203]

    Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. 2024. Monkey: Image resolution and text label are important things for large multi-modal models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern R...

  197. [204]

    Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. 2024. Vila: On pre-training for visual language models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 26689–26699

  198. [205]

    Luyang Lin, Lingzhi Wang, Jinsong Guo, and Kam-Fai Wong. 2024. Investigating bias in llm-based bias detection: Disparities between llms and human perception. arXiv preprint arXiv:2403.14896 (2024)

  199. [206]

    Sam Lin, Wenyue Hua, Zhenting Wang, Mingyu Jin, Lizhou Fan, and Yongfeng Zhang. 2025. EmojiPrompt: Generative Prompt Obfuscation for Privacy-Preserving Communication with Cloud-based LLMs. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Asso...

  200. [207]

    Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. 2023. Mitigating hallucination in large multi-modal models via robust instruction tuning. InThe Twelfth International Conference on Learning Representations

  201. [208]

    Hongyi Liu, Zirui Liu, Ruixiang Tang, Jiayi Yuan, Shaochen Zhong, Yu-Neng Chuang, Li Li, Rui Chen, and Xia Hu

  202. [209]

    Jingyu Liu, Jiaen Lin, and Yong Liu. 2024. How much can rag help the reasoning of llm?arXiv preprint arXiv:2410.02338 (2024)

  203. [211]

    Bingchang Liu, Chaoyu Chen, Zi Gong, Cong Liao, Huan Wang, Zhichao Lei, Ming Liang, Dajun Chen, Min Shen, Hailian Zhou, et al. 2024. Mftcoder: Boosting code llms with multitask fine-tuning. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining...

  204. [212]

    Chengyuan Liu, Yangyang Kang, Shihang Wang, Lizhi Qing, Fubang Zhao, Changlong Sun, Kun Kuang, and Fei Wu

  205. [213]

    arXiv preprint arXiv:2405.17830 (2024)

    More than catastrophic forgetting: Integrating general capabilities for domain-specific llms. arXiv preprint arXiv:2405.17830 (2024)

  206. [214]

    Dong Liu. 2024. Contemporary Model Compression on Large Language Models Inference. arXiv preprint arXiv:2409.01990 (2024)

  207. [215]

    Qidong Liu, Xian Wu, Xiangyu Zhao, Yuanshao Zhu, Derong Xu, Feng Tian, and Yefeng Zheng. 2024. When moe meets llms: Parameter efficient fine-tuning for multi-task medical applications. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in...

  208. [216]

    Shigang Liu, Di Cao, Junae Kim, Tamas Abraham, Paul Montague, Seyit Camtepe, Jun Zhang, and Yang Xiang. 2024. {EaTVul}:{ChatGPT-based} Evasion Attack Against Software Vulnerability Detection. In 33rd USENIX Security Symposium (USENIX Security 24) . 7357–7374

  209. [217]

    arXiv preprint arXiv:2403.00108 (2024)

    LoRA-as-an-Attack! Piercing LLM Safety Under The Share-and-Play Scenario. arXiv preprint arXiv:2403.00108 (2024)

  210. [218]

    Tianyang Liu, Canwen Xu, and Julian McAuley. 2023. Repobench: Benchmarking repository-level code auto- completion systems. arXiv preprint arXiv:2306.03091 (2023)

  211. [219]

    Junwei Liu, Kaixin Wang, Yixuan Chen, Xin Peng, Zhenpeng Chen, Lingming Zhang, and Yiling Lou. 2024. Large language model-based agents for software engineering: A survey. arXiv preprint arXiv:2409.02977 (2024)

  212. [220]

    Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2024. Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems 36 (2024)

  213. [221]

    Lian Liu, Haimeng Ren, Long Cheng, Zhaohui Xu, Yudong Pan, Mengdi Wang, Xiaowei Li, Yinhe Han, and Ying Wang. 2024. COMET: Towards Partical W4A4KV4 LLMs Serving. arXiv preprint arXiv:2410.12168 (2024)

  214. [222]

    Liang Liu, Dong Zhang, Shoushan Li, Guodong Zhou, and Erik Cambria. 2024. Two Heads are Better than One: Zero-shot Cognitive Reasoning via Multi-LLM Knowledge Fusion. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management . 1462–1472

  215. [223]

    Minghao Liu, Zonglin Di, Jiaheng Wei, Zhongruo Wang, Hengxiang Zhang, Ruixuan Xiao, Haoyu Wang, Jinlong Pang, Hao Chen, Ankit Shah, et al. 2024. Automatic dataset construction (adc): Sample collection, data curation, and beyond. arXiv preprint arXiv:2408.11338 (2024)

  216. [224]

    Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra. 2023. Llm-qat: Data-free quantization aware training for large language models. arXiv preprint arXiv:2305.17888 (2023)

  217. [225]

    Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu. 2024. Kivi: A tuning-free asymmetric 2bit quantization for kv cache. arXiv preprint arXiv:2402.02750 (2024)

  218. [226]

    Tong Liu, Zizhuang Deng, Guozhu Meng, Yuekang Li, and Kai Chen. 2024. Demystifying rce vulnerabilities in llm-integrated apps. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security . 1716–1730

  219. [227]

    Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, et al. 2023. A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity. arXiv preprint...

  220. [228]

    Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao. 2025. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. In European Conference on Computer Vision . Springer, 386–403

  221. [229]

    Yang Liu, Jiahuan Cao, Chongyu Liu, Kai Ding, and Lianwen Jin. 2024. Datasets for large language models: A comprehensive survey. arXiv preprint arXiv:2402.18041 (2024)

  222. [230]

    Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Zihao Wang, Xiaofeng Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, et al. 2023. Prompt Injection attack against LLM-integrated Applications. arXiv preprint arXiv:2306.05499 (2023)

  223. [231]

    Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong. 2024. Formalizing and benchmarking prompt injection attacks and defenses. In 33rd USENIX Security Symposium (USENIX Security 24) . 1831–1847. , Vol. 1, No. 1, Article . Publication date: August 2025. Soft...

  224. [232]

    Yilun Liu, Shimin Tao, Xiaofeng Zhao, Ming Zhu, Wenbing Ma, Junhao Zhu, Chang Su, Yutai Hou, Miao Zhang, Min Zhang, et al. 2024. Coachlm: Automatic instruction revisions improve the data quality in llm instruction tuning. In 2024 IEEE 40th International Conference on Data Engi...

  225. [233]

    Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Lifeng Dong, Ruiping Wang, Jilong Xue, and Furu Wei. 2024. The era of 1-bit llms: All large language models are in 1.58 bits. arXiv preprint arXiv:2402.17764 1 (2024)

  226. [234]

    Wanqin Ma, Chenyang Yang, and Christian Kästner. 2024. (Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs. InProceedings of the IEEE/ACM 3rd International Conference on AI Engineering-Software Engineering for AI. 166–171

  227. [235]

    Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, and Haobo Wang. 2024. On llms-driven synthetic data generation, curation, and evaluation: A survey. arXiv preprint arXiv:2406.15126 (2024)

  228. [236]

    Zilin Ma, Yiyang Mei, Krzysztof Z Gajos, and Ian Arawjo. 2024. Schrödinger’s Update: User Perceptions of Uncertainties in Proprietary Large Language Model Updates. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–9

  229. [237]

    Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, et al. 2024. Starcoder 2 and the stack v2: The next generation. arXiv preprint arXiv:2402.19173 (2024)

  230. [238]

    Daqin Luo, Chengjian Feng, Yuxuan Nong, and Yiqing Shen. 2024. Autom3l: An automated multimodal machine learning framework with large language models. InProceedings of the 32nd ACM International Conference on Multimedia. 8586–8594

  231. [239]

    Junyu Luo, Weizhi Zhang, Ye Yuan, Yusheng Zhao, Junwei Yang, Yiyang Gu, Bohan Wu, Binqi Chen, Ziyue Qiao, Qingqing Long, et al. 2025. Large Language Model Agent: A Survey on Methodology, Applications and Challenges. arXiv preprint arXiv:2503.21460 (2025)

  232. [240]

    Xiang Luo, Zhiwen Tang, Jin Wang, and Xuejie Zhang. 2024. Zero-Shot Cross-Domain Dialogue State Tracking via Dual Low-Rank Adaptation. arXiv preprint arXiv:2407.21633 (2024)

  233. [241]

    Alisia Lupidi, Carlos Gemmell, Nicola Cancedda, Jane Dwivedi-Yu, Jason Weston, Jakob Foerster, Roberta Raileanu, and Maria Lomeli. 2024. Source2synth: Synthetic data generation and curation grounded in real data sources. arXiv preprint arXiv:2409.08239 (2024)

  234. [242]

    Arsalan Masoudifard, Mohammad Mowlavi Sorond, Moein Madadi, Mohammad Sabokrou, and Elahe Habibi. 2024. Leveraging Graph-RAG and Prompt Engineering to Enhance LLM-Based Automated Requirement Traceability and Compliance Checks. arXiv preprint arXiv:2412.08593 (2024)

  235. [243]

    Timothy R McIntosh, Teo Susnjak, Nalin Arachchilage, Tong Liu, Paul Watters, and Malka N Halgamuge. 2024. Inadequacies of large language model benchmarks in the era of generative artificial intelligence. arXiv preprint arXiv:2402.09880 (2024). , Vol. 1, No. 1, Article . Public...

  236. [244]

    Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023. Llm-pruner: On the structural pruning of large language models. Advances in neural information processing systems 36 (2023), 21702–21720

  237. [245]

    Ahmed Menshawy, Zeeshan Nawaz, and Mahmoud Fahmy. 2024. Navigating Challenges and Technical Debt in Large Language Models Deployment. In Proceedings of the 4th Workshop on Machine Learning and Systems . 192–199

  238. [246]

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems 36 (2023), 46534–46594

  239. [247]

    Pratyush Maini, Hengrui Jia, Nicolas Papernot, and Adam Dziedzic. 2024. LLM Dataset Inference: Did you train on my dataset? arXiv preprint arXiv:2406.06443 (2024)

  240. [248]

    Pratyush Maini, Hengrui Jia, Nicolas Papernot, and Adam Dziedzic. 2025. LLM Dataset Inference: Did you train on my dataset? Advances in Neural Information Processing Systems 37 (2025), 124069–124092

  241. [249]

    Paul Joe Maliakel, Shashikant Ilager, and Ivona Brandic. 2025. Investigating Energy Efficiency and Performance Trade-offs in LLM Inference Across Tasks and DVFS Settings. arXiv preprint arXiv:2501.08219 (2025)

  242. [250]

    Vorbe\c{s} ti Rom\ˆ ane\c{s} te?

    Mihai Masala, Denis C Ilie-Ablachim, Alexandru Dima, Dragos Corlatescu, Miruna Zavelca, Ovio Olaru, Simina Terian, Andrei Terian, Marius Leordeanu, Horia Velicu, et al. 2024. " Vorbe\c{s} ti Rom\ˆ ane\c{s} te?" A Recipe to Train Powerful Romanian LLMs with English Instructions...

  243. [251]

    Ida Momennejad, Hosein Hasanbeig, Felipe Vieira Frujeri, Hiteshi Sharma, Nebojsa Jojic, Hamid Palangi, Robert Ness, and Jonathan Larson. 2024. Evaluating cognitive maps and planning in large language models with CogEval. Advances in Neural Information Processing Systems 36 (2024)

  244. [252]

    Mahdi Morafah, Vyacheslav Kungurtsev, Hojin Chang, Chen Chen, and Bill Lin. 2024. Towards Diverse Device Heterogeneous Federated Learning via Task Arithmetic Knowledge Integration. arXiv preprint arXiv:2409.18461 (2024)

  245. [253]

    Matthieu Meeus, Shubham Jain, Marek Rei, and Yves-Alexandre de Montjoye. 2024. Did the neurons read your book? document-level membership inference for large language models. In 33rd USENIX Security Symposium (USENIX Security 24). 2369–2385

  246. [254]

    Vineeth Sai Narajala and Idan Habler. 2025. Enterprise-Grade Security for the Model Context Protocol (MCP): Frameworks and Mitigation Strategies. arXiv preprint arXiv:2504.08623 (2025)

  247. [255]

    Microsoft. 2025. Azure OpenAI Service. https://azure.microsoft.com/en-us/products/ai-services/openai-service

  248. [256]

    Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. 2021. Accelerating sparse deep neural networks. arXiv preprint arXiv:2104.08378 (2021)

  249. [257]

    Yuhong Mo, Hao Qin, Yushan Dong, Ziyi Zhu, and Zhenglin Li. 2024. Large language model (llm) ai text generation detection based on transformer deep learning algorithm. arXiv preprint arXiv:2405.06652 (2024)

  250. [258]

    Apoorve Mohan, Mengmei Ye, Hubertus Franke, Mudhakar Srivatsa, Zhuoran Liu, and Nelson Mimura Gonzalez. 2024. Securing AI Inference in the Cloud: Is CPU-GPU Confidential Computing Ready?. In 2024 IEEE 17th International Conference on Cloud Computing (CLOUD) . IEEE, 164–175

  251. [259]

    Ahmad Mohsin, Helge Janicke, Adrian Wood, Iqbal H Sarker, Leandros Maglaras, and Naeem Janjua. 2024. Can we trust large language models generated code? a framework for in-context learning, security patterns, and code evaluations across diverse llms. arXiv preprint arXiv:2406.1...

  252. [260]

    Blaschko, Guohao Dai, Huazhong Yang, and Yu Wang

    Xuefei Ning, Zifu Wang, Shiyao Li, Zinan Lin, Peiran Yao, Tianyu Fu, Matthew B. Blaschko, Guohao Dai, Huazhong Yang, and Yu Wang. 2024. Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study. arXiv:2406.14629 [cs.CL] https://arxiv.org/abs/2406.14629

  253. [261]

    Kosuke Nishida, Kyosuke Nishida, and Kuniko Saito. 2024. Initialization of large language models via reparameteriza- tion to mitigate loss spikes. arXiv preprint arXiv:2410.05052 (2024)

  254. [262]

    Alhassan Mumuni and Fuseini Mumuni. 2025. Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches. arXiv preprint arXiv:2501.03151 (2025)

  255. [263]

    NVIDIA. 2025. TensorRT: High-Performance Deep Learning Inference Library. https://github.com/NVIDIA/TensorRT

  256. [264]

    Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, et al. 2021. Efficient large-scale language model training on gpu clusters using megatron-lm. In Pro...

  257. [265]

    Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. 2023. A comprehensive overview of large language models. arXiv preprint arXiv:2307.06435 (2023)

  258. [266]

    Nihal V Nayak, Yiyang Nan, Avi Trost, and Stephen H Bach. 2024. Learning to generate instruction tuning datasets for zero-shot task adaptation. arXiv preprint arXiv:2402.18334 (2024)

  259. [267]

    Mahmoud Nazzal, Issa Khalil, Abdallah Khreishah, and NhatHai Phan. 2024. PromSec: Prompt Optimization for Secure Generation of Functional Source Code with Large Language Models (LLMs). In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security ...

  260. [268]

    Shiwen Ni, Dingwei Chen, Chengming Li, Xiping Hu, Ruifeng Xu, and Min Yang. 2023. Forgetting before learning: Utilizing parametric arithmetic for knowledge updating in large language models. arXiv preprint arXiv:2311.08011 (2023)

  261. [269]

    Malte Ostendorff, Pedro Ortiz Suarez, Lucas Fonseca Lage, and Georg Rehm. 2024. LLM-Datasets: An Open Framework for Pretraining Datasets of Large Language Models. In First Conference on Language Modeling

  262. [270]

    Tiago P Pagano, Rafael B Loureiro, Fernanda VN Lisboa, Rodrigo M Peixoto, Guilherme AS Guimarães, Gustavo OR Cruz, Maira M Araujo, Lucas L Santos, Marco AS Cruz, Ewerton LS Oliveira, et al . 2023. Bias and unfairness in machine learning models: a systematic review on datasets,...

  263. [271]

    Dominic Novado, Eliyah Cohen, and Jacob Foster. 2024. Multi-tier privacy protection for large language models using differential privacy. Authorea Preprints (2024)

  264. [272]

    Leyi Pan, Aiwei Liu, Shiyu Huang, Yijian Lu, Xuming Hu, Lijie Wen, Irwin King, and Philip S Yu. 2025. Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation? arXiv preprint arXiv:2502.11598 (2025)

  265. [273]

    Hyungjun Oh, Kihong Kim, Jaemin Kim, Sungkyun Kim, Junyeol Lee, Du-seong Chang, and Jiwon Seo. 2024. Exegpt: Constraint-aware resource scheduling for llm inference. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and O...

  266. [274]

    Ollama. 2025. Ollama Official Website. https://ollama.com/. Accessed: 2025-05-15

  267. [275]

    OpenAI. 2023. OpenAI Platform API Reference. https://platform.openai.com/docs/api-reference/introduction

  268. [276]

    OpenAI. 2025. Introducing GPT-4.5. https://openai.com/index/introducing-gpt-4-5/ Accessed: 2025-04-20

  269. [277]

    OpenAI. 2025. OpenAI API Platform. https://openai.com/api/. , Vol. 1, No. 1, Article . Publication date: August 2025. Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead 53

  270. [278]

    Bhrij Patel, Vishnu Sashank Dorbala, and Amrit Singh Bedi. 2024. Embodied Question Answering via Multi-LLM Systems. arXiv preprint arXiv:2406.10918 (2024)

  271. [279]

    Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini

  272. [280]

    Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, et al. 2024. Byte Latent Transformer: Patches Scale Better Than Tokens. arXiv preprint arXiv:2412.09871 (2024)

  273. [281]

    Rodrigo Pedro, Miguel E Coimbra, Daniel Castro, Paulo Carreira, and Nuno Santos. 2024. Prompt-to-SQL Injections in LLM-Integrated Web Applications: Risks and Defenses. In2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE Computer Society, 76–88

  274. [282]

    Jinlong Pang, Jiaheng Wei, Ankit Parag Shah, Zhaowei Zhu, Yaxuan Wang, Chen Qian, Yang Liu, Yujia Bao, and Wei Wei. 2024. Improving Data Efficiency via Curating LLM-Driven Rating Systems. arXiv preprint arXiv:2410.10877 (2024)

  275. [283]

    Qi Pang, Shengyuan Hu, Wenting Zheng, and Virginia Smith. 2024. Attacking llm watermarks by exploiting their strengths. In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models

  276. [284]

    Xiaoyi Pang, Jiahui Hu, Peng Sun, Ju Ren, and Zhibo Wang. 2024. When Federated Learning Meets Knowledge Distillation. IEEE Wireless Communications 31, 5 (2024), 208–214

  277. [285]

    Samuel Panterino and Matthew Fellington. 2024. Dynamic moving target defense for mitigating targeted llm prompt injection. Authorea Preprints (2024)

  278. [286]

    David Pape, Sina Mavali, Thorsten Eisenhofer, and Lea Schönherr. 2024. Prompt obfuscation for large language models. arXiv preprint arXiv:2409.11026 (2024)

  279. [287]

    Max Ploner and Alan Akbik. 2024. Parameter-Efficient Fine-Tuning: Is There An Optimal Subset of Parameters to Tune?. In Findings of the Association for Computational Linguistics: EACL 2024 . 1743–1759

  280. [288]

    PyTorch Team. 2023. Executorch Overview. https://pytorch.org/executorch-overview

  281. [289]

    In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA)

    Splitwise: Efficient generative llm inference using phase splitting. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) . IEEE, 118–132

  282. [290]

    Pankayaraj Pathmanathan, Souradip Chakraborty, Xiangyu Liu, Yongyuan Liang, and Furong Huang. 2024. Is poisoning a real threat to LLM alignment? Maybe more so than you think. arXiv preprint arXiv:2406.12091 (2024)

  283. [292]

    Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Bap- tiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only. ...

  284. [293]

    Letian Peng, Zilong Wang, Feng Yao, Zihan Wang, and Jingbo Shang. 2024. Metaie: Distilling a meta model from llm for all kinds of information extraction tasks. arXiv preprint arXiv:2404.00457 (2024)

  285. [294]

    Nikhil Pesati. 2024. Security Considerations for Large Language Model Use: Implementation Research in Securing LLM-Integrated Applications. A vailable at SSRN 4962370 (2024)

  286. [295]

    Matthew E Peters, Sebastian Ruder, and Noah A Smith. 2019. To tune or not to tune? adapting pretrained representa- tions to diverse tasks. arXiv preprint arXiv:1903.05987 (2019)

  287. [296]

    Julien Piet, Maha Alrashed, Chawin Sitawarin, Sizhe Chen, Zeming Wei, Elizabeth Sun, Basel Alomair, and David Wagner. 2024. Jatmo: Prompt injection defense by task-specific finetuning. In European Symposium on Research in Computer Security. Springer, 105–124

  288. [299]

    Shuang Qiao, Haiyang Xu, Chenhong Cao, Wei Gong, Si Chen, and Jiangchuan Liu. 2025. PrismPrompt: Layering prompt-enhanced cloud-edge collaborative language model towards healthcare. IEEE Network (2025)

  289. [300]

    Laiqiao Qin, Tianqing Zhu, Wanlei Zhou, and Philip S Yu. 2024. Knowledge distillation in federated learning: A survey on long lasting challenges and new solutions. arXiv preprint arXiv:2406.10861 (2024)

  290. [2023]

    arXiv preprint arXiv:2303.03004 (2023)

    xcodeeval: A large scale multilingual multitask benchmark for code understanding, generation, translation and retrieval. arXiv preprint arXiv:2303.03004 (2023)

  291. [2024]

    arXiv preprint arXiv:2410.07383 (2024)

    SparseGrad: A Selective Method for Efficient Fine-tuning of MLP Layers. arXiv preprint arXiv:2410.07383 (2024)

  292. [2025]

    ACM SIGKDD Explorations Newsletter 26, 2 (2025), 76–88

    Ft-shield: A watermark against unauthorized fine-tuning in text-to-image diffusion models. ACM SIGKDD Explorations Newsletter 26, 2 (2025), 76–88

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.