Pith. sign in

REVIEW 2 major objections 2 minor 2 cited by

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda

T0 review · 2 major / 2 minor · reviewed 2026-05-10 · grok-4.3

Pith's one-line read Large language models depend on cloud-native and distributed systems for efficient scaling.

desk verdict This is a standard research agenda that lists familiar LLM scaling problems and names cloud trends without new mechanisms, data, or resolved questions. read the letter →

arxiv 2604.17227 v1 submitted 2026-04-19 cs.DC

classification cs.DC
keywords largelanguagemodelscloud-nativesystemsdistributedcomputingscalabilityefficiencyserverlessinferencequantumfederatedlearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The authors aim to establish that cloud-native and distributed systems must be integrated into LLM workflows because traditional setups fall short on the needed scale and efficiency. This matters because LLMs underpin growing AI uses from text processing to code writing, and bottlenecks in computing could slow their development. The paper details practical issues like managing data and optimizing resources, then highlights tools such as automatic scaling and hybrid cloud setups as solutions. It also surveys newer ideas including serverless inference and federated learning as future paths. Finally, it provides a roadmap stressing the value of shared standards and teamwork across fields.

What carries the argument

Cloud-native and distributed architectures, which enable handling of LLM computational demands through features like microservices, autoscaling, and hybrid cloud-edge solutions.

What would settle it

A demonstration that a large-scale LLM can be trained and deployed efficiently using only traditional non-cloud, non-distributed infrastructure.

Watch

Extended reading notes

Core claim

The paper claims that cloud platforms and distributed systems play a key role in supporting the scalability, efficiency, and optimization of large language models. The complexities of LLM deployment involve data management, resource optimization, and the adoption of microservices, autoscaling, and hybrid cloud-edge solutions. Emerging research trends such as serverless inference, quantum computing, and federated learning hold potential to advance LLM capabilities further. A roadmap for future work calls for ongoing research, standardization efforts, and collaboration between sectors to support LLM expansion in research and enterprise settings.

Load-bearing premise

That traditional systems are unable to meet the computational requirements of large language models and that the listed emerging trends will drive the next phase of innovation.

Editorial extensions

If this is right

  • LLM deployment can incorporate data management and resource optimization techniques from cloud systems.
  • Hybrid cloud-edge solutions will address specific deployment complexities.
  • Serverless inference, quantum computing, and federated learning represent promising directions for next innovations.
  • Continued research, standardization, and cross-sector collaboration are required to sustain LLM growth.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Organizations without large data centers could leverage these systems to develop competitive LLMs.
  • The emphasis on federated learning may enable privacy-preserving model training across distributed data sources.
  • Quantum computing integration could transform energy efficiency in model operations if realized.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper claims that the computational demands of Large Language Models (LLMs) necessitate the integration of cloud-native and distributed architectures, as traditional systems are often unable to meet these requirements. It explores complexities in LLM deployment such as data management, resource optimization, microservices, autoscaling, and hybrid cloud-edge solutions. The manuscript examines emerging research trends including serverless inference, quantum computing, and federated learning, and concludes with a roadmap for future developments emphasizing continued research, standardization, and cross-sector collaboration.

Significance. If pursued, this agenda could help direct research toward more scalable and efficient LLM systems by highlighting the intersection of distributed computing and AI challenges. The paper's strength is in providing a high-level synthesis of deployment issues and forward-looking trends, which may stimulate targeted follow-up studies, though its impact as a position paper rests on the actionability of the proposed directions rather than new derivations or data.

major comments (2)
  1. [Abstract] Abstract: The foundational assertion that 'Traditional systems are often unable to meet these requirements' is presented without any specific benchmarks, scaling examples, or citations to LLM performance bottlenecks, which is load-bearing for the motivation to integrate cloud-native architectures.
  2. [Abstract] Abstract (emerging trends paragraph): The statement that trends such as serverless inference, quantum computing, and federated learning 'will drive the next phase of LLM innovation' is made without analysis of their current maturity, specific applicability to LLM training/inference, or preliminary evidence, which underpins the credibility of the concluding roadmap.
minor comments (2)
  1. The discussion of deployment complexities would benefit from clearer organization, such as explicit subsection headings for data management versus resource optimization, to improve readability.
  2. Additional citations to recent work on distributed LLM training frameworks would help ground the high-level claims in existing literature.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback and the recommendation for minor revision. We address the two points on the abstract below and will incorporate targeted clarifications to strengthen the grounding of our claims while preserving the high-level nature of this research agenda paper.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The foundational assertion that 'Traditional systems are often unable to meet these requirements' is presented without any specific benchmarks, scaling examples, or citations to LLM performance bottlenecks, which is load-bearing for the motivation to integrate cloud-native architectures.

    Authors: We agree that the abstract would benefit from explicit support for this statement. The full manuscript already references scaling challenges such as quadratic attention complexity and the multi-node requirements for training models with hundreds of billions of parameters. In revision, we will add one or two concise citations (e.g., to transformer scaling laws and reported training infrastructure needs) directly in the abstract to anchor the claim without expanding its length or shifting the paper's position-paper character. revision: yes

  2. Referee: [Abstract] Abstract (emerging trends paragraph): The statement that trends such as serverless inference, quantum computing, and federated learning 'will drive the next phase of LLM innovation' is made without analysis of their current maturity, specific applicability to LLM training/inference, or preliminary evidence, which underpins the credibility of the concluding roadmap.

    Authors: This observation is valid for the abstract's brevity. The body of the manuscript discusses maturity levels and applicability in more detail (e.g., serverless for inference workloads, federated learning for privacy-preserving training). For the revision, we will soften the phrasing to 'emerging trends with the potential to drive...' and insert a short qualifier noting their varying stages of readiness, thereby improving credibility while keeping the abstract concise and forward-looking. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity

full rationale

The manuscript is a descriptive research agenda surveying LLM challenges and advocating cloud-native/distributed approaches plus emerging trends. It contains no equations, formal derivations, fitted parameters, or predictions that reduce to inputs by construction. Central claims are high-level motivations and calls for future work rather than asserted results resting on self-referential steps, self-citations, or renamed empirical patterns. The text is therefore self-contained with no load-bearing circular reductions.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

The paper is a research agenda without new technical claims, so it introduces no free parameters, axioms, or invented entities beyond standard assumptions in cloud computing literature.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda." pith.science (2026). https://pith.science/paper/2604.17227

@misc{pith2026260417227,
  author       = {Pith},
  title        = {Pith review of: Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2604.17227}},
  note         = {Machine review of arXiv:2604.17227}
}
read the original abstract

The rapid rise of Large Language Models (LLMs) has revolutionized various artificial intelligence (AI) applications, from natural language processing to code generation. However, the computational demands of these models, particularly in training and inference, present significant challenges. Traditional systems are often unable to meet these requirements, necessitating the integration of cloud-native and distributed architectures. This paper explores the role of cloud platforms and distributed systems in supporting the scalability, efficiency, and optimization of LLMs. We discuss the complexities of LLM deployment, including data management, resource optimization, and the need for microservices, autoscaling, and hybrid cloud-edge solutions. Additionally, we examine emerging research trends, such as serverless inference, quantum computing, and federated learning, and their potential to drive the next phase of LLM innovation. The paper concludes with a roadmap for future developments, emphasizing the need for continued research, standardization, and cross-sector collaboration to sustain the growth of LLMs in both research and enterprise applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving

    cs.DC 2026-07 conditional novelty 6.5 of 10

    Layer-wise data-parallel replication of hot Transformer layers onto reclaimed idle GPUs reduces LLM serving cold-start latency 97.9–99.3% and average latency 20.7–28.1% while attaining 100% SLO on production traces.

  2. BrownoutMoE: Structure-Aware Expert Grouping for Efficient and Accurate LLM Web-based Services

    cs.DC 2026-07 conditional novelty 5.5 of 10

    GRPO-optimized expert grouping plus grouping-consistent distillation reduces MoE brownout accuracy degradation by up to 71.4% and raises throughput by up to 2.24× versus sequential grouping.

Reference graph

Works this paper leans on

151 extracted references · 151 canonical work pages · cited by 2 Pith papers

  1. [1]

    In: Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation

    ZhongY,LiuS,ChenJ,etal.DistServe:Disaggregatingprefillanddecodingforgoodput-optimizedlargelanguagemodel serving. In: Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation. 2024:193– 210

  2. [2]

    Cloud container technologies: a state-of-the-art review.IEEE Transactions on Cloud Computing.2017;7(3):677–692

    Pahl C, Brogi A, Soldani J, Jamshidi P. Cloud container technologies: a state-of-the-art review.IEEE Transactions on Cloud Computing.2017;7(3):677–692

  3. [3]

    Attention is all you need

    Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. In: Proceedings of the Advances in neural information processing systems. 2017

  4. [4]

    Scaling Laws for Neural Language Models

    Kaplan J, McCandlish S, Henighan T, et al. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361. 2020

  5. [5]

    Deep learning.Nature.2015;521(7553):436–444

    LeCun Y, Bengio Y, Hinton G. Deep learning.Nature.2015;521(7553):436–444

  6. [6]

    IsaevM,McDonaldN,VuducR.Scalinginfrastructuretosupportmulti-trillionparameterLLMtraining.In:Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 2023

  7. [7]

    Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.; 2020

    Shoeybi M, Patwary M, Puri R, LeGresley P, Casper J, Catanzaro B. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.; 2020

  8. [8]

    Curran Associates, Inc

    ZhangY,YouY.SpeedLoader:AnI/OefficientschemeforheterogeneousanddistributedLLMoperation.In:Proceedings of the Advances in Neural Information Processing Systems. Curran Associates, Inc. 2024:34637–34655

Show all 151 references
  1. [9]

    Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

    Aminabadi RY, Rajbhandari S, Awan AA, et al. Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 2022:1–15

  2. [10]

    FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

    Dao T, Fu D, Ermon S, Rudra A, Ré C. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In: Proceedings of the Advances in Neural Information Processing Systems. Curran Associates, Inc. 2022:16344–16359

  3. [11]

    Towards end-to-end optimization of llm-based applications with ayo

    Tan X, Jiang Y, Yang Y, Xu H. Towards end-to-end optimization of llm-based applications with ayo. In: Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. 2025:1302–1316

  4. [12]

    In-datacenter performance analysis of a tensor processing unit

    Jouppi NP, Young C, Patil N, et al. In-datacenter performance analysis of a tensor processing unit. In: Proceedings of the 44th annual international symposium on computer architecture. 2017:1–12

  5. [13]

    Alpa: Automating Inter-and Intra-Operator Parallelism for Distributed Deep Learning

    Zheng L, Li Z, Zhang H, et al. Alpa: Automating Inter-and Intra-Operator Parallelism for Distributed Deep Learning. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation. 2022. 38 Xu ET AL

  6. [14]

    Efficient memory management for large language model serving with pagedattention

    Kwon W, Li Z, Zhuang S, et al. Efficient memory management for large language model serving with pagedattention. In: Proceedings of the 29th symposium on operating systems principles. 2023:611–626

  7. [15]

    Open issues in scheduling microservices in the cloud.IEEE Cloud Computing.2016;3(5):81–88

    Fazio M, Celesti A, Ranjan R, Liu C, Chen L, Villari M. Open issues in scheduling microservices in the cloud.IEEE Cloud Computing.2016;3(5):81–88

  8. [16]

    TheKubernetesAuthors.Kubernetes:Production-GradeContainerOrchestration.https://kubernetes.io;.Accessed:2026- 04-08

  9. [17]

    Flexgen: High-throughput generative inference of large language models with a single gpu

    Sheng Y, Zheng L, Yuan B, et al. Flexgen: High-throughput generative inference of large language models with a single gpu. In: Proceedings of the International Conference on Machine Learning. 2023:31094–31116

  10. [18]

    In: Proceedings of the 16th USENIX symposium on operating systems design and implementation

    YuGI,JeongJS,KimGW,KimS,ChunBG.Orca:AdistributedservingsystemforTransformer-Basedgenerativemodels. In: Proceedings of the 16th USENIX symposium on operating systems design and implementation. 2022:521–538

  11. [19]

    Evaluation and benchmarking of llm agents: A survey

    Mohammadi M, Li Y, Lo J, Yip W. Evaluation and benchmarking of llm agents: A survey. In: Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 2025:6129–6139

  12. [20]

    Towards High-Goodput LLM Serving with Prefill-decode Multiplexing

    Chen Y, Cui W, Zhao H, et al. Towards High-Goodput LLM Serving with Prefill-decode Multiplexing. In: Proceedings ofthe31stACMInternationalConferenceonArchitecturalSupportforProgrammingLanguagesandOperatingSystems, Volume 2. 2026:2030–2047

  13. [21]

    Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

    Rombaut B, Masoumzadeh S, Vasilevski K, Lin D, Hassan AE. Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents. In: Proceedings of the 40th IEEE/ACM International Conference on Automated Software Engineering. 2025:739–751

  14. [22]

    Process modeling in web applications.ACM Transactions on Software Engineering and Methodology.2006;15(4):360–409

    Brambilla M, Ceri S, Fraternali P, Manolescu I. Process modeling in web applications.ACM Transactions on Software Engineering and Methodology.2006;15(4):360–409

  15. [23]

    Challenges in deployment and configuration management in cyber physical system

    Jha DN, Li Y, Jayaraman PP, et al. Challenges in deployment and configuration management in cyber physical system. Handbook of Integration of Cloud Computing, Cyber Physical Systems and Internet of Things.2020:215–235

  16. [24]

    Parallel processing systems for big data: a survey

    Zhang Y, Cao T, Li S, et al. Parallel processing systems for big data: a survey. Proceedings of the IEEE. 2016;104(11):2114–2136

  17. [25]

    LiuX,BuyyaR.Resourcemanagementandschedulingindistributedstreamprocessingsystems:ataxonomy,review,and future directions.ACM Computing Surveys.2020;53(3):1–41

  18. [26]

    Deep learning workload scheduling in gpu datacenters: A survey.ACM Computing Surveys

    Ye Z, Gao W, Hu Q, et al. Deep learning workload scheduling in gpu datacenters: A survey.ACM Computing Surveys. 2024;56(6):1–38

  19. [27]

    StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications.ACM Trans

    Wen L, Xu M, Gill SS, et al. StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications.ACM Trans. Auton. Adapt. Syst..2025;20(1). doi: 10.1145/3686253

  20. [28]

    LLM Inference Scheduling: A Survey of Techniques, Frameworks, and Trade-offs.Authorea Preprints.2025

    Heisler M, Yousefijamarani Z, Wang X, et al. LLM Inference Scheduling: A Survey of Techniques, Frameworks, and Trade-offs.Authorea Preprints.2025

  21. [29]

    Cloud native system for llm inference serving.arXiv preprint arXiv:2507.18007

    Xu M, Liao J, Wu J, He Y, Ye K, Xu C. Cloud native system for llm inference serving.arXiv preprint arXiv:2507.18007. 2025

  22. [30]

    {NanoFlow}:Towardsoptimallargelanguagemodelservingthroughput.In:Proceedingsof the 19th USENIX Symposium on Operating Systems Design and Implementation

    ZhuK,GaoY,ZhaoY,etal. {NanoFlow}:Towardsoptimallargelanguagemodelservingthroughput.In:Proceedingsof the 19th USENIX Symposium on Operating Systems Design and Implementation. 2025:749–765

  23. [31]

    Flashinfer: Efficient and customizable attention engine for llm inference serving.Proceedings of Machine Learning and Systems.2025;7

    Ye Z, Chen L, Lai R, et al. Flashinfer: Efficient and customizable attention engine for llm inference serving.Proceedings of Machine Learning and Systems.2025;7

  24. [32]

    2025:446–461

    ZhangC,DuK,LiuS,etal.JENGA:EffectivememorymanagementforservingLLMwithheterogeneity.In:Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles. 2025:446–461. Xu ET AL. 39

  25. [33]

    throttll’em: Predictive gpu throttling for energy effi- cient llm inference serving

    Kakolyris AK, Masouros D, Vavaroutsos P, Xydis S, Soudris D. throttll’em: Predictive gpu throttling for energy effi- cient llm inference serving. In: Proceedings of the 2025 IEEE International Symposium on High Performance Computer Architecture. 2025:1363–1378

  26. [34]

    Collaborative inference and learning between edge slms and cloud LLMs: A survey of algorithms, execution, and open challenges.arXiv preprint arXiv:2507.16731.2025

    Li S, Wang H, Xu W, et al. Collaborative inference and learning between edge slms and cloud LLMs: A survey of algorithms, execution, and open challenges.arXiv preprint arXiv:2507.16731.2025

  27. [35]

    Extracting training data from large language models

    Carlini N, Tramer F, Wallace E, et al. Extracting training data from large language models. In: Proceedings of the 30th USENIX security symposium. 2021:2633–2650

  28. [36]

    Quantifying memorization across neural language models

    Carlini N, Ippolito D, Jagielski M, Lee K, Tramer F, Zhang C. Quantifying memorization across neural language models. In: Proceedings of the 11th International Conference on Learning Representations. 2022

  29. [37]

    Deep learning with differential privacy

    Abadi M, Chu A, Goodfellow I, et al. Deep learning with differential privacy. In: Proceedings of the 2016 ACM SIGSAC conference on computer and communications security. 2016:308–318

  30. [38]

    Communication-efficient learning of deep networks from decentralized data

    McMahan B, Moore E, Ramage D, Hampson S, Arcas yBA. Communication-efficient learning of deep networks from decentralized data. In: Artificial intelligence and statistics. 2017:1273–1282

  31. [39]

    Oblivious {Multi-Party} machine learning on trusted processors

    Ohrimenko O, Schuster F, Fournet C, et al. Oblivious {Multi-Party} machine learning on trusted processors. In: Proceedings of the 25th USENIX Security Symposium. 2016:619–636

  32. [40]

    Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection

    Greshake K, Abdelnabi S, Mishra S, Endres C, Holz T, Fritz M. Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection. In: Proceedings of the 16th ACM workshop on artificial intelligence and security. 2023:79–90

  33. [41]

    Stealing machine learning models via prediction APIs

    Tramèr F, Zhang F, Juels A, Reiter MK, Ristenpart T. Stealing machine learning models via prediction APIs. In: Proceedings of the 25th USENIX security symposium. 2016:601–618

  34. [42]

    Universal and transferable adversarial attacks on aligned language models.arXiv preprint arXiv:2307.15043.2023

    Zou A, Wang Z, Carlini N, Nasr M, Kolter JZ, Fredrikson M. Universal and transferable adversarial attacks on aligned language models.arXiv preprint arXiv:2307.15043.2023

  35. [43]

    ACMComputingSurveys

    DasBC,AminiMH,WuY.Securityandprivacychallengesoflargelanguagemodels:Asurvey. ACMComputingSurveys. 2025;57(6):1–39

  36. [44]

    IEEEInternet of Things Journal.2022;9(11):8364–8386

    BianJ,AlArafatA,XiongH,etal.Machinelearninginreal-timeInternetofThings(IoT)systems:Asurvey. IEEEInternet of Things Journal.2022;9(11):8364–8386

  37. [45]

    Spatial big data architecture: from data warehouses and data lakes to the Lakehouse.Journal of Parallel and Distributed Computing.2023;176:70–79

    Ait Errami S, Hajji H, Ait El Kadi K, Badir H. Spatial big data architecture: from data warehouses and data lakes to the Lakehouse.Journal of Parallel and Distributed Computing.2023;176:70–79

  38. [46]

    IEEETransactionsonKnowledgeand Data Engineering.2023;35(12):12571–12590

    HaiR,KoutrasC,QuixC,JarkeM.Datalakes:Asurveyoffunctionsandsystems. IEEETransactionsonKnowledgeand Data Engineering.2023;35(12):12571–12590

  39. [47]

    Retrieval augmented generation (rag) and beyond: A comprehensive survey on how to make your llms use external data more wisely.arXiv preprint arXiv:2409.14924.2024

    Zhao S, Yang Y, Wang Z, He Z, Qiu LK, Qiu L. Retrieval augmented generation (rag) and beyond: A comprehensive survey on how to make your llms use external data more wisely.arXiv preprint arXiv:2409.14924.2024

  40. [48]

    Computational cxl-memory solution for accelerating memory-intensive applications.IEEE Computer Architecture Letters.2022;22(1):5–8

    Sim J, Ahn S, Ahn T, et al. Computational cxl-memory solution for accelerating memory-intensive applications.IEEE Computer Architecture Letters.2022;22(1):5–8

  41. [49]

    Gu R, Xu Z, Che Y, et al. High-level data abstraction and elastic data caching for data-intensive ai applications on cloud- native platforms.IEEE Transactions on Parallel and Distributed Systems.2023;34(11):2946–2964

  42. [50]

    Retrieval-augmented generation for large language models: A survey.arXiv preprint arXiv:2312.10997.2023;2(1):32

    Gao Y, Xiong Y, Gao X, et al. Retrieval-augmented generation for large language models: A survey.arXiv preprint arXiv:2312.10997.2023;2(1):32

  43. [51]

    New Trends in High-D Vector Similarity Search: AI-driven, Progressive, and Distributed

    Echihabi K, Palpanas T, Zoumpatianos K. New Trends in High-D Vector Similarity Search: AI-driven, Progressive, and Distributed.. Proceedings of the VLDB Endowment.2021;14(12):3198–3201. 40 Xu ET AL

  44. [52]

    Unleash llms potential for sequential recommendation by coordinating dual dynamic index mechanism

    Yin J, Zeng Z, Li M, et al. Unleash llms potential for sequential recommendation by coordinating dual dynamic index mechanism. In: Proceedings of the 2025 ACM on Web Conference. 2025:216–227

  45. [53]

    LiH,FuF,LinS,etal.Hydraulis:BalancingLargeTransformerModelTrainingviaCo-designingParallelStrategiesand Data Assignment.Proceedings of the ACM on Management of Data.2025;3(6):1–30

  46. [54]

    A generative caching system for large language models.arXiv preprint arXiv:2503.17603.2025

    Iyengar A, Kundu A, Kompella R, Mamidi SN. A generative caching system for large language models.arXiv preprint arXiv:2503.17603.2025

  47. [55]

    Tracing the Data Trail: A Survey of Data Provenance, Transparency and Traceability in LLMs.arXiv preprint arXiv:2601.14311.2026

    Hohensinner R, Mutlu B, Estrada IGM, Vukovic M, Kopeinik S, Kern R. Tracing the Data Trail: A Survey of Data Provenance, Transparency and Traceability in LLMs.arXiv preprint arXiv:2601.14311.2026

  48. [56]

    In: Proceedings of the 41st IEEE/ACM International Conference on Computer-Aided Design

    AgostiniNB,CurzelS,AmatyaV,etal.AnMLIR-basedcompilerflowforsystem-leveldesignandhardwareacceleration. In: Proceedings of the 41st IEEE/ACM International Conference on Computer-Aided Design. 2022:1–9

  49. [57]

    2023:189–201

    MajumderK,BondhugulaU.Hir:Anmlir-basedintermediaterepresentationforhardwareacceleratordescription.In:Pro- ceedingsofthe28thACMInternationalConferenceonArchitecturalSupportforProgrammingLanguagesandOperating Systems, Volume 4. 2023:189–201

  50. [58]

    Packt Publishing Ltd, 2022

    Masood F, Brigoli R.Machine Learning on Kubernetes: A practical handbook for building and using a complete open source machine learning platform on Kubernetes. Packt Publishing Ltd, 2022

  51. [59]

    Llm-pilot: Characterize and optimize performance of your llm inference services

    Lazuka M, Anghel A, Parnell T. Llm-pilot: Characterize and optimize performance of your llm inference services. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE. 2024:1–18

  52. [60]

    Jegham N,Abdelatti M,Koh CY,Elmoubarki L,Hendawi A.How hungryis ai?benchmarking energy,water, andcarbon footprint of llm inference.arXiv preprint arXiv:2505.09598.2025

  53. [61]

    A survey on federated fine-tuning of large language models.arXiv preprint arXiv:2503.12016

    Wu Y, Tian C, Li J, et al. A survey on federated fine-tuning of large language models.arXiv preprint arXiv:2503.12016. 2025

  54. [62]

    IEEE Transactions on Computers.2026;75(4):1636-1649

    HuJ,XuM,YeK,XuC.BrownoutServe:SLO-AwareInferenceServingUnderBurstyWorkloadsforMoE-BasedLLMs. IEEE Transactions on Computers.2026;75(4):1636-1649. doi: 10.1109/TC.2026.3655019

  55. [63]

    DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving.IEEE Transactions on Services Computing.2026;19(2):1134-1147

    Liao J, Xu M, Zheng W, et al. DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving.IEEE Transactions on Services Computing.2026;19(2):1134-1147. doi: 10.1109/TSC.2026.3670011

  56. [64]

    ShuffleInfer: Disaggregate LLM inference for mixed downstream workloads.ACM Transactions on Architecture and Code Optimization.2025;22(2):1–24

    Hu C, Huang H, Xu L, et al. ShuffleInfer: Disaggregate LLM inference for mixed downstream workloads.ACM Transactions on Architecture and Code Optimization.2025;22(2):1–24

  57. [65]

    Splitwise: Efficient generative LLM inference using phase splitting

    Patel P, Choukse E, Zhang C, et al. Splitwise: Efficient generative LLM inference using phase splitting. In: Proceedings of the 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture. 2024:118–132

  58. [66]

    Fast distributed inference serving for large language models

    Wu B, Zhong Y, Zhang Z, et al. Fast distributed inference serving for large language models. arXiv preprint arXiv:2305.05920.2023

  59. [67]

    Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference.arXiv preprint arXiv:2508.19559.2025

    Li R, Du R, Chu Z, et al. Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference.arXiv preprint arXiv:2508.19559.2025

  60. [68]

    TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity.arXiv preprint arXiv:2512.03416.2025

    Lai R, Liu H, Lu C, et al. TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity.arXiv preprint arXiv:2512.03416.2025

  61. [69]

    LoongServe: Efficiently serving long-context large language models with elastic sequence parallelism

    Wu B, Liu S, Zhong Y, Sun P, Liu X, Jin X. LoongServe: Efficiently serving long-context large language models with elastic sequence parallelism. In: Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles. 2024:640–654. Xu ET AL. 41

  62. [70]

    dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving

    Wu B, Zhu R, Zhang Z, Sun P, Liu X, Jin X. dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving. In: Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation. 2024:911– 927

  63. [71]

    MegaScale-Infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism

    Zhu R, Jiang Z, Jin C, et al. MegaScale-Infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism. In: Proceedings of the ACM SIGCOMM 2025 Conference. 2025:592–608

  64. [72]

    Punica: Multi-tenant LoRA serving.Proceedings of Machine Learning and Systems.2024;6:1–13

    Chen L, Ye Z, Wu Y, Zhuo D, Ceze L, Krishnamurthy A. Punica: Multi-tenant LoRA serving.Proceedings of Machine Learning and Systems.2024;6:1–13

  65. [73]

    AlpaServe: Statistical multiplexing with model parallelism for deep learning serving

    Li Z, Zheng L, Zhong Y, et al. AlpaServe: Statistical multiplexing with model parallelism for deep learning serving. In: Proceedings of the 17th USENIX Symposium on Operating Systems Design and Implementation. 2023:663–679

  66. [74]

    arXivpreprint arXiv:2503.22562.2025

    GoelK,MohanJ,KwatraN,AnupindiRS,RamjeeR.Niyama:BreakingthesilosofLLMinferenceserving. arXivpreprint arXiv:2503.22562.2025

  67. [75]

    BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Dis- aggregated LLM Serving in AI Infrastructure

    He Y, Xu M, Wu J, et al. BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Dis- aggregated LLM Serving in AI Infrastructure. Software: Practice and Experience. 2026;56(4):424-444. doi: https://doi.org/10.1002/spe.70054

  68. [76]

    DynamoLLM: Designing LLM Inference Clusters for Perfor- manceandEnergyEfficiency.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture

    Stojkovic J, Zhang C, Goiri I, Torrellas J, Choukse E. DynamoLLM: Designing LLM Inference Clusters for Perfor- manceandEnergyEfficiency.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture. 2025:1348-1362

  69. [77]

    2024:207–222

    PatelP,ChoukseE,ZhangC,etal.CharacterizingpowermanagementopportunitiesforLLMsinthecloud.In:Proceedings ofthe29thACMInternationalConferenceonArchitecturalSupportforProgrammingLanguagesandOperatingSystems, Volume 3. 2024:207–222

  70. [78]

    In: Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2

    StojkovicJ,ZhangC,GoiriI,etal.TAPAS:Thermal-andPower-AwareSchedulingforLLMInferenceinCloudPlatforms. In: Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. Association for Computing Machinery...

  71. [79]

    AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding.arXiv preprint arXiv:2501.12162.2026

    Zikun L, Zhuofu C, Remi D, et al. AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding.arXiv preprint arXiv:2501.12162.2026

  72. [80]

    Fairness in serving large language models

    Sheng Y, Cao S, Li D, et al. Fairness in serving large language models. In: Proceedings of the 18th USENIX Conference on Operating Systems Design and Implementation. USENIX Association 2024; USA

  73. [81]

    DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-Based Clusters.IEEE Transactions on Services Computing.2024;17(6):3473-3484

    Bai H, Xu M, Ye K, Buyya R, Xu C. DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-Based Clusters.IEEE Transactions on Services Computing.2024;17(6):3473-3484. doi: 10.1109/TSC.2024.3433388

  74. [82]

    In: Proceedings of the 2023 USENIX Annual Technical Conference

    QiuH,MaoW,WangC,etal.AWARE:Automateworkloadautoscalingwithreinforcementlearninginproductioncloud systems. In: Proceedings of the 2023 USENIX Annual Technical Conference. 2023:387–402

  75. [83]

    2024:929–945

    LinC,HanZ,ZhangC,etal.Parrot:EfficientservingofLLM-basedapplicationswithsemanticvariable.In:Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation. 2024:929–945

  76. [84]

    Sun B, Pinciroli R, Casale G, Smirni E. DeepBAT: Performance and Cost Optimization of Serverless Inference Using Transformers.In:Proceedingsofthe2025IEEEInternationalParallelandDistributedProcessingSymposium.2025:335- 346

  77. [85]

    In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis

    MoZ,LiaoJ,XuH,ZhouZ,XuC.Hetis:ServingLLMsinHeterogeneousGPUClusterswithFine-grainedandDynamic Parallelism. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. Association for Computing Machinery 2025; New York, NY, ...

  78. [86]

    USENIX Association 2024; Santa Clara, CA:135–153

    FuY,XueL,HuangY,etal.ServerlessLLM:Low-LatencyServerlessInferenceforLargeLanguageModels.In:Proceed- ings of the 18th USENIX Symposium on Operating Systems Design and Implementation. USENIX Association 2024; Santa Clara, CA:135–153. 42 Xu ET AL

  79. [87]

    TAPAS: Thermal-and power-aware scheduling for LLM inference in cloud platforms

    Stojkovic J, Zhang C, Goiri I, et al. TAPAS: Thermal-and power-aware scheduling for LLM inference in cloud platforms. In: Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. 2025:1266–1281

  80. [88]

    throttLL’eM: Predictive GPU Throttling for Energy Effi- cientLLMInferenceServing.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture

    Kakolyris AK, Masouros D, Vavaroutsos P, Xydis S, Soudris D. throttLL’eM: Predictive GPU Throttling for Energy Effi- cientLLMInferenceServing.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture. 2025:1363-1378

  81. [89]

    Medusa: Accelerating Serverless LLM Inference with Materialization

    Zeng S, Xie M, Gao S, Chen Y, Lu Y. Medusa: Accelerating Serverless LLM Inference with Materialization. In: Pro- ceedingsofthe30thACMInternationalConferenceonArchitecturalSupportforProgrammingLanguagesandOperating Systems, Volume 1. Association for Computing Machinery 2025; Ne...

  82. [90]

    BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching

    Zhang D, Wang H, Liu Y, et al. BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching. In: Proceedings of the 19th USENIX Symposium on Operating Systems Design and Implementation. 2025:275–293

  83. [91]

    2025:415–430

    GimI,MaZ,LeeSs,ZhongL.Pie:Aprogrammableservingsystemforemergingllmapplications.In:Proceedingsofthe ACM SIGOPS 31st Symposium on Operating Systems Principles. 2025:415–430

  84. [92]

    HanY,KimI,KimJ,MoonGE.TensorCore-AdaptedSparseMatrixMultiplicationforAcceleratingSparseDeepNeural Networks.Electronics.2024;13(20):3981

  85. [93]

    FPGA-Based Sparse Matrix Multiplication Accelerators: From State-of-the-Art to Future Opportunities

    Liu Y, Chen R, Li S, others . FPGA-Based Sparse Matrix Multiplication Accelerators: From State-of-the-Art to Future Opportunities. ACM Transactions on Reconfigurable Technology and Systems.2024;17(4)

  86. [94]

    A survey on hardware accelerators for large language models.Applied Sciences.2025;15(2):586

    Kachris C. A survey on hardware accelerators for large language models.Applied Sciences.2025;15(2):586

  87. [95]

    arXiv preprint arXiv:2307.08191.2023

    LiangZ,ChengJ,YangR,etal.Unleashingthepotentialofllmsforquantumcomputing:Astudyinquantumarchitecture design. arXiv preprint arXiv:2307.08191.2023

  88. [96]

    Efficient large language models: A survey.arXiv preprint arXiv:2312.03863.2023

    Wan Z, Wang X, Liu C, et al. Efficient large language models: A survey.arXiv preprint arXiv:2312.03863.2023

  89. [97]

    HeY,FangJ,YuFR,LeungVC.Largelanguagemodels(LLMs)inferenceoffloadingandresourceallocationincloud-edge computing: An active inference approach.IEEE Transactions on Mobile Computing.2024;23(12):11253–11264

  90. [98]

    TangX,LiuF,XuD,etal.LLM-assistedreinforcementlearning:leveraginglightweightlargelanguagemodelcapabilities for efficient task scheduling in multi-cloud environment.IEEE Transactions on Consumer Electronics.2025;71(2):5631– 5644

  91. [99]

    LLM-based cost-aware task scheduling for cloud computing systems.Journal of Cloud Computing

    Pei H, Gu Y, Sun Y, et al. LLM-based cost-aware task scheduling for cloud computing systems.Journal of Cloud Computing. 2025;14(1):81

  92. [100]

    Green scheduling for LLM workloads with model and data reuse across geo-distributed data centers.Digital Communications and Networks.2025

    Liu H, Hu X, Wang R, Hao J, Wu Q, Zhang H. Green scheduling for LLM workloads with model and data reuse across geo-distributed data centers.Digital Communications and Networks.2025

  93. [101]

    Reliable and Efficient LLM Inference on Resource-Constrained Mobile Devices via Dynamic Scheduling.IEEE Internet of Things Journal.2026

    Liang A, Hong C, Li Y, Zhu H. Reliable and Efficient LLM Inference on Resource-Constrained Mobile Devices via Dynamic Scheduling.IEEE Internet of Things Journal.2026

  94. [102]

    Scalable Resource Provisioning Framework for Fog Computing Using LLM-Guided Q- Learning Approach.Algorithms

    Krishnamurthy B, Shiva SG. Scalable Resource Provisioning Framework for Fog Computing Using LLM-Guided Q- Learning Approach.Algorithms. 2025;18(4):230

  95. [103]

    Distributed training of large language models: A survey.Natural Language Processing Journal

    Zeng F, Gan W, Wang Y, Yu PS. Distributed training of large language models: A survey.Natural Language Processing Journal. 2025:100174

  96. [104]

    Atom: Low-bit quantization for efficient and accurate llm serving.Proceedings of Machine Learning and Systems.2024;6:196–209

    Zhao Y, Lin CY, Zhu K, et al. Atom: Low-bit quantization for efficient and accurate llm serving.Proceedings of Machine Learning and Systems.2024;6:196–209

  97. [105]

    Ladder: Enabling efficient{Low-Precision} deep learning computing through hardware- aware tensor transformation

    Wang L, Ma L, Cao S, et al. Ladder: Enabling efficient{Low-Precision} deep learning computing through hardware- aware tensor transformation. In: Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation. 2024:307–323. Xu ET AL. 43

  98. [106]

    Energy Efficient Software Hardware CoDesign for Machine Learning: From TinyML to Large Language Models.arXiv preprint arXiv:2603.23668.2026

    Vahdatpour MS, Zhang Y. Energy Efficient Software Hardware CoDesign for Machine Learning: From TinyML to Large Language Models.arXiv preprint arXiv:2603.23668.2026

  99. [107]

    arXivpreprintarXiv:2404.14294

    ZhouZ,NingX,HongK,etal.Asurveyonefficientinferenceforlargelanguagemodels. arXivpreprintarXiv:2404.14294. 2024

  100. [108]

    Renewable and Sustainable Energy Reviews.2026;225:116159

    JiZ,JiangM.Asystematicreviewofelectricitydemandforlargelanguagemodels:evaluations,challenges,andsolutions. Renewable and Sustainable Energy Reviews.2026;225:116159

  101. [109]

    In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    FernandezJ,NaC,TiwariV,BiskY,LuccioniS,StrubellE.Energyconsiderationsoflargelanguagemodelinferenceand efficiency optimizations. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025:32556–32569

  102. [110]

    Energy-Efficient Training and Inference in Large Language Models: Optimizing Computational and Energy Costs.International Journal of Computer Applications.2025;975:8887

    Narsepalle KR. Energy-Efficient Training and Inference in Large Language Models: Optimizing Computational and Energy Costs.International Journal of Computer Applications.2025;975:8887

  103. [111]

    WangG,CaiS,LiW,LyuD,HeG.OFQ-LLM:Outlier-FlexingQuantizationforEfficientLow-BitLargeLanguageModel Acceleration.IEEE Transactions on Circuits and Systems I: Regular Papers.2025

  104. [112]

    2025:285–309

    XuZ,LiF,ChenAC,WangX.ProCut:LLMPromptCompressionviaAttributionEstimation.In:Proceedingsofthe2025 Conference on Empirical Methods in Natural Language Processing: Industry Track. 2025:285–309

  105. [113]

    Parse trees guided LLM prompt compression.IEEE Transactions on Pattern Analysis and Machine Intelligence.2025

    Mao W, Hou C, Zhang T, Lin X, Tang K, Lv H. Parse trees guided LLM prompt compression.IEEE Transactions on Pattern Analysis and Machine Intelligence.2025

  106. [114]

    Promptware Engineering: Software Engineering for Prompt-Enabled Systems.ACM Transactions on Software Engineering and Methodology.2026

    Chen Z, Wang C, Sun W, Liu X, Zhang JM, Liu Y. Promptware Engineering: Software Engineering for Prompt-Enabled Systems.ACM Transactions on Software Engineering and Methodology.2026

  107. [115]

    Sufficient context: A new lens on retrieval augmented generation systems.arXiv preprint arXiv:2411.06037.2024

    Joren H, Zhang J, Ferng CS, Juan DC, Taly A, Rashtchian C. Sufficient context: A new lens on retrieval augmented generation systems.arXiv preprint arXiv:2411.06037.2024

  108. [116]

    Rankrag: Unifying context ranking with retrieval-augmented generation in llms.Proceedings of the Advances in Neural Information Processing Systems.2024;37:121156–121184

    Yu Y, Ping W, Liu Z, et al. Rankrag: Unifying context ranking with retrieval-augmented generation in llms.Proceedings of the Advances in Neural Information Processing Systems.2024;37:121156–121184

  109. [117]

    In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track

    LiZ,LiC,ZhangM,MeiQ,BenderskyM.Retrievalaugmentedgenerationorlong-contextllms?acomprehensivestudy and hybrid approach. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track. 2024:881–893

  110. [118]

    arXivpreprintarXiv:2502.12110

    XuW,LiangZ,MeiK,GaoH,TanJ,ZhangY.A-mem:Agenticmemoryforllmagents. arXivpreprintarXiv:2502.12110. 2025

  111. [119]

    Aiopslab: A holistic framework to evaluate ai agents for enabling autonomous clouds

    Chen Y, Shetty M, Somashekar G, et al. Aiopslab: A holistic framework to evaluate ai agents for enabling autonomous clouds. Proceedings of Machine Learning and Systems.2025;7

  112. [120]

    Lads: Leveraging llms for ai-driven devops.arXiv preprint arXiv:2502.20825

    Khan AF, Khan AA, Mohamed A, et al. Lads: Leveraging llms for ai-driven devops.arXiv preprint arXiv:2502.20825. 2025

  113. [121]

    The Multi-Agent Fault Localization System Based on Monte Carlo Tree Search Approach.arXiv preprint arXiv:2507.22800.2025

    Ren R. The Multi-Agent Fault Localization System Based on Monte Carlo Tree Search Approach.arXiv preprint arXiv:2507.22800.2025

  114. [122]

    Parallax: Efficient llm inference service over decentralized environment.arXiv preprint arXiv:2509.26182.2025

    Tong C, Jiang Y, Chen G, et al. Parallax: Efficient llm inference service over decentralized environment.arXiv preprint arXiv:2509.26182.2025

  115. [123]

    Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput.arXiv preprint arXiv:2511.11733.2025

    Song J, Chen W, Song X, et al. Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput.arXiv preprint arXiv:2511.11733.2025

  116. [124]

    IEEEComputerArchitectureLetters

    WangX,SunX,LiW,etal.Low-LatencyPIMAcceleratorforEdgeLLMInference. IEEEComputerArchitectureLetters. 2025. 44 Xu ET AL

  117. [125]

    Decentralized llm inference over edge networks with energy harvesting

    Khoshsirat A, Perin G, Rossi M. Decentralized llm inference over edge networks with energy harvesting. In: Proceedings of the IEEE Global Communications Conference. 2024:3703–3708

  118. [126]

    Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic.arXiv preprint arXiv:2601.21972.2026

    Liu S, Chen T, Amiri R, Amato C. Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic.arXiv preprint arXiv:2601.21972.2026

  119. [127]

    arXivpreprintarXiv:2106.09685

    HuEJ,ShenY,WallisP,etal.LoRA:Low-RankAdaptationofLargeLanguageModels. arXivpreprintarXiv:2106.09685. 2021

  120. [128]

    Safe-FedLLM: Delving into the Safety of Federated Large Language Models

    Tao M, Tian Y, Tu W, Yang Y, Yang X, Tang X. Safe-FedLLM: Delving into the Safety of Federated Large Language Models. arXiv preprint arXiv:2601.07177.2026

  121. [129]

    Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference.arXiv preprint arXiv:2512.16317.2025

    Tian A, Ding A, Chen F, Wu A, Chan A, Zhang B. Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference.arXiv preprint arXiv:2512.16317.2025

  122. [130]

    Findingsofthe Association for Computational Linguistics: EMNLP.2025:12841–12854

    MoonH,SeoJ,KooS,etal.LimaCost:DataValuationforInstructionTuningofLargeLanguageModels. Findingsofthe Association for Computational Linguistics: EMNLP.2025:12841–12854

  123. [131]

    Fairshare Data Pricing via Data Valuation for Large Language Models.arXiv preprint arXiv:2502.00198.2025

    Zhang L, Jiao C, Li B, Xiong C. Fairshare Data Pricing via Data Valuation for Large Language Models.arXiv preprint arXiv:2502.00198.2025

  124. [132]

    Incentivizing permissionless distributed learning of llms

    Lidin J, Sarfi A, Pappas E, Dare S, Belilovsky E, Steeves J. Incentivizing permissionless distributed learning of llms. In: Proceedings of the 2025 7th International Conference on Distributed Artificial Intelligence. 2025:12–18

  125. [133]

    Incentive mechanism of foundation model enabled cross-silo federated learning

    Zhang N, Xu X, Liu X, Wu J, Tang H. Incentive mechanism of foundation model enabled cross-silo federated learning. Scientific Reports.2025;15(1):24181

  126. [134]

    WinFLoRA: Incentivizing Client-Adaptive Aggregation in Federated LoRA under Privacy Heterogeneity.arXiv preprint arXiv:2602.01126.2026

    Kou M, Xia X, Wang Z, et al. WinFLoRA: Incentivizing Client-Adaptive Aggregation in Federated LoRA under Privacy Heterogeneity.arXiv preprint arXiv:2602.01126.2026

  127. [135]

    ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction.arXiv preprint arXiv:2602.12247.2026

    Ferguson N, Pennington J, Beghian N, et al. ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction.arXiv preprint arXiv:2602.12247.2026

  128. [136]

    2025:2749– 2763

    ShrimalA,JainA,ChowdhuryS,YenigallaP.PARSE:LLMdrivenschemaoptimizationforreliableentityextraction.In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track. 2025:2749– 2763

  129. [137]

    DELM: a Python toolkit for Data Extraction with Language Models

    Fithian E, Skobelev K. DELM: a Python toolkit for Data Extraction with Language Models. arXiv preprint arXiv:2509.20617.2025

  130. [138]

    LLMStructBench: Benchmarking Large Language Model Structured Data Extraction.arXiv preprint arXiv:2602.14743.2026

    Tenckhoff S, Koddenbrock M, Rodner E. LLMStructBench: Benchmarking Large Language Model Structured Data Extraction.arXiv preprint arXiv:2602.14743.2026

  131. [139]

    Structure-Aware Decoding Mechanisms for Complex Entity Extraction with Large-Scale Language Models.arXiv preprint arXiv:2512.13980.2025

    Qiu Z, Wu D, Liu F, Wang Y. Structure-Aware Decoding Mechanisms for Complex Entity Extraction with Large-Scale Language Models.arXiv preprint arXiv:2512.13980.2025

  132. [140]

    Manalyzer:End-to-end Automated Meta-analysis withMulti-agent System.arXiv preprint arXiv:2505.20310.2025

    Xu W, ZhangW, Ling F,et al. Manalyzer:End-to-end Automated Meta-analysis withMulti-agent System.arXiv preprint arXiv:2505.20310.2025

  133. [141]

    Reducing hallucinations of large language models via hierarchical semantic piece.Complex & Intelligent Systems.2025;11(5):231

    Liu Y, Yang Q, Tang J, et al. Reducing hallucinations of large language models via hierarchical semantic piece.Complex & Intelligent Systems.2025;11(5):231

  134. [142]

    In: Proceedings of the 10th Workshop on Noisy and User-generated Text

    BirkmoseR,ReeceNM,NorvinEH,BjervaJ,ZhangM.On-deviceLLMsforhomeassistant:Dualroleinintentdetection and response generation. In: Proceedings of the 10th Workshop on Noisy and User-generated Text. 2025:57–67

  135. [143]

    Vega: LLM-driven intelligent chatbot platform for internet of things control and development.Sensors.2025;25(12):3809

    Al-Safi H, Ibrahim H, Steenson P. Vega: LLM-driven intelligent chatbot platform for internet of things control and development.Sensors.2025;25(12):3809

  136. [144]

    IEEE Open Journal of the Computer Society.2025

    LinYW,LinYB,WuYF,ShenPH.Voicetalk:Ano-codeapproachforcreatingvoice-controlledsmarthomeapplications. IEEE Open Journal of the Computer Society.2025. Xu ET AL. 45

  137. [145]

    Introducing LEAF: LLM Edge Assessment Framework for Generative AI on the Edge

    Abdulkadhim M, Repas SR. Introducing LEAF: LLM Edge Assessment Framework for Generative AI on the Edge. Machine Learning and Knowledge Extraction.2026;8(2):48

  138. [146]

    arXivpreprintarXiv:2409.00088

    XuJ,LiZ,ChenW,etal.On-devicelanguagemodels:Acomprehensivereview. arXivpreprintarXiv:2409.00088. 2024

  139. [147]

    LLM-Driven Auto Configuration for Transient IoT Device Collaboration

    Shastri H, Hanafy W, Wu L, Irwin D, Srivastava M, Shenoy P. LLM-Driven Auto Configuration for Transient IoT Device Collaboration. In: Proceedings of the 10th ACM/IEEE Symposium on Edge Computing. 2025:1–17

  140. [148]

    Instant Personalized Large Language Model Adaptation via Hypernetwork.arXiv preprint arXiv:2510.16282.2025

    Tan Z, Zhang Z, Wen H, et al. Instant Personalized Large Language Model Adaptation via Hypernetwork.arXiv preprint arXiv:2510.16282.2025

  141. [149]

    Embedding-to-Prefix: Parameter-efficient personalization for pre-trained large language models.arXiv preprint arXiv:2505.17051.2025

    Huber B, Fazelnia G, Damianou A, et al. Embedding-to-Prefix: Parameter-efficient personalization for pre-trained large language models.arXiv preprint arXiv:2505.17051.2025

  142. [150]

    Ensuring fair llm serving amid diverse applications.arXiv preprint arXiv:2411.15997

    Khan RIS, Jain K, Shen H, et al. Ensuring fair llm serving amid diverse applications.arXiv preprint arXiv:2411.15997. 2024

  143. [151]

    T-POP: Test-Time Personalization with Online Preference Feedback.arXiv preprint arXiv:2509.24696.2025

    Qu Z, Zhang M, Kong M, et al. T-POP: Test-Time Personalization with Online Preference Feedback.arXiv preprint arXiv:2509.24696.2025

Pith tools

Reviewed May 10, 2026 · model on record in the stance chip above.