REVIEW 2 major objections 2 minor 2 cited by
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
T0 review · 2 major / 2 minor · reviewed 2026-05-10 · grok-4.3
Pith's one-line read Large language models depend on cloud-native and distributed systems for efficient scaling.
desk verdict This is a standard research agenda that lists familiar LLM scaling problems and names cloud trends without new mechanisms, data, or resolved questions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Cloud-native and distributed architectures, which enable handling of LLM computational demands through features like microservices, autoscaling, and hybrid cloud-edge solutions.
What would settle it
A demonstration that a large-scale LLM can be trained and deployed efficiently using only traditional non-cloud, non-distributed infrastructure.
Extended reading notes
Core claim
The paper claims that cloud platforms and distributed systems play a key role in supporting the scalability, efficiency, and optimization of large language models. The complexities of LLM deployment involve data management, resource optimization, and the adoption of microservices, autoscaling, and hybrid cloud-edge solutions. Emerging research trends such as serverless inference, quantum computing, and federated learning hold potential to advance LLM capabilities further. A roadmap for future work calls for ongoing research, standardization efforts, and collaboration between sectors to support LLM expansion in research and enterprise settings.
Load-bearing premise
That traditional systems are unable to meet the computational requirements of large language models and that the listed emerging trends will drive the next phase of innovation.
Editorial extensions
If this is right
- LLM deployment can incorporate data management and resource optimization techniques from cloud systems.
- Hybrid cloud-edge solutions will address specific deployment complexities.
- Serverless inference, quantum computing, and federated learning represent promising directions for next innovations.
- Continued research, standardization, and cross-sector collaboration are required to sustain LLM growth.
Reading between the lines
- Organizations without large data centers could leverage these systems to develop competitive LLMs.
- The emphasis on federated learning may enable privacy-preserving model training across distributed data sources.
- Quantum computing integration could transform energy efficiency in model operations if realized.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that the computational demands of Large Language Models (LLMs) necessitate the integration of cloud-native and distributed architectures, as traditional systems are often unable to meet these requirements. It explores complexities in LLM deployment such as data management, resource optimization, microservices, autoscaling, and hybrid cloud-edge solutions. The manuscript examines emerging research trends including serverless inference, quantum computing, and federated learning, and concludes with a roadmap for future developments emphasizing continued research, standardization, and cross-sector collaboration.
Significance. If pursued, this agenda could help direct research toward more scalable and efficient LLM systems by highlighting the intersection of distributed computing and AI challenges. The paper's strength is in providing a high-level synthesis of deployment issues and forward-looking trends, which may stimulate targeted follow-up studies, though its impact as a position paper rests on the actionability of the proposed directions rather than new derivations or data.
major comments (2)
- [Abstract] Abstract: The foundational assertion that 'Traditional systems are often unable to meet these requirements' is presented without any specific benchmarks, scaling examples, or citations to LLM performance bottlenecks, which is load-bearing for the motivation to integrate cloud-native architectures.
- [Abstract] Abstract (emerging trends paragraph): The statement that trends such as serverless inference, quantum computing, and federated learning 'will drive the next phase of LLM innovation' is made without analysis of their current maturity, specific applicability to LLM training/inference, or preliminary evidence, which underpins the credibility of the concluding roadmap.
minor comments (2)
- The discussion of deployment complexities would benefit from clearer organization, such as explicit subsection headings for data management versus resource optimization, to improve readability.
- Additional citations to recent work on distributed LLM training frameworks would help ground the high-level claims in existing literature.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback and the recommendation for minor revision. We address the two points on the abstract below and will incorporate targeted clarifications to strengthen the grounding of our claims while preserving the high-level nature of this research agenda paper.
read point-by-point responses
-
Referee: [Abstract] Abstract: The foundational assertion that 'Traditional systems are often unable to meet these requirements' is presented without any specific benchmarks, scaling examples, or citations to LLM performance bottlenecks, which is load-bearing for the motivation to integrate cloud-native architectures.
Authors: We agree that the abstract would benefit from explicit support for this statement. The full manuscript already references scaling challenges such as quadratic attention complexity and the multi-node requirements for training models with hundreds of billions of parameters. In revision, we will add one or two concise citations (e.g., to transformer scaling laws and reported training infrastructure needs) directly in the abstract to anchor the claim without expanding its length or shifting the paper's position-paper character. revision: yes
-
Referee: [Abstract] Abstract (emerging trends paragraph): The statement that trends such as serverless inference, quantum computing, and federated learning 'will drive the next phase of LLM innovation' is made without analysis of their current maturity, specific applicability to LLM training/inference, or preliminary evidence, which underpins the credibility of the concluding roadmap.
Authors: This observation is valid for the abstract's brevity. The body of the manuscript discusses maturity levels and applicability in more detail (e.g., serverless for inference workloads, federated learning for privacy-preserving training). For the revision, we will soften the phrasing to 'emerging trends with the potential to drive...' and insert a short qualifier noting their varying stages of readiness, thereby improving credibility while keeping the abstract concise and forward-looking. revision: partial
Circularity Check
No significant circularity
full rationale
The manuscript is a descriptive research agenda surveying LLM challenges and advocating cloud-native/distributed approaches plus emerging trends. It contains no equations, formal derivations, fitted parameters, or predictions that reduce to inputs by construction. Central claims are high-level motivations and calls for future work rather than asserted results resting on self-referential steps, self-citations, or renamed empirical patterns. The text is therefore self-contained with no load-bearing circular reductions.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda." pith.science (2026). https://pith.science/paper/2604.17227
@misc{pith2026260417227,
author = {Pith},
title = {Pith review of: Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda},
year = {2026},
howpublished = {\url{https://pith.science/paper/2604.17227}},
note = {Machine review of arXiv:2604.17227}
}
read the original abstract
The rapid rise of Large Language Models (LLMs) has revolutionized various artificial intelligence (AI) applications, from natural language processing to code generation. However, the computational demands of these models, particularly in training and inference, present significant challenges. Traditional systems are often unable to meet these requirements, necessitating the integration of cloud-native and distributed architectures. This paper explores the role of cloud platforms and distributed systems in supporting the scalability, efficiency, and optimization of LLMs. We discuss the complexities of LLM deployment, including data management, resource optimization, and the need for microservices, autoscaling, and hybrid cloud-edge solutions. Additionally, we examine emerging research trends, such as serverless inference, quantum computing, and federated learning, and their potential to drive the next phase of LLM innovation. The paper concludes with a roadmap for future developments, emphasizing the need for continued research, standardization, and cross-sector collaboration to sustain the growth of LLMs in both research and enterprise applications.
Forward citations
Cited by 2 Pith papers
-
CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving
Layer-wise data-parallel replication of hot Transformer layers onto reclaimed idle GPUs reduces LLM serving cold-start latency 97.9–99.3% and average latency 20.7–28.1% while attaining 100% SLO on production traces.
-
BrownoutMoE: Structure-Aware Expert Grouping for Efficient and Accurate LLM Web-based Services
GRPO-optimized expert grouping plus grouping-consistent distillation reduces MoE brownout accuracy degradation by up to 71.4% and raises throughput by up to 2.24× versus sequential grouping.
Reference graph
Works this paper leans on
-
[1]
In: Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation
ZhongY,LiuS,ChenJ,etal.DistServe:Disaggregatingprefillanddecodingforgoodput-optimizedlargelanguagemodel serving. In: Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation. 2024:193– 210
work page 2024
-
[2]
Pahl C, Brogi A, Soldani J, Jamshidi P. Cloud container technologies: a state-of-the-art review.IEEE Transactions on Cloud Computing.2017;7(3):677–692
work page 2017
-
[3]
Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. In: Proceedings of the Advances in neural information processing systems. 2017
work page 2017
-
[4]
Scaling Laws for Neural Language Models
Kaplan J, McCandlish S, Henighan T, et al. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361. 2020
work page Pith review arXiv 2001
-
[5]
Deep learning.Nature.2015;521(7553):436–444
LeCun Y, Bengio Y, Hinton G. Deep learning.Nature.2015;521(7553):436–444
work page 2015
-
[6]
IsaevM,McDonaldN,VuducR.Scalinginfrastructuretosupportmulti-trillionparameterLLMtraining.In:Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 2023
work page 2023
-
[7]
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.; 2020
Shoeybi M, Patwary M, Puri R, LeGresley P, Casper J, Catanzaro B. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.; 2020
work page 2020
-
[8]
ZhangY,YouY.SpeedLoader:AnI/OefficientschemeforheterogeneousanddistributedLLMoperation.In:Proceedings of the Advances in Neural Information Processing Systems. Curran Associates, Inc. 2024:34637–34655
work page 2024
Show all 151 references
-
[9]
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Aminabadi RY, Rajbhandari S, Awan AA, et al. Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 2022:1–15
2022
-
[10]
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Dao T, Fu D, Ermon S, Rudra A, Ré C. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In: Proceedings of the Advances in Neural Information Processing Systems. Curran Associates, Inc. 2022:16344–16359
2022
-
[11]
Towards end-to-end optimization of llm-based applications with ayo
Tan X, Jiang Y, Yang Y, Xu H. Towards end-to-end optimization of llm-based applications with ayo. In: Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. 2025:1302–1316
2025
-
[12]
In-datacenter performance analysis of a tensor processing unit
Jouppi NP, Young C, Patil N, et al. In-datacenter performance analysis of a tensor processing unit. In: Proceedings of the 44th annual international symposium on computer architecture. 2017:1–12
2017
-
[13]
Alpa: Automating Inter-and Intra-Operator Parallelism for Distributed Deep Learning
Zheng L, Li Z, Zhang H, et al. Alpa: Automating Inter-and Intra-Operator Parallelism for Distributed Deep Learning. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation. 2022. 38 Xu ET AL
2022
-
[14]
Efficient memory management for large language model serving with pagedattention
Kwon W, Li Z, Zhuang S, et al. Efficient memory management for large language model serving with pagedattention. In: Proceedings of the 29th symposium on operating systems principles. 2023:611–626
2023
-
[15]
Open issues in scheduling microservices in the cloud.IEEE Cloud Computing.2016;3(5):81–88
Fazio M, Celesti A, Ranjan R, Liu C, Chen L, Villari M. Open issues in scheduling microservices in the cloud.IEEE Cloud Computing.2016;3(5):81–88
2016
-
[16]
TheKubernetesAuthors.Kubernetes:Production-GradeContainerOrchestration.https://kubernetes.io;.Accessed:2026- 04-08
2026
-
[17]
Flexgen: High-throughput generative inference of large language models with a single gpu
Sheng Y, Zheng L, Yuan B, et al. Flexgen: High-throughput generative inference of large language models with a single gpu. In: Proceedings of the International Conference on Machine Learning. 2023:31094–31116
2023
-
[18]
In: Proceedings of the 16th USENIX symposium on operating systems design and implementation
YuGI,JeongJS,KimGW,KimS,ChunBG.Orca:AdistributedservingsystemforTransformer-Basedgenerativemodels. In: Proceedings of the 16th USENIX symposium on operating systems design and implementation. 2022:521–538
2022
-
[19]
Evaluation and benchmarking of llm agents: A survey
Mohammadi M, Li Y, Lo J, Yip W. Evaluation and benchmarking of llm agents: A survey. In: Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 2025:6129–6139
2025
-
[20]
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
Chen Y, Cui W, Zhao H, et al. Towards High-Goodput LLM Serving with Prefill-decode Multiplexing. In: Proceedings ofthe31stACMInternationalConferenceonArchitecturalSupportforProgrammingLanguagesandOperatingSystems, Volume 2. 2026:2030–2047
2026
-
[21]
Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
Rombaut B, Masoumzadeh S, Vasilevski K, Lin D, Hassan AE. Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents. In: Proceedings of the 40th IEEE/ACM International Conference on Automated Software Engineering. 2025:739–751
2025
-
[22]
Process modeling in web applications.ACM Transactions on Software Engineering and Methodology.2006;15(4):360–409
Brambilla M, Ceri S, Fraternali P, Manolescu I. Process modeling in web applications.ACM Transactions on Software Engineering and Methodology.2006;15(4):360–409
2006
-
[23]
Challenges in deployment and configuration management in cyber physical system
Jha DN, Li Y, Jayaraman PP, et al. Challenges in deployment and configuration management in cyber physical system. Handbook of Integration of Cloud Computing, Cyber Physical Systems and Internet of Things.2020:215–235
2020
-
[24]
Parallel processing systems for big data: a survey
Zhang Y, Cao T, Li S, et al. Parallel processing systems for big data: a survey. Proceedings of the IEEE. 2016;104(11):2114–2136
2016
-
[25]
LiuX,BuyyaR.Resourcemanagementandschedulingindistributedstreamprocessingsystems:ataxonomy,review,and future directions.ACM Computing Surveys.2020;53(3):1–41
2020
-
[26]
Deep learning workload scheduling in gpu datacenters: A survey.ACM Computing Surveys
Ye Z, Gao W, Hu Q, et al. Deep learning workload scheduling in gpu datacenters: A survey.ACM Computing Surveys. 2024;56(6):1–38
2024
-
[27]
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications.ACM Trans
Wen L, Xu M, Gill SS, et al. StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications.ACM Trans. Auton. Adapt. Syst..2025;20(1). doi: 10.1145/3686253
2025 doi
-
[28]
LLM Inference Scheduling: A Survey of Techniques, Frameworks, and Trade-offs.Authorea Preprints.2025
Heisler M, Yousefijamarani Z, Wang X, et al. LLM Inference Scheduling: A Survey of Techniques, Frameworks, and Trade-offs.Authorea Preprints.2025
2025
-
[29]
Cloud native system for llm inference serving.arXiv preprint arXiv:2507.18007
Xu M, Liao J, Wu J, He Y, Ye K, Xu C. Cloud native system for llm inference serving.arXiv preprint arXiv:2507.18007. 2025
2025
-
[30]
{NanoFlow}:Towardsoptimallargelanguagemodelservingthroughput.In:Proceedingsof the 19th USENIX Symposium on Operating Systems Design and Implementation
ZhuK,GaoY,ZhaoY,etal. {NanoFlow}:Towardsoptimallargelanguagemodelservingthroughput.In:Proceedingsof the 19th USENIX Symposium on Operating Systems Design and Implementation. 2025:749–765
2025
-
[31]
Flashinfer: Efficient and customizable attention engine for llm inference serving.Proceedings of Machine Learning and Systems.2025;7
Ye Z, Chen L, Lai R, et al. Flashinfer: Efficient and customizable attention engine for llm inference serving.Proceedings of Machine Learning and Systems.2025;7
2025
-
[32]
2025:446–461
ZhangC,DuK,LiuS,etal.JENGA:EffectivememorymanagementforservingLLMwithheterogeneity.In:Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles. 2025:446–461. Xu ET AL. 39
2025
-
[33]
throttll’em: Predictive gpu throttling for energy effi- cient llm inference serving
Kakolyris AK, Masouros D, Vavaroutsos P, Xydis S, Soudris D. throttll’em: Predictive gpu throttling for energy effi- cient llm inference serving. In: Proceedings of the 2025 IEEE International Symposium on High Performance Computer Architecture. 2025:1363–1378
2025
-
[34]
Collaborative inference and learning between edge slms and cloud LLMs: A survey of algorithms, execution, and open challenges.arXiv preprint arXiv:2507.16731.2025
Li S, Wang H, Xu W, et al. Collaborative inference and learning between edge slms and cloud LLMs: A survey of algorithms, execution, and open challenges.arXiv preprint arXiv:2507.16731.2025
2025
-
[35]
Extracting training data from large language models
Carlini N, Tramer F, Wallace E, et al. Extracting training data from large language models. In: Proceedings of the 30th USENIX security symposium. 2021:2633–2650
2021
-
[36]
Quantifying memorization across neural language models
Carlini N, Ippolito D, Jagielski M, Lee K, Tramer F, Zhang C. Quantifying memorization across neural language models. In: Proceedings of the 11th International Conference on Learning Representations. 2022
2022
-
[37]
Deep learning with differential privacy
Abadi M, Chu A, Goodfellow I, et al. Deep learning with differential privacy. In: Proceedings of the 2016 ACM SIGSAC conference on computer and communications security. 2016:308–318
2016
-
[38]
Communication-efficient learning of deep networks from decentralized data
McMahan B, Moore E, Ramage D, Hampson S, Arcas yBA. Communication-efficient learning of deep networks from decentralized data. In: Artificial intelligence and statistics. 2017:1273–1282
2017
-
[39]
Oblivious {Multi-Party} machine learning on trusted processors
Ohrimenko O, Schuster F, Fournet C, et al. Oblivious {Multi-Party} machine learning on trusted processors. In: Proceedings of the 25th USENIX Security Symposium. 2016:619–636
2016
-
[40]
Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection
Greshake K, Abdelnabi S, Mishra S, Endres C, Holz T, Fritz M. Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection. In: Proceedings of the 16th ACM workshop on artificial intelligence and security. 2023:79–90
2023
-
[41]
Stealing machine learning models via prediction APIs
Tramèr F, Zhang F, Juels A, Reiter MK, Ristenpart T. Stealing machine learning models via prediction APIs. In: Proceedings of the 25th USENIX security symposium. 2016:601–618
2016
-
[42]
Universal and transferable adversarial attacks on aligned language models.arXiv preprint arXiv:2307.15043.2023
Zou A, Wang Z, Carlini N, Nasr M, Kolter JZ, Fredrikson M. Universal and transferable adversarial attacks on aligned language models.arXiv preprint arXiv:2307.15043.2023
2023 arXiv
-
[43]
ACMComputingSurveys
DasBC,AminiMH,WuY.Securityandprivacychallengesoflargelanguagemodels:Asurvey. ACMComputingSurveys. 2025;57(6):1–39
2025
-
[44]
IEEEInternet of Things Journal.2022;9(11):8364–8386
BianJ,AlArafatA,XiongH,etal.Machinelearninginreal-timeInternetofThings(IoT)systems:Asurvey. IEEEInternet of Things Journal.2022;9(11):8364–8386
2022
-
[45]
Spatial big data architecture: from data warehouses and data lakes to the Lakehouse.Journal of Parallel and Distributed Computing.2023;176:70–79
Ait Errami S, Hajji H, Ait El Kadi K, Badir H. Spatial big data architecture: from data warehouses and data lakes to the Lakehouse.Journal of Parallel and Distributed Computing.2023;176:70–79
2023
-
[46]
IEEETransactionsonKnowledgeand Data Engineering.2023;35(12):12571–12590
HaiR,KoutrasC,QuixC,JarkeM.Datalakes:Asurveyoffunctionsandsystems. IEEETransactionsonKnowledgeand Data Engineering.2023;35(12):12571–12590
2023
-
[47]
Retrieval augmented generation (rag) and beyond: A comprehensive survey on how to make your llms use external data more wisely.arXiv preprint arXiv:2409.14924.2024
Zhao S, Yang Y, Wang Z, He Z, Qiu LK, Qiu L. Retrieval augmented generation (rag) and beyond: A comprehensive survey on how to make your llms use external data more wisely.arXiv preprint arXiv:2409.14924.2024
2024
-
[48]
Computational cxl-memory solution for accelerating memory-intensive applications.IEEE Computer Architecture Letters.2022;22(1):5–8
Sim J, Ahn S, Ahn T, et al. Computational cxl-memory solution for accelerating memory-intensive applications.IEEE Computer Architecture Letters.2022;22(1):5–8
2022
-
[49]
Gu R, Xu Z, Che Y, et al. High-level data abstraction and elastic data caching for data-intensive ai applications on cloud- native platforms.IEEE Transactions on Parallel and Distributed Systems.2023;34(11):2946–2964
2023
-
[50]
Retrieval-augmented generation for large language models: A survey.arXiv preprint arXiv:2312.10997.2023;2(1):32
Gao Y, Xiong Y, Gao X, et al. Retrieval-augmented generation for large language models: A survey.arXiv preprint arXiv:2312.10997.2023;2(1):32
2023 arXiv
-
[51]
New Trends in High-D Vector Similarity Search: AI-driven, Progressive, and Distributed
Echihabi K, Palpanas T, Zoumpatianos K. New Trends in High-D Vector Similarity Search: AI-driven, Progressive, and Distributed.. Proceedings of the VLDB Endowment.2021;14(12):3198–3201. 40 Xu ET AL
2021
-
[52]
Unleash llms potential for sequential recommendation by coordinating dual dynamic index mechanism
Yin J, Zeng Z, Li M, et al. Unleash llms potential for sequential recommendation by coordinating dual dynamic index mechanism. In: Proceedings of the 2025 ACM on Web Conference. 2025:216–227
2025
-
[53]
LiH,FuF,LinS,etal.Hydraulis:BalancingLargeTransformerModelTrainingviaCo-designingParallelStrategiesand Data Assignment.Proceedings of the ACM on Management of Data.2025;3(6):1–30
2025
-
[54]
A generative caching system for large language models.arXiv preprint arXiv:2503.17603.2025
Iyengar A, Kundu A, Kompella R, Mamidi SN. A generative caching system for large language models.arXiv preprint arXiv:2503.17603.2025
2025
-
[55]
Tracing the Data Trail: A Survey of Data Provenance, Transparency and Traceability in LLMs.arXiv preprint arXiv:2601.14311.2026
Hohensinner R, Mutlu B, Estrada IGM, Vukovic M, Kopeinik S, Kern R. Tracing the Data Trail: A Survey of Data Provenance, Transparency and Traceability in LLMs.arXiv preprint arXiv:2601.14311.2026
2026
-
[56]
In: Proceedings of the 41st IEEE/ACM International Conference on Computer-Aided Design
AgostiniNB,CurzelS,AmatyaV,etal.AnMLIR-basedcompilerflowforsystem-leveldesignandhardwareacceleration. In: Proceedings of the 41st IEEE/ACM International Conference on Computer-Aided Design. 2022:1–9
2022
-
[57]
2023:189–201
MajumderK,BondhugulaU.Hir:Anmlir-basedintermediaterepresentationforhardwareacceleratordescription.In:Pro- ceedingsofthe28thACMInternationalConferenceonArchitecturalSupportforProgrammingLanguagesandOperating Systems, Volume 4. 2023:189–201
2023
-
[58]
Packt Publishing Ltd, 2022
Masood F, Brigoli R.Machine Learning on Kubernetes: A practical handbook for building and using a complete open source machine learning platform on Kubernetes. Packt Publishing Ltd, 2022
2022
-
[59]
Llm-pilot: Characterize and optimize performance of your llm inference services
Lazuka M, Anghel A, Parnell T. Llm-pilot: Characterize and optimize performance of your llm inference services. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE. 2024:1–18
2024
-
[60]
Jegham N,Abdelatti M,Koh CY,Elmoubarki L,Hendawi A.How hungryis ai?benchmarking energy,water, andcarbon footprint of llm inference.arXiv preprint arXiv:2505.09598.2025
2025
-
[61]
A survey on federated fine-tuning of large language models.arXiv preprint arXiv:2503.12016
Wu Y, Tian C, Li J, et al. A survey on federated fine-tuning of large language models.arXiv preprint arXiv:2503.12016. 2025
2025
-
[62]
IEEE Transactions on Computers.2026;75(4):1636-1649
HuJ,XuM,YeK,XuC.BrownoutServe:SLO-AwareInferenceServingUnderBurstyWorkloadsforMoE-BasedLLMs. IEEE Transactions on Computers.2026;75(4):1636-1649. doi: 10.1109/TC.2026.3655019
2026 doi
-
[63]
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving.IEEE Transactions on Services Computing.2026;19(2):1134-1147
Liao J, Xu M, Zheng W, et al. DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving.IEEE Transactions on Services Computing.2026;19(2):1134-1147. doi: 10.1109/TSC.2026.3670011
2026 doi
-
[64]
ShuffleInfer: Disaggregate LLM inference for mixed downstream workloads.ACM Transactions on Architecture and Code Optimization.2025;22(2):1–24
Hu C, Huang H, Xu L, et al. ShuffleInfer: Disaggregate LLM inference for mixed downstream workloads.ACM Transactions on Architecture and Code Optimization.2025;22(2):1–24
2025
-
[65]
Splitwise: Efficient generative LLM inference using phase splitting
Patel P, Choukse E, Zhang C, et al. Splitwise: Efficient generative LLM inference using phase splitting. In: Proceedings of the 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture. 2024:118–132
2024
-
[66]
Fast distributed inference serving for large language models
Wu B, Zhong Y, Zhang Z, et al. Fast distributed inference serving for large language models. arXiv preprint arXiv:2305.05920.2023
2023
-
[67]
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference.arXiv preprint arXiv:2508.19559.2025
Li R, Du R, Chu Z, et al. Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference.arXiv preprint arXiv:2508.19559.2025
2025
-
[68]
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity.arXiv preprint arXiv:2512.03416.2025
Lai R, Liu H, Lu C, et al. TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity.arXiv preprint arXiv:2512.03416.2025
2025
-
[69]
LoongServe: Efficiently serving long-context large language models with elastic sequence parallelism
Wu B, Liu S, Zhong Y, Sun P, Liu X, Jin X. LoongServe: Efficiently serving long-context large language models with elastic sequence parallelism. In: Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles. 2024:640–654. Xu ET AL. 41
2024
-
[70]
dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving
Wu B, Zhu R, Zhang Z, Sun P, Liu X, Jin X. dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving. In: Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation. 2024:911– 927
2024
-
[71]
MegaScale-Infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism
Zhu R, Jiang Z, Jin C, et al. MegaScale-Infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism. In: Proceedings of the ACM SIGCOMM 2025 Conference. 2025:592–608
2025
-
[72]
Punica: Multi-tenant LoRA serving.Proceedings of Machine Learning and Systems.2024;6:1–13
Chen L, Ye Z, Wu Y, Zhuo D, Ceze L, Krishnamurthy A. Punica: Multi-tenant LoRA serving.Proceedings of Machine Learning and Systems.2024;6:1–13
2024
-
[73]
AlpaServe: Statistical multiplexing with model parallelism for deep learning serving
Li Z, Zheng L, Zhong Y, et al. AlpaServe: Statistical multiplexing with model parallelism for deep learning serving. In: Proceedings of the 17th USENIX Symposium on Operating Systems Design and Implementation. 2023:663–679
2023
-
[74]
arXivpreprint arXiv:2503.22562.2025
GoelK,MohanJ,KwatraN,AnupindiRS,RamjeeR.Niyama:BreakingthesilosofLLMinferenceserving. arXivpreprint arXiv:2503.22562.2025
2025
-
[75]
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Dis- aggregated LLM Serving in AI Infrastructure
He Y, Xu M, Wu J, et al. BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Dis- aggregated LLM Serving in AI Infrastructure. Software: Practice and Experience. 2026;56(4):424-444. doi: https://doi.org/10.1002/spe.70054
2026 doi
-
[76]
DynamoLLM: Designing LLM Inference Clusters for Perfor- manceandEnergyEfficiency.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture
Stojkovic J, Zhang C, Goiri I, Torrellas J, Choukse E. DynamoLLM: Designing LLM Inference Clusters for Perfor- manceandEnergyEfficiency.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture. 2025:1348-1362
2025
-
[77]
2024:207–222
PatelP,ChoukseE,ZhangC,etal.CharacterizingpowermanagementopportunitiesforLLMsinthecloud.In:Proceedings ofthe29thACMInternationalConferenceonArchitecturalSupportforProgrammingLanguagesandOperatingSystems, Volume 3. 2024:207–222
2024
-
[78]
In: Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2
StojkovicJ,ZhangC,GoiriI,etal.TAPAS:Thermal-andPower-AwareSchedulingforLLMInferenceinCloudPlatforms. In: Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. Association for Computing Machinery...
2025
-
[79]
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding.arXiv preprint arXiv:2501.12162.2026
Zikun L, Zhuofu C, Remi D, et al. AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding.arXiv preprint arXiv:2501.12162.2026
2026
-
[80]
Fairness in serving large language models
Sheng Y, Cao S, Li D, et al. Fairness in serving large language models. In: Proceedings of the 18th USENIX Conference on Operating Systems Design and Implementation. USENIX Association 2024; USA
2024
-
[81]
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-Based Clusters.IEEE Transactions on Services Computing.2024;17(6):3473-3484
Bai H, Xu M, Ye K, Buyya R, Xu C. DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-Based Clusters.IEEE Transactions on Services Computing.2024;17(6):3473-3484. doi: 10.1109/TSC.2024.3433388
2024 doi
-
[82]
In: Proceedings of the 2023 USENIX Annual Technical Conference
QiuH,MaoW,WangC,etal.AWARE:Automateworkloadautoscalingwithreinforcementlearninginproductioncloud systems. In: Proceedings of the 2023 USENIX Annual Technical Conference. 2023:387–402
2023
-
[83]
2024:929–945
LinC,HanZ,ZhangC,etal.Parrot:EfficientservingofLLM-basedapplicationswithsemanticvariable.In:Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation. 2024:929–945
2024
-
[84]
Sun B, Pinciroli R, Casale G, Smirni E. DeepBAT: Performance and Cost Optimization of Serverless Inference Using Transformers.In:Proceedingsofthe2025IEEEInternationalParallelandDistributedProcessingSymposium.2025:335- 346
2025
-
[85]
In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
MoZ,LiaoJ,XuH,ZhouZ,XuC.Hetis:ServingLLMsinHeterogeneousGPUClusterswithFine-grainedandDynamic Parallelism. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. Association for Computing Machinery 2025; New York, NY, ...
2025
-
[86]
USENIX Association 2024; Santa Clara, CA:135–153
FuY,XueL,HuangY,etal.ServerlessLLM:Low-LatencyServerlessInferenceforLargeLanguageModels.In:Proceed- ings of the 18th USENIX Symposium on Operating Systems Design and Implementation. USENIX Association 2024; Santa Clara, CA:135–153. 42 Xu ET AL
2024
-
[87]
TAPAS: Thermal-and power-aware scheduling for LLM inference in cloud platforms
Stojkovic J, Zhang C, Goiri I, et al. TAPAS: Thermal-and power-aware scheduling for LLM inference in cloud platforms. In: Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. 2025:1266–1281
2025
-
[88]
throttLL’eM: Predictive GPU Throttling for Energy Effi- cientLLMInferenceServing.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture
Kakolyris AK, Masouros D, Vavaroutsos P, Xydis S, Soudris D. throttLL’eM: Predictive GPU Throttling for Energy Effi- cientLLMInferenceServing.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture. 2025:1363-1378
2025
-
[89]
Medusa: Accelerating Serverless LLM Inference with Materialization
Zeng S, Xie M, Gao S, Chen Y, Lu Y. Medusa: Accelerating Serverless LLM Inference with Materialization. In: Pro- ceedingsofthe30thACMInternationalConferenceonArchitecturalSupportforProgrammingLanguagesandOperating Systems, Volume 1. Association for Computing Machinery 2025; Ne...
2025
-
[90]
BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching
Zhang D, Wang H, Liu Y, et al. BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching. In: Proceedings of the 19th USENIX Symposium on Operating Systems Design and Implementation. 2025:275–293
2025
-
[91]
2025:415–430
GimI,MaZ,LeeSs,ZhongL.Pie:Aprogrammableservingsystemforemergingllmapplications.In:Proceedingsofthe ACM SIGOPS 31st Symposium on Operating Systems Principles. 2025:415–430
2025
-
[92]
HanY,KimI,KimJ,MoonGE.TensorCore-AdaptedSparseMatrixMultiplicationforAcceleratingSparseDeepNeural Networks.Electronics.2024;13(20):3981
2024
-
[93]
FPGA-Based Sparse Matrix Multiplication Accelerators: From State-of-the-Art to Future Opportunities
Liu Y, Chen R, Li S, others . FPGA-Based Sparse Matrix Multiplication Accelerators: From State-of-the-Art to Future Opportunities. ACM Transactions on Reconfigurable Technology and Systems.2024;17(4)
2024
-
[94]
A survey on hardware accelerators for large language models.Applied Sciences.2025;15(2):586
Kachris C. A survey on hardware accelerators for large language models.Applied Sciences.2025;15(2):586
2025
-
[95]
arXiv preprint arXiv:2307.08191.2023
LiangZ,ChengJ,YangR,etal.Unleashingthepotentialofllmsforquantumcomputing:Astudyinquantumarchitecture design. arXiv preprint arXiv:2307.08191.2023
2023
-
[96]
Efficient large language models: A survey.arXiv preprint arXiv:2312.03863.2023
Wan Z, Wang X, Liu C, et al. Efficient large language models: A survey.arXiv preprint arXiv:2312.03863.2023
2023
-
[97]
HeY,FangJ,YuFR,LeungVC.Largelanguagemodels(LLMs)inferenceoffloadingandresourceallocationincloud-edge computing: An active inference approach.IEEE Transactions on Mobile Computing.2024;23(12):11253–11264
2024
-
[98]
TangX,LiuF,XuD,etal.LLM-assistedreinforcementlearning:leveraginglightweightlargelanguagemodelcapabilities for efficient task scheduling in multi-cloud environment.IEEE Transactions on Consumer Electronics.2025;71(2):5631– 5644
2025
-
[99]
LLM-based cost-aware task scheduling for cloud computing systems.Journal of Cloud Computing
Pei H, Gu Y, Sun Y, et al. LLM-based cost-aware task scheduling for cloud computing systems.Journal of Cloud Computing. 2025;14(1):81
2025
-
[100]
Green scheduling for LLM workloads with model and data reuse across geo-distributed data centers.Digital Communications and Networks.2025
Liu H, Hu X, Wang R, Hao J, Wu Q, Zhang H. Green scheduling for LLM workloads with model and data reuse across geo-distributed data centers.Digital Communications and Networks.2025
2025
-
[101]
Reliable and Efficient LLM Inference on Resource-Constrained Mobile Devices via Dynamic Scheduling.IEEE Internet of Things Journal.2026
Liang A, Hong C, Li Y, Zhu H. Reliable and Efficient LLM Inference on Resource-Constrained Mobile Devices via Dynamic Scheduling.IEEE Internet of Things Journal.2026
2026
-
[102]
Scalable Resource Provisioning Framework for Fog Computing Using LLM-Guided Q- Learning Approach.Algorithms
Krishnamurthy B, Shiva SG. Scalable Resource Provisioning Framework for Fog Computing Using LLM-Guided Q- Learning Approach.Algorithms. 2025;18(4):230
2025
-
[103]
Distributed training of large language models: A survey.Natural Language Processing Journal
Zeng F, Gan W, Wang Y, Yu PS. Distributed training of large language models: A survey.Natural Language Processing Journal. 2025:100174
2025
-
[104]
Atom: Low-bit quantization for efficient and accurate llm serving.Proceedings of Machine Learning and Systems.2024;6:196–209
Zhao Y, Lin CY, Zhu K, et al. Atom: Low-bit quantization for efficient and accurate llm serving.Proceedings of Machine Learning and Systems.2024;6:196–209
2024
-
[105]
Ladder: Enabling efficient{Low-Precision} deep learning computing through hardware- aware tensor transformation
Wang L, Ma L, Cao S, et al. Ladder: Enabling efficient{Low-Precision} deep learning computing through hardware- aware tensor transformation. In: Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation. 2024:307–323. Xu ET AL. 43
2024
-
[106]
Energy Efficient Software Hardware CoDesign for Machine Learning: From TinyML to Large Language Models.arXiv preprint arXiv:2603.23668.2026
Vahdatpour MS, Zhang Y. Energy Efficient Software Hardware CoDesign for Machine Learning: From TinyML to Large Language Models.arXiv preprint arXiv:2603.23668.2026
2026
-
[107]
arXivpreprintarXiv:2404.14294
ZhouZ,NingX,HongK,etal.Asurveyonefficientinferenceforlargelanguagemodels. arXivpreprintarXiv:2404.14294. 2024
2024
-
[108]
Renewable and Sustainable Energy Reviews.2026;225:116159
JiZ,JiangM.Asystematicreviewofelectricitydemandforlargelanguagemodels:evaluations,challenges,andsolutions. Renewable and Sustainable Energy Reviews.2026;225:116159
2026
-
[109]
In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
FernandezJ,NaC,TiwariV,BiskY,LuccioniS,StrubellE.Energyconsiderationsoflargelanguagemodelinferenceand efficiency optimizations. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025:32556–32569
2025
-
[110]
Energy-Efficient Training and Inference in Large Language Models: Optimizing Computational and Energy Costs.International Journal of Computer Applications.2025;975:8887
Narsepalle KR. Energy-Efficient Training and Inference in Large Language Models: Optimizing Computational and Energy Costs.International Journal of Computer Applications.2025;975:8887
2025
-
[111]
WangG,CaiS,LiW,LyuD,HeG.OFQ-LLM:Outlier-FlexingQuantizationforEfficientLow-BitLargeLanguageModel Acceleration.IEEE Transactions on Circuits and Systems I: Regular Papers.2025
2025
-
[112]
2025:285–309
XuZ,LiF,ChenAC,WangX.ProCut:LLMPromptCompressionviaAttributionEstimation.In:Proceedingsofthe2025 Conference on Empirical Methods in Natural Language Processing: Industry Track. 2025:285–309
2025
-
[113]
Parse trees guided LLM prompt compression.IEEE Transactions on Pattern Analysis and Machine Intelligence.2025
Mao W, Hou C, Zhang T, Lin X, Tang K, Lv H. Parse trees guided LLM prompt compression.IEEE Transactions on Pattern Analysis and Machine Intelligence.2025
2025
-
[114]
Promptware Engineering: Software Engineering for Prompt-Enabled Systems.ACM Transactions on Software Engineering and Methodology.2026
Chen Z, Wang C, Sun W, Liu X, Zhang JM, Liu Y. Promptware Engineering: Software Engineering for Prompt-Enabled Systems.ACM Transactions on Software Engineering and Methodology.2026
2026
-
[115]
Sufficient context: A new lens on retrieval augmented generation systems.arXiv preprint arXiv:2411.06037.2024
Joren H, Zhang J, Ferng CS, Juan DC, Taly A, Rashtchian C. Sufficient context: A new lens on retrieval augmented generation systems.arXiv preprint arXiv:2411.06037.2024
2024
-
[116]
Rankrag: Unifying context ranking with retrieval-augmented generation in llms.Proceedings of the Advances in Neural Information Processing Systems.2024;37:121156–121184
Yu Y, Ping W, Liu Z, et al. Rankrag: Unifying context ranking with retrieval-augmented generation in llms.Proceedings of the Advances in Neural Information Processing Systems.2024;37:121156–121184
2024
-
[117]
In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track
LiZ,LiC,ZhangM,MeiQ,BenderskyM.Retrievalaugmentedgenerationorlong-contextllms?acomprehensivestudy and hybrid approach. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track. 2024:881–893
2024
-
[118]
arXivpreprintarXiv:2502.12110
XuW,LiangZ,MeiK,GaoH,TanJ,ZhangY.A-mem:Agenticmemoryforllmagents. arXivpreprintarXiv:2502.12110. 2025
2025
-
[119]
Aiopslab: A holistic framework to evaluate ai agents for enabling autonomous clouds
Chen Y, Shetty M, Somashekar G, et al. Aiopslab: A holistic framework to evaluate ai agents for enabling autonomous clouds. Proceedings of Machine Learning and Systems.2025;7
2025
-
[120]
Lads: Leveraging llms for ai-driven devops.arXiv preprint arXiv:2502.20825
Khan AF, Khan AA, Mohamed A, et al. Lads: Leveraging llms for ai-driven devops.arXiv preprint arXiv:2502.20825. 2025
2025
-
[121]
The Multi-Agent Fault Localization System Based on Monte Carlo Tree Search Approach.arXiv preprint arXiv:2507.22800.2025
Ren R. The Multi-Agent Fault Localization System Based on Monte Carlo Tree Search Approach.arXiv preprint arXiv:2507.22800.2025
2025
-
[122]
Parallax: Efficient llm inference service over decentralized environment.arXiv preprint arXiv:2509.26182.2025
Tong C, Jiang Y, Chen G, et al. Parallax: Efficient llm inference service over decentralized environment.arXiv preprint arXiv:2509.26182.2025
2025
-
[123]
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput.arXiv preprint arXiv:2511.11733.2025
Song J, Chen W, Song X, et al. Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput.arXiv preprint arXiv:2511.11733.2025
2025
-
[124]
IEEEComputerArchitectureLetters
WangX,SunX,LiW,etal.Low-LatencyPIMAcceleratorforEdgeLLMInference. IEEEComputerArchitectureLetters. 2025. 44 Xu ET AL
2025
-
[125]
Decentralized llm inference over edge networks with energy harvesting
Khoshsirat A, Perin G, Rossi M. Decentralized llm inference over edge networks with energy harvesting. In: Proceedings of the IEEE Global Communications Conference. 2024:3703–3708
2024
-
[126]
Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic.arXiv preprint arXiv:2601.21972.2026
Liu S, Chen T, Amiri R, Amato C. Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic.arXiv preprint arXiv:2601.21972.2026
2026 arXiv
-
[127]
arXivpreprintarXiv:2106.09685
HuEJ,ShenY,WallisP,etal.LoRA:Low-RankAdaptationofLargeLanguageModels. arXivpreprintarXiv:2106.09685. 2021
2021 arXiv
-
[128]
Safe-FedLLM: Delving into the Safety of Federated Large Language Models
Tao M, Tian Y, Tu W, Yang Y, Yang X, Tang X. Safe-FedLLM: Delving into the Safety of Federated Large Language Models. arXiv preprint arXiv:2601.07177.2026
2026 arXiv
-
[129]
Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference.arXiv preprint arXiv:2512.16317.2025
Tian A, Ding A, Chen F, Wu A, Chan A, Zhang B. Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference.arXiv preprint arXiv:2512.16317.2025
2025
-
[130]
Findingsofthe Association for Computational Linguistics: EMNLP.2025:12841–12854
MoonH,SeoJ,KooS,etal.LimaCost:DataValuationforInstructionTuningofLargeLanguageModels. Findingsofthe Association for Computational Linguistics: EMNLP.2025:12841–12854
2025
-
[131]
Fairshare Data Pricing via Data Valuation for Large Language Models.arXiv preprint arXiv:2502.00198.2025
Zhang L, Jiao C, Li B, Xiong C. Fairshare Data Pricing via Data Valuation for Large Language Models.arXiv preprint arXiv:2502.00198.2025
2025
-
[132]
Incentivizing permissionless distributed learning of llms
Lidin J, Sarfi A, Pappas E, Dare S, Belilovsky E, Steeves J. Incentivizing permissionless distributed learning of llms. In: Proceedings of the 2025 7th International Conference on Distributed Artificial Intelligence. 2025:12–18
2025
-
[133]
Incentive mechanism of foundation model enabled cross-silo federated learning
Zhang N, Xu X, Liu X, Wu J, Tang H. Incentive mechanism of foundation model enabled cross-silo federated learning. Scientific Reports.2025;15(1):24181
2025
-
[134]
WinFLoRA: Incentivizing Client-Adaptive Aggregation in Federated LoRA under Privacy Heterogeneity.arXiv preprint arXiv:2602.01126.2026
Kou M, Xia X, Wang Z, et al. WinFLoRA: Incentivizing Client-Adaptive Aggregation in Federated LoRA under Privacy Heterogeneity.arXiv preprint arXiv:2602.01126.2026
2026
-
[135]
ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction.arXiv preprint arXiv:2602.12247.2026
Ferguson N, Pennington J, Beghian N, et al. ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction.arXiv preprint arXiv:2602.12247.2026
2026
-
[136]
2025:2749– 2763
ShrimalA,JainA,ChowdhuryS,YenigallaP.PARSE:LLMdrivenschemaoptimizationforreliableentityextraction.In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track. 2025:2749– 2763
2025
-
[137]
DELM: a Python toolkit for Data Extraction with Language Models
Fithian E, Skobelev K. DELM: a Python toolkit for Data Extraction with Language Models. arXiv preprint arXiv:2509.20617.2025
2025
-
[138]
LLMStructBench: Benchmarking Large Language Model Structured Data Extraction.arXiv preprint arXiv:2602.14743.2026
Tenckhoff S, Koddenbrock M, Rodner E. LLMStructBench: Benchmarking Large Language Model Structured Data Extraction.arXiv preprint arXiv:2602.14743.2026
2026
-
[139]
Structure-Aware Decoding Mechanisms for Complex Entity Extraction with Large-Scale Language Models.arXiv preprint arXiv:2512.13980.2025
Qiu Z, Wu D, Liu F, Wang Y. Structure-Aware Decoding Mechanisms for Complex Entity Extraction with Large-Scale Language Models.arXiv preprint arXiv:2512.13980.2025
2025
-
[140]
Manalyzer:End-to-end Automated Meta-analysis withMulti-agent System.arXiv preprint arXiv:2505.20310.2025
Xu W, ZhangW, Ling F,et al. Manalyzer:End-to-end Automated Meta-analysis withMulti-agent System.arXiv preprint arXiv:2505.20310.2025
2025
-
[141]
Reducing hallucinations of large language models via hierarchical semantic piece.Complex & Intelligent Systems.2025;11(5):231
Liu Y, Yang Q, Tang J, et al. Reducing hallucinations of large language models via hierarchical semantic piece.Complex & Intelligent Systems.2025;11(5):231
2025
-
[142]
In: Proceedings of the 10th Workshop on Noisy and User-generated Text
BirkmoseR,ReeceNM,NorvinEH,BjervaJ,ZhangM.On-deviceLLMsforhomeassistant:Dualroleinintentdetection and response generation. In: Proceedings of the 10th Workshop on Noisy and User-generated Text. 2025:57–67
2025
-
[143]
Vega: LLM-driven intelligent chatbot platform for internet of things control and development.Sensors.2025;25(12):3809
Al-Safi H, Ibrahim H, Steenson P. Vega: LLM-driven intelligent chatbot platform for internet of things control and development.Sensors.2025;25(12):3809
2025
-
[144]
IEEE Open Journal of the Computer Society.2025
LinYW,LinYB,WuYF,ShenPH.Voicetalk:Ano-codeapproachforcreatingvoice-controlledsmarthomeapplications. IEEE Open Journal of the Computer Society.2025. Xu ET AL. 45
2025
-
[145]
Introducing LEAF: LLM Edge Assessment Framework for Generative AI on the Edge
Abdulkadhim M, Repas SR. Introducing LEAF: LLM Edge Assessment Framework for Generative AI on the Edge. Machine Learning and Knowledge Extraction.2026;8(2):48
2026
-
[146]
arXivpreprintarXiv:2409.00088
XuJ,LiZ,ChenW,etal.On-devicelanguagemodels:Acomprehensivereview. arXivpreprintarXiv:2409.00088. 2024
2024
-
[147]
LLM-Driven Auto Configuration for Transient IoT Device Collaboration
Shastri H, Hanafy W, Wu L, Irwin D, Srivastava M, Shenoy P. LLM-Driven Auto Configuration for Transient IoT Device Collaboration. In: Proceedings of the 10th ACM/IEEE Symposium on Edge Computing. 2025:1–17
2025
-
[148]
Instant Personalized Large Language Model Adaptation via Hypernetwork.arXiv preprint arXiv:2510.16282.2025
Tan Z, Zhang Z, Wen H, et al. Instant Personalized Large Language Model Adaptation via Hypernetwork.arXiv preprint arXiv:2510.16282.2025
2025
-
[149]
Embedding-to-Prefix: Parameter-efficient personalization for pre-trained large language models.arXiv preprint arXiv:2505.17051.2025
Huber B, Fazelnia G, Damianou A, et al. Embedding-to-Prefix: Parameter-efficient personalization for pre-trained large language models.arXiv preprint arXiv:2505.17051.2025
2025
-
[150]
Ensuring fair llm serving amid diverse applications.arXiv preprint arXiv:2411.15997
Khan RIS, Jain K, Shen H, et al. Ensuring fair llm serving amid diverse applications.arXiv preprint arXiv:2411.15997. 2024
2024
-
[151]
T-POP: Test-Time Personalization with Online Preference Feedback.arXiv preprint arXiv:2509.24696.2025
Qu Z, Zhang M, Kong M, et al. T-POP: Test-Time Personalization with Online Preference Feedback.arXiv preprint arXiv:2509.24696.2025
2025
Reviewed May 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.