Pith. sign in

Paper Citation Record · LEDGER

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda

As of 5 August 2026, this Paper Citation Record lists 100 of 151 outbound references and 2 inbound Pith citation observations for arXiv:2604.17227.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.17227 v1

Coverage vector

measured 100 of 151 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T06:27:23.580445Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T21:14:49.980623Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 151 outbound references displayed

  • verified exact18
  • verified fuzzy69
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c48d021b-c541-4a70-97a1-acbdd17d45ab · outbound

This paper cites In: Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.366677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:548ab63f5942556d97a49d3bbf27e793248754b669507e083b28bc2032cf4212

Observation cb4b2a59-1731-492c-8188-b74ba53c7337 · outbound

This paper cites Cloud container technologies: a state-of-the-art review.IEEE Transactions on Cloud Computing.2017;7(3):677–692.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Cloud container technologies: a state-of-the-art review.IEEE Transactions on Cloud Computing.2017;7(3):677–692

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.338435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:e1fef1cd20e2fd6cd2036ab6b481141a16e4f493b5758af3bf3cbca64fc85cbe

Observation 87d57ad5-9bb0-4aa3-b1f4-e7e99f4f2c17 · outbound

This paper cites Attention is all you need.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Attention is all you need

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.326453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:730aabc8a4311300a5667180b7ad325ef535a0c3e422ba80cf49b5f0bb5d03aa

Observation cfbd3ce2-6896-4075-b0a1-30ca5b82bf50 · outbound

This paper cites Scaling Laws for Neural Language Models.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Scaling Laws for Neural Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-10T06:31:30.880784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:2a382805ba1a19daff5686d9ccc484dbfabbf6d7c1576cee0e6915034a142b3b

Observation 921ccf4f-c350-484b-8d0b-9fc327f4461e · outbound

This paper cites Deep learning.Nature.2015;521(7553):436–444.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Deep learning.Nature.2015;521(7553):436–444

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.340832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:cd2fad6033b2c0993184dcae5a6e454c8ccea75d240042b6dee00625f1898b38

Observation 42cc0610-3505-46a8-94f3-434d6cea221d · outbound

This paper cites an unresolved cited work.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-21T16:34:17.362597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:e21c7925b2e28b2f93aa44019fe71c905927bc39f893cfffb2e6dd32192af23d

Observation a4c780b8-a6aa-4ea1-9cde-96d2911ed3d5 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.; 2020.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.; 2020

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.285487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:81641657ecebfd8e41b6036d019719476062c49dfd680e27a8c997154dcc29c8

Observation e4414f9a-de00-4c4a-8ab8-d5105e610d3c · outbound

This paper cites Curran Associates, Inc.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Curran Associates, Inc

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.275458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:63d22906217eb5437f544cc7d0312156c8f9f06c3b93129e346df4c92914557b

Observation edd5c27a-92dd-4792-8e0c-2ec5752cc93f · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.316645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:aa13a1cf7f148baa5c8dbd1846f8d2d90a1bcae5dca26579d35867447dc1d3c0

Observation 7d1ee9dd-94c6-4734-a7a8-0f3a543aff82 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.360767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:285dcc1c2307700e0c635d2e481a496448554d9f75f7959962c1548c52ad1729

Observation abbd93f5-bdde-4ef1-86ff-452886a8afc1 · outbound

This paper cites Towards end-to-end optimization of llm-based applications with ayo.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Towards end-to-end optimization of llm-based applications with ayo

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.356085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:b31853b63b804d83a837d186f8dfeb06a0367bf720e4f448484c39038c239be5

Observation 191292db-c5f5-46bb-a001-ca180fef44a1 · outbound

This paper cites In-datacenter performance analysis of a tensor processing unit.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In-datacenter performance analysis of a tensor processing unit

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.322518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:e6db83ea1a2934df7f7854abf39077536aa4a8df87293f8696280c47dbbababd

Observation 71e6585a-b4c3-484f-8bc6-464319c0caf5 · outbound

This paper cites Alpa: Automating Inter-and Intra-Operator Parallelism for Distributed Deep Learning.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Alpa: Automating Inter-and Intra-Operator Parallelism for Distributed Deep Learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.259827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:ab3772d0f4405db65626242440a5f9772c9a8e081964d228d665b2d97fadaf3f

Observation 76f87055-c3fd-41ea-8f8c-b79c8a5a175f · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Efficient memory management for large language model serving with pagedattention

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.328335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:9d1050c4cf20ac1a57e49d1dfa7ada30399891b39dcd3f0e83aeef09494d838a

Observation 00b1b090-037b-4b7b-a68e-63993c042d0b · outbound

This paper cites Open issues in scheduling microservices in the cloud.IEEE Cloud Computing.2016;3(5):81–88.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Open issues in scheduling microservices in the cloud.IEEE Cloud Computing.2016;3(5):81–88

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.330406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:d60107110035bdac5d9e9be1f9e25c015c9bebc80b3f2ea4433972daf5b81a53

Observation c6832a14-856f-4a83-8538-6dd4db7e69ee · outbound

This paper cites an unresolved cited work.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-21T16:34:17.344166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:5340d331f69440d6f45351afcf9ddd04c358a49d5024616dc5e669bc4faf7080

Observation 04b02cac-0fc8-495f-88f3-a0b7207abf99 · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:35:22.390830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:10b52d33b4602029d4bc03b46d693145451ccb672d5311fd02d5d6ce1e4ac398

Observation 31c66a56-e34f-4b51-9dca-7a324c2a4559 · outbound

This paper cites In: Proceedings of the 16th USENIX symposium on operating systems design and implementation.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the 16th USENIX symposium on operating systems design and implementation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.312290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:f8823dd6a0d2d35b39ad0ceb0681d66e27fe8949f759706c56f1ce8321dc536c

Observation 7592b4d9-ee42-4e58-99f8-b94283ad7729 · outbound

This paper cites Evaluation and benchmarking of llm agents: A survey.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Evaluation and benchmarking of llm agents: A survey

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.380668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:9054a9437295639469de335272f5ef5c2efec73f4d164c7a69c23835aa194c90

Observation f62d3047-60a2-4612-a197-c53594eb37af · outbound

This paper cites Towards High-Goodput LLM Serving with Prefill-decode Multiplexing.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Towards High-Goodput LLM Serving with Prefill-decode Multiplexing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.300201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:0d1a8cf499954929dd7d1c9276a77457ba6318ce9525967d36fbdbd8bfaaba4a

Observation d0dcb709-9ef9-4b4f-8ceb-a3b11b8cbe72 · outbound

This paper cites Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.302362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:da8d41f063d962df9897820937cb828a82b0862901431cfcbc91343acac70f6f

Observation 7de80109-f300-4869-8291-f4851a0591e6 · outbound

This paper cites Process modeling in web applications.ACM Transactions on Software Engineering and Methodology.2006;15(4):360–409.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Process modeling in web applications.ACM Transactions on Software Engineering and Methodology.2006;15(4):360–409

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.310408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:590e3b76a20769ee089ce6f328ec07667ef9a64d9dc82852f06937c7752829f9

Observation 674efc27-9095-43ba-8783-20e44dd606c0 · outbound

This paper cites Challenges in deployment and configuration management in cyber physical system.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Challenges in deployment and configuration management in cyber physical system

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.296029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:0293c11a1106b79aa98aceafecbcda22c321d7403844176e3584eb448ed8e806

Observation 3e8760e8-7ba6-4e9a-989d-1cdc00472b7d · outbound

This paper cites Parallel processing systems for big data: a survey.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Parallel processing systems for big data: a survey

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.384760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:37d65949d5dc33cad1eb845493f77cda2193b81bd1286c5a1d0b7b4d574d01bd

Observation 87b97a7f-290d-432b-9472-6b79a89b36a8 · outbound

This paper cites an unresolved cited work.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-21T16:34:17.308428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:d9f5ee834b55f1fa6d249df05cfddff4a7bbf919affa31ed649bd68c4d589fee

Observation 3266407b-5c1d-4cd9-a523-4b9831d03494 · outbound

This paper cites Deep learning workload scheduling in gpu datacenters: A survey.ACM Computing Surveys.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Deep learning workload scheduling in gpu datacenters: A survey.ACM Computing Surveys

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.304441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:4135ea855264acef2638ca9ffce6bdf054b92344f5580753cd60d9e5c0566531

Observation c5b06068-db71-42aa-a74e-63855aa317e8 · outbound

This paper cites StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications.ACM Trans.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications.ACM Trans

Reference 27

Resolution
verified exact
doi, observed 2026-05-10T06:31:30.372822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:c53feb5c5561ebd05562a7d92367f294206bd2021938ca477b52e347d6469cd3

Observation aee1276b-26da-4cb1-9be2-40515646f7fe · outbound

This paper cites LLM Inference Scheduling: A Survey of Techniques, Frameworks, and Trade-offs.Authorea Preprints.2025.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda LLM Inference Scheduling: A Survey of Techniques, Frameworks, and Trade-offs.Authorea Preprints.2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.291688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:49aa801078aa1596e176936f45d183e1063a54a7d31c6efd1c72756c2c6843af

Observation 9ed6541f-855a-4c8a-ad05-0de53dbfd62c · outbound

This paper cites Cloud Native System for LLM Inference Serving.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Cloud Native System for LLM Inference Serving

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.777592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:98a5b7e7b03285ba36403ae0729b0b4bcde6932c164cc29463da33897f6c6168

Observation c6219f06-6e3a-45c4-bb18-ba43623f8455 · outbound

This paper cites {NanoFlow}:Towardsoptimallargelanguagemodelservingthroughput.In:Proceedingsof the 19th USENIX Symposium on Operating Systems Design and Implementation.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda {NanoFlow}:Towardsoptimallargelanguagemodelservingthroughput.In:Proceedingsof the 19th USENIX Symposium on Operating Systems Design and Implementation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.318701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:fea34d93622e1a7701439d53682ba39b38f1efe37d1308d06bfd7e30ab4fc206

Observation 8a1a0b92-22c1-45d6-9c25-9a4134562f43 · outbound

This paper cites Flashinfer: Efficient and customizable attention engine for llm inference serving.Proceedings of Machine Learning and Systems.2025;7.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Flashinfer: Efficient and customizable attention engine for llm inference serving.Proceedings of Machine Learning and Systems.2025;7

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.279709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:f23a146e36563dd537d76b5b8aa65a3d6f8bfde2ee75b89f874d9d3814c109dd

Observation b70ee6af-dddb-4229-9208-ca4bc3086d6c · outbound

This paper cites 2025:446–461.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2025:446–461

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.283410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:a848230be022d3bad728c09c5f30522be907c20af0f9322f62ae36c931b97489

Observation a32948a6-e47b-4542-8725-c4ae40f46be8 · outbound

This paper cites throttll’em: Predictive gpu throttling for energy effi- cient llm inference serving.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda throttll’em: Predictive gpu throttling for energy effi- cient llm inference serving

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.350080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:a810738d6b49f862367e824a7f89589075735069e90206398cc91bb7db518b15

Observation 136e5946-975c-43d6-948e-22a5ef67db2d · outbound

This paper cites Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.807812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:e14e8b8739013a529bd2004d5f7b531160150ba236d42766840d72efb6d3cf27

Observation 3f61e7be-90b1-4fc5-a2e0-3b76f985d0dc · outbound

This paper cites Extracting training data from large language models.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Extracting training data from large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.324489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:583ec0f6aad36934ec2995c426a1a752494722bcfb8e336b7e3c42149ffcd694

Observation 52b95004-ad1a-4f20-b241-b1c70af874c5 · outbound

This paper cites Quantifying memorization across neural language models.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Quantifying memorization across neural language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.346215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:1d055b5ab2f68b99a714e0139981f136024b96e351ed28eda9972b7c6812e28d

Observation 0d4861b5-39cc-417a-9dce-1fd76660fbb7 · outbound

This paper cites Deep learning with differential privacy.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Deep learning with differential privacy

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.348168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:1922c6179bb5ad9276501a510c192a718c763fa4543af0eca89abca5fba85517

Observation 79988a33-79ae-4a19-802b-1b733f770298 · outbound

This paper cites Communication-efficient learning of deep networks from decentralized data.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Communication-efficient learning of deep networks from decentralized data

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.253289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:bb200b73199b9759f4fc761e1278ce08e51ce87faf1ca7a6612a8197fb2e2abc

Observation 9fb2a544-f8fc-4640-8e94-99a86ad8a14a · outbound

This paper cites Oblivious {Multi-Party} machine learning on trusted processors.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Oblivious {Multi-Party} machine learning on trusted processors

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.266586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:0c9a8a9d293e5e9c60b1327abb8f5d8f97871db4501810f0e234232cb447ca47

Observation 84cbbcb9-0f0e-4a32-858f-d4609f09995c · outbound

This paper cites Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.289684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:7cac3363b218d777bf4dc7794fa37875c99e84355b529937d1675d16650ca956

Observation b7f52107-b027-4e1e-90c6-7726ee0fbdc8 · outbound

This paper cites Stealing machine learning models via prediction APIs.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Stealing machine learning models via prediction APIs

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.334407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:a35684d90568f62cb2d5e25db3306520f4da93801711013a639b81befb916fb2

Observation 3897ff2d-81d3-4c5b-a502-d1e187ecc349 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-10T06:31:30.831485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:0f42a434ac501426f011e711803fe28fb7bc45d1fb2efd9ccab20f31eb107cca

Observation 867c3056-3558-43b0-9978-e4746d394dbf · outbound

This paper cites ACMComputingSurveys.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda ACMComputingSurveys

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.277744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:191830b557e95f7ea2c1b9bdab8db5be313bb588bc12683d6a5960e2fce14005

Observation 483a2e16-488a-47b8-a178-8cebbb345a8d · outbound

This paper cites IEEEInternet of Things Journal.2022;9(11):8364–8386.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda IEEEInternet of Things Journal.2022;9(11):8364–8386

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.261925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:c76acb81d384294af68865b48f7bc3e8915f0b6dfa761ec054cd44fb95b8731a

Observation d42d3bd2-9270-46c3-acf1-06fc45c39d89 · outbound

This paper cites Spatial big data architecture: from data warehouses and data lakes to the Lakehouse.Journal of Parallel and Distributed Computing.2023;176:70–79.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Spatial big data architecture: from data warehouses and data lakes to the Lakehouse.Journal of Parallel and Distributed Computing.2023;176:70–79

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.306573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:3aa2e1f5bc573818c5944c78314575b70b9693babc38cbf7a4c59f95e93769c7

Observation feaab007-1c53-4b10-aeed-098ea0d3f8c6 · outbound

This paper cites IEEETransactionsonKnowledgeand Data Engineering.2023;35(12):12571–12590.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda IEEETransactionsonKnowledgeand Data Engineering.2023;35(12):12571–12590

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.861097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:a24d0015a586046c99ac5c4d26200746cf843f3c74e9ee925ac25040810e2469

Observation 90ed30bf-8d6e-483d-9703-f031ef266de1 · outbound

This paper cites Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:31:30.820771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:e0d4e56bd39e28d93808a1fb09930733ca07dd4ace5955ee8f1dd8576b341ea8

Observation 56a3e7f3-5af7-45b2-890e-e365d97dc068 · outbound

This paper cites Computational cxl-memory solution for accelerating memory-intensive applications.IEEE Computer Architecture Letters.2022;22(1):5–8.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Computational cxl-memory solution for accelerating memory-intensive applications.IEEE Computer Architecture Letters.2022;22(1):5–8

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.243986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:37b6654dbbefe01df698807d97ec6eb2f4c838b80089422beb7606bb231b7431

Observation 39e63fcf-045e-4efd-83d7-d7d1c1292230 · outbound

This paper cites an unresolved cited work.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-21T16:34:17.298169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:7c8b089cc76f632e19af84beda5fe643759e5689a714974e0597a435faa7daff

Observation a28dbc49-7171-4db8-b8d2-9f37d0ef3ead · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-10T06:31:30.817937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:eaa3d91480e4a20a68d3f09b8413c20de6e3836065553b82187ef965665d4813

Observation 139a995e-2eeb-4bd5-859c-619552c03f31 · outbound

This paper cites New Trends in High-D Vector Similarity Search: AI-driven, Progressive, and Distributed.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda New Trends in High-D Vector Similarity Search: AI-driven, Progressive, and Distributed

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.368750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:685917d157ee4aba7afec6cca24e3e00564745683b5dc3749b443d764bd56b29

Observation 6f894ca1-b279-46c6-a7cd-50a60b4ee277 · outbound

This paper cites Unleash llms potential for sequential recommendation by coordinating dual dynamic index mechanism.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unleash llms potential for sequential recommendation by coordinating dual dynamic index mechanism

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.358378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:d649fa892b534e6200c2ac5e72b237b38167d9e6957d2a383a9b82e4248ebd10

Observation 5a4194cc-fc9f-4b09-be6f-094e7a7c59d4 · outbound

This paper cites an unresolved cited work.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-21T16:34:17.268661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:53c5cf18a5ab2d47186beb23201129ef5d54e292b115267cf335946f26a49cc3

Observation 1ddf85f7-28e1-48f3-baad-46b027d7885f · outbound

This paper cites A Generative Caching System for Large Language Models.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda A Generative Caching System for Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.823360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:49e75efbb76201669627c7ed8b959e8887e6d702caf1e1e398e89e04809560b6

Observation e54daba5-3418-4a40-b476-013c1f153ae3 · outbound

This paper cites Tracing the Data Trail: A Survey of Data Provenance, Transparency and Traceability in LLMs.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Tracing the Data Trail: A Survey of Data Provenance, Transparency and Traceability in LLMs

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.834032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:6b1f8ad3589ad9af77a91f19ac666f94da85f8ed77c49ac64d7c2d4d525ea803

Observation 95e1ea6d-628d-4f72-81cb-e5e5314f365b · outbound

This paper cites In: Proceedings of the 41st IEEE/ACM International Conference on Computer-Aided Design.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the 41st IEEE/ACM International Conference on Computer-Aided Design

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.314502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:dc26dee26e06cbf7a1c2f2012485607023bd4101c5648146e85120b5d4dfd3bf

Observation 5a6e6b58-1395-41da-809e-8aa2ee6b3b5a · outbound

This paper cites 2023:189–201.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2023:189–201

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.950777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:0dd7ef9dedc8d33c13b030ed920b9e8f8db09de5fdac2669d6e3cad677d72c29

Observation cda81a75-03f0-4941-9fff-fc5a3bcba3a8 · outbound

This paper cites Packt Publishing Ltd.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Packt Publishing Ltd

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.844474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:63c663bdb2655f55769576353bea2e7d19acd166e21f59d75452d4a1d48f0370

Observation 5acdf7b7-5833-4142-93cd-19bb6cea4ad3 · outbound

This paper cites Llm-pilot: Characterize and optimize performance of your llm inference services.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Llm-pilot: Characterize and optimize performance of your llm inference services

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.364636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:8e7f017dc09f0983dfe162c6d0fdec842a5706e9bc8ab0b11c69423898837efc

Observation 4c10931d-66ce-424e-9566-8d6835130bad · outbound

This paper cites How hungry is ai? benchmarking energy, water, and carbon footprint of llm inference.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda How hungry is ai? benchmarking energy, water, and carbon footprint of llm inference

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.849607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:da9c6f7da466878ca3a5989c3652e1affc405bfcf8500d563390d2808bdd4cf0

Observation a560c185-448e-4cc3-bbb9-8a339d009fe9 · outbound

This paper cites A survey on federated fine-tuning of large language models.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda A survey on federated fine-tuning of large language models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.873235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:42e1bb5af03bd6dfd943f4bb287785514e35af2dfc5255165d85d56c72ca485a

Observation d9bde166-b12b-4be7-b4bb-c6e04bb594d3 · outbound

This paper cites IEEE Transactions on Computers.2026;75(4):1636-1649.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda IEEE Transactions on Computers.2026;75(4):1636-1649

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.363657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:9130c7245baf53b160dd4734a84d79f4fb34b557c3505a37939ceb8c90fb9c9e

Observation 470a91df-ce65-4816-aec1-95220eeb9386 · outbound

This paper cites 2026 , issn =.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2026 , issn =

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:31:30.366337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:799ce030a69e18dc6d3b2dee44bc74e7872610c04a9ac491149bd1c649e179bb

Observation 4cdca6a3-738d-4eb3-b33a-e108ece95454 · outbound

This paper cites ShuffleInfer: Disaggregate LLM inference for mixed downstream workloads.ACM Transactions on Architecture and Code Optimization.2025;22(2):1–24.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda ShuffleInfer: Disaggregate LLM inference for mixed downstream workloads.ACM Transactions on Architecture and Code Optimization.2025;22(2):1–24

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:35:22.378027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:a4db249f1334ee9fbb1c81573b4bc21034741faa5a332c68f69c361b9dfbfd01

Observation 4b8d8b10-2c56-495d-b1c0-603b5941ac09 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Splitwise: Efficient generative LLM inference using phase splitting

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.374709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:de4ecb163e41cb1365f73600fa49b0b1b9052ea3ce59cc7de695d1670a1c8de7

Observation c02a5c08-c33e-4da5-b4cf-530662a421fc · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Fast Distributed Inference Serving for Large Language Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:23:46.844514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:d2c2d9395b1e2798eafc9213ab4edbfb3adeb4026628233739a09a1a00856ce3

Observation cf870d12-d482-4314-8454-4d57614fcc5a · outbound

This paper cites Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.854890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:18d20315b9c4d00a436124c39a9c58906f36cab8216566cdd121298ec22c431c

Observation 1e2d305a-eb41-4beb-b51b-bc43a6fd9001 · outbound

This paper cites TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity.arXiv preprint arXiv:2512.03416.2025.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity.arXiv preprint arXiv:2512.03416.2025

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.847190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:b6ac1bb3109542010b9d83f44b3a6e104b847495675a75d987d55e8656aaed2b

Observation 0ebbb320-0c7b-4aa0-9864-c55af95745a8 · outbound

This paper cites LoongServe: Efficiently serving long-context large language models with elastic sequence parallelism.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda LoongServe: Efficiently serving long-context large language models with elastic sequence parallelism

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.958636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:717d921e045a58b58b6e99c66b3f0480dee2ff6e26368285989de3caecc3aaa8

Observation 6c720b1a-abd1-4dfe-bb2a-3d0af8021e15 · outbound

This paper cites dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.956536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:257fb393506e84125de82db61aef1de98a8e77c0b8212608cba52097f539bb41

Observation 07cb87e2-5ea5-4db6-9d49-76b258da44c6 · outbound

This paper cites MegaScale-Infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda MegaScale-Infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.831402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:a49f76c73b5f385673edb7e0e0b88e924ed459de4963597a58b2cf017233ad56

Observation e6f1961f-65db-4611-baea-e25520c1d722 · outbound

This paper cites Punica: Multi-tenant LoRA serving.Proceedings of Machine Learning and Systems.2024;6:1–13.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Punica: Multi-tenant LoRA serving.Proceedings of Machine Learning and Systems.2024;6:1–13

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.970694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:add711a55435802cd9bd0e45b9e5a1662e2365713101fa486b64c2ea3777b892

Observation 252ce8d0-2450-4d20-ab5a-a1d0bc76e187 · outbound

This paper cites AlpaServe: Statistical multiplexing with model parallelism for deep learning serving.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda AlpaServe: Statistical multiplexing with model parallelism for deep learning serving

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.823625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:d48051fc0afde646b88b3bd5d668d8fbaf8eeefc0809281fae9589735b6b88b2

Observation 89e93491-6fd0-4de9-9676-507bdd68670d · outbound

This paper cites Niyama : Breaking the Silos of LLM Inference Serving.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Niyama : Breaking the Silos of LLM Inference Serving

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.826535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:b1190ad228563ecf7df9de8ee148d9013c2effcdd3986d0d3e998e0d6665b526

Observation b768dfc0-cc88-4817-ac1e-a89250b66ff3 · outbound

This paper cites BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Dis- aggregated LLM Serving in AI Infrastructure.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Dis- aggregated LLM Serving in AI Infrastructure

Reference 75

Resolution
verified exact
doi, observed 2026-05-10T06:31:30.371034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:786e7685ccbb32b2e5e734588fcc47e4594e8ff503944a9bae86a9a71f051c9a

Observation d9e69d42-00ab-4fce-92f1-1ebd1daf33cd · outbound

This paper cites DynamoLLM: Designing LLM Inference Clusters for Perfor- manceandEnergyEfficiency.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda DynamoLLM: Designing LLM Inference Clusters for Perfor- manceandEnergyEfficiency.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.372914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:794b7296178a797a94274b503901d88c0b0856beab6d1719f83eb0c54af17d32

Observation b769b3bd-5ab1-4d73-b565-4ea4b0e547f8 · outbound

This paper cites 2024:207–222.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2024:207–222

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.833945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:2bc55b36815c8b4589a8029dfc0d2467f514c6b01347ed5315696d1a1bc0c142

Observation 34368300-537d-4d58-9c1d-09b685e1df89 · outbound

This paper cites In: Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.854056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:f6dbb594862be686c4b5f07c0a28c1168a4be5f376fb46f47cba7d52eb12ce7c

Observation 7e021acf-ec94-43f0-ad98-bddc56040aaa · outbound

This paper cites AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.784210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:b41cda573a12d89988fa2da0772c1e7279b8aa19aa67505e6461f98f74c47432

Observation 48d79405-9708-478b-8961-ebd6a4bdf4a2 · outbound

This paper cites Fairness in serving large language models.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Fairness in serving large language models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.889999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:6367d8cbdce675f49b78f958e4d314c9a25d608d4e0678382a6906b5cbf6b3ce

Observation 4db4c170-dc7f-4925-9d9c-198787d691ab · outbound

This paper cites DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-Based Clusters.IEEE Transactions on Services Computing.2024;17(6):3473-3484.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-Based Clusters.IEEE Transactions on Services Computing.2024;17(6):3473-3484

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:31:30.369104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:8be2ae8a21cd8b3b460e85f0ce11bf3b7e999a008eb193bea43444eb2808c4a3

Observation 4f59ea1a-7706-4412-8db4-9cf016b23707 · outbound

This paper cites In: Proceedings of the 2023 USENIX Annual Technical Conference.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the 2023 USENIX Annual Technical Conference

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.980492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:260d9ee4f372032e868b922195d6b3f436a666a969bbccc0943e217b7abeaf22

Observation 0855d5e9-181c-4cf3-b976-b5aa139b186d · outbound

This paper cites 2024:929–945.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2024:929–945

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:35:22.386322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:9f229b9b3f5f5e04f18ede539d351df4c6efe1dde2e32859d7ef5e2f8f146f73

Observation f6b3a9a2-364a-4b87-a085-8f8509fef4a1 · outbound

This paper cites an unresolved cited work.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-05-21T16:35:22.388478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:df0699c5ed8f3c843a01cd2997ced50eaa41aa52744b1a845f657d2cc8d0616d

Observation 80a2d610-850d-461a-b19e-f84086bd8d0a · outbound

This paper cites In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.994671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:fb94436320641e63c85a24f1177ddceeebdf9420a34e589653395b868069c80f

Observation fa266403-5627-4fff-a767-dca513225439 · outbound

This paper cites USENIX Association 2024; Santa Clara, CA:135–153.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda USENIX Association 2024; Santa Clara, CA:135–153

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.997301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:1aac30806eeaae10f3ce45446825b64722f012e2af94f408531bdace228ca796

Observation a1edc71a-d655-4c6e-a2bd-beff4a77fe74 · outbound

This paper cites TAPAS: Thermal-and power-aware scheduling for LLM inference in cloud platforms.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda TAPAS: Thermal-and power-aware scheduling for LLM inference in cloud platforms

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:35:22.384097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:73740829bee12e0c0d2f80b11bceb787dfbc5c0c43872a5485151ce81955f88f

Observation 26b0ea8f-3cbd-4e8c-a4c6-80e7b21ac568 · outbound

This paper cites throttLL’eM: Predictive GPU Throttling for Energy Effi- cientLLMInferenceServing.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda throttLL’eM: Predictive GPU Throttling for Energy Effi- cientLLMInferenceServing.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:35:22.381085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:b76c72ff6ee149aad0947f11e1920076e3309f1c775cd24d5f5f07ebe2474488

Observation b59494f5-6e33-4c66-a087-a3e4edb67882 · outbound

This paper cites Medusa: Accelerating Serverless LLM Inference with Materialization.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Medusa: Accelerating Serverless LLM Inference with Materialization

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.828861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:1edba523fa726364daf9789103567e070ae1c5186995dc00cf8ff97f88d7f93c

Observation 54a4aea7-0576-4baa-b0db-10a051ecd40c · outbound

This paper cites BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.836715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:48c70991caed7a8a588a3f9fdaadf44739c336f101b123ff38905e32c2115fdf

Observation 04ae0b47-ae62-40d0-b1e8-1a3e30cfa5e5 · outbound

This paper cites 2025:415–430.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2025:415–430

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.952789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:ed8c42c867cc3e41d3cee301db7fb56b04295315821a8d6ed639a070b577907e

Observation 72bfd0a9-e6db-497c-a0ff-9edcd31fa148 · outbound

This paper cites an unresolved cited work.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-05-21T16:34:16.965952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:e60f6f7f7251bfa79a1e734f4948bd3f7c7f888876ba3d4d2c0f71f1d2d5e4e8

Observation 0c6c0afc-6793-4665-89fe-b29eadc0bca2 · outbound

This paper cites FPGA-Based Sparse Matrix Multiplication Accelerators: From State-of-the-Art to Future Opportunities.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda FPGA-Based Sparse Matrix Multiplication Accelerators: From State-of-the-Art to Future Opportunities

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.946047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:d60382b860b8357af56eac1f5f19ca361287995a74c9497339baffdc8ebc7305

Observation 2a152618-5713-4a2f-8ee9-7e371ce6e873 · outbound

This paper cites A survey on hardware accelerators for large language models.Applied Sciences.2025;15(2):586.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda A survey on hardware accelerators for large language models.Applied Sciences.2025;15(2):586

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.984399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:990486a2c5b46353406c3c8bfbaa34875253b030315fed21766a4fc19f1c32cd

Observation 0f49a0d9-8768-41c1-a0af-0ce3f0c5d290 · outbound

This paper cites Unleashing the Potential of LLMs for Quantum Computing: A Study in Quantum Architecture Design.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unleashing the Potential of LLMs for Quantum Computing: A Study in Quantum Architecture Design

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:31:30.829116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:cf4f4ac3035c156931b32b2b0188db2bd94c5f2954d73ee0955355522d0eeb2f

Observation 0b92a5b2-d53e-4cbb-90aa-981378b81a51 · outbound

This paper cites Efficient Large Language Models: A Survey.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Efficient Large Language Models: A Survey

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.836748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:a54b796daa40b7e62a254d1f7bc2c0847775f18133edb6084bce01b19c6d48a3

Observation f1402783-282e-4a23-b33e-67340aeed24a · outbound

This paper cites an unresolved cited work.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-05-21T16:34:17.393500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:7affa6a95723266a8a2a7416c0b38a0fa828682d9264ad18340958eaa3f88aec

Observation 5152f1ce-90b0-4831-8771-e6593bd6f7a3 · outbound

This paper cites an unresolved cited work.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-05-21T16:34:16.940624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:35a2158448e87d6467bd0d6c2cc6dc76ce72b06e16147f3e7ccf050eb30ed289

Observation e806a45e-cdc7-457f-a409-f22f2c480eb0 · outbound

This paper cites LLM-based cost-aware task scheduling for cloud computing systems.Journal of Cloud Computing.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda LLM-based cost-aware task scheduling for cloud computing systems.Journal of Cloud Computing

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:16.943139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:2017c15e0eb1160684ebb74ecc469f9f7b70fca7b49b1e015168490ddb029d10

Observation 32bdcac2-74ec-4c92-80f4-a718d4355b33 · outbound

This paper cites Green scheduling for LLM workloads with model and data reuse across geo-distributed data centers.Digital Communications and Networks.2025.

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Green scheduling for LLM workloads with model and data reuse across geo-distributed data centers.Digital Communications and Networks.2025

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T16:34:17.388523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:27:23.580445Z digest=sha256:d19e6422bd00911bd7e18662f76ebff7d536d1c300d9ba6532fd2b2fac0c7bbd

Pith citing papers

Observation 4804a744-832c-4fc8-acef-932896006489 · inbound

BrownoutMoE: Structure-Aware Expert Grouping for Efficient and Accurate LLM Web-based Services cites this paper.

BrownoutMoE: Structure-Aware Expert Grouping for Efficient and Accurate LLM Web-based Services Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T21:14:49.980623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:14:49.980623Z digest=sha256:24fdd08cc70434e43a2b35bdab57e0a46672a79442f23078450e4ceabe3118a9

Observation ab4d15bd-1b8b-4b4a-8b7c-eedd876d3522 · inbound

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving cites this paper.

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T21:08:24.159706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:08:24.159706Z digest=sha256:508edc50099e3fc404a6cbf2332a447d7322e3fd788007c6568b2bb04b91cdb5