Pith. sign in

Paper Citation Record · LEDGER

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters

As of 4 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2604.16400.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.16400 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T10:33:00.445749Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T12:16:10.904456Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-12T18:15:00.917874Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact7
  • verified fuzzy32
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 775ea295-3671-4c55-a06b-644f12684d2b · outbound

This paper cites (2023) Github copilot.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters (2023) Github copilot

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.876878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:6a8d01c04ad403f1b7aa03b31416eec3971564bad735b2abad9db3e4b03ce625

Observation 9192094e-7af7-40e9-a642-b62d659bdfcd · outbound

This paper cites an unresolved cited work.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-21T10:34:07.873951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:cac529c5a39b6add704e6072cc910743c7b97aa45e64aea975f8ded9508f1b48

Observation ed653408-5948-4088-9d6f-0c33e71ca280 · outbound

This paper cites (2022) Chatgpt: Optimizing language models for dialogue.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters (2022) Chatgpt: Optimizing language models for dialogue

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.871447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:0103ff11dcc6d89b3144356096b7cde545550067f218f6e77cdf8c0fe5ae096c

Observation 65a5ec62-2864-4bc9-b10d-8c05786e360f · outbound

This paper cites A review on edge large language models: Design, execution, and applications.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters A review on edge large language models: Design, execution, and applications

Reference 4

Resolution
verified exact
doi, observed 2026-05-21T10:34:06.931096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:f437852325e915e0948da12b1d12185d759c3ab8c7e822bc965da7d02e802af9

Observation f2113891-e3dd-47f7-8a76-d74edbc4d43e · outbound

This paper cites Mobile edge intelligence for large language models: A contemporary survey.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Mobile edge intelligence for large language models: A contemporary survey

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.861644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:75acd5ed4319f8df7b13055389b1181e3d92da10b34134c193ee59f61cc10655

Observation f6bbf74a-4cbd-4060-bf8d-106b9bad304c · outbound

This paper cites Exploring parameter-efficient fine-tuning to enable foundation models in feder- ated learning.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Exploring parameter-efficient fine-tuning to enable foundation models in feder- ated learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.851757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:8bb2f07af8af227bcb2ba390b52c998f92975938c4ab06e59eec95ae9225a6d3

Observation 96d713be-7495-44c1-8ddb-92f6f4c5a614 · outbound

This paper cites Heterogeneous LoRA for federated fine-tuning of on-device foundation models.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Heterogeneous LoRA for federated fine-tuning of on-device foundation models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.854222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:7a3ecd33f06898ebdd1676ba3a456675d120e9a6b87b56b36ee4e2022e46d77e

Observation e1195fcc-cfbb-4e19-969c-390963878770 · outbound

This paper cites Federated fine-tuning for pre-trained foundation models over wireless networks.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Federated fine-tuning for pre-trained foundation models over wireless networks

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:34:06.923091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:a5fb81e3f5743a192574e66c662b2c84600c412c53f9ca06e8cac6506a612b81

Observation 2ddfaebd-12b1-497d-8eda-e1f240aff3bf · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters LoRA: Low-rank adaptation of large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.848824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:6ec5b64b8a2bb0c054e477282aed51198313823ca25b58008c1fda98fe9dbe57

Observation c3cbe103-acca-4c20-a093-1928c8fbbe31 · outbound

This paper cites Beyond scale: the diversity coefficient as a data quality metric demonstrates llms are pre-trained on formally diverse data.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Beyond scale: the diversity coefficient as a data quality metric demonstrates llms are pre-trained on formally diverse data

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.843236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:6c1f31d383630dd4e0d1f0d1c0ca1c0d0d8db486027f071691c811a9d11d3ae7

Observation 49bcd51f-07f6-4fe6-8c07-ffcd311d5c5a · outbound

This paper cites DeepBoot: Dynamic Scheduling System for Training and Inference Deep Learning Tasks in GPU Cluster.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters DeepBoot: Dynamic Scheduling System for Training and Inference Deep Learning Tasks in GPU Cluster

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.845891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:d839d6a5a5589a28c4d1b8b8b8658a8550f354dfbe3929220f9623bf08e0b746

Observation 3ddc8af0-704b-4d46-997c-109aed6be3f6 · outbound

This paper cites Multiplexing dynamic deep learning workloads with slo-awareness in gpu clusters.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Multiplexing dynamic deep learning workloads with slo-awareness in gpu clusters

Reference 12

Resolution
verified exact
doi, observed 2026-05-21T10:34:06.934184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:8d95293feb76e3c7822e3f39c08c57cf2640c212ecdc8dfa665170a73fd60e85

Observation 4b8318f0-0267-4140-b2ee-3a1106ac983d · outbound

This paper cites Scaling Laws for Neural Language Models.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Scaling Laws for Neural Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T10:34:07.208947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:5a5a22f9b7c2a4606b929b7e8981549703d6f8a1ed649d445695d1bd825811fe

Observation db49b4a8-191e-4a31-b0ac-7b94630a49fe · outbound

This paper cites Lyra: Elastic scheduling for deep learning clusters.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Lyra: Elastic scheduling for deep learning clusters

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:34:06.930067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:7788b3d6a43849e27baf95a9609f21c36ff0d987824afb8202a0aaf56082b217

Observation 2c5c4f4c-e305-4ee6-9ee7-b5077943bad7 · outbound

This paper cites Serving hetero- geneous machine learning models on Multi-GPU servers with Spatio- Temporal sharing.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Serving hetero- geneous machine learning models on Multi-GPU servers with Spatio- Temporal sharing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.834860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:9b2a194069f1a531a506abfd7659c78662784aafde58f195314512c94ad8a75a

Observation 2a03124e-a738-473a-b4ff-8092765c21fd · outbound

This paper cites (2023) Multi-process service (mps).

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters (2023) Multi-process service (mps)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.837262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:2b2199459b97f1441ce76b3bb9fb695b8d84b14a32991b1b028826a8e50a80b0

Observation 5e589751-e464-4d9e-87e2-0527db0ee71a · outbound

This paper cites Shepherd : Serving DNNs in the Wild.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Shepherd : Serving DNNs in the Wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.831997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:bec2c2274814584460a8340c5cae73e8e5054bb69edffef889a3b5c3c3443210

Observation 77f0ca7c-c4e4-40c6-bed2-4fa9fa9740ae · outbound

This paper cites Federated Learning while Pro- viding Model as a Service: Joint Training and Inference Optimization.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Federated Learning while Pro- viding Model as a Service: Joint Training and Inference Optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.827137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:16d53be7aa7fa2cd480b6e9e3ec408df4562680b0e57e411ae7c775fa4694308

Observation cba83386-c1de-4b08-9c7d-0c460e8cc593 · outbound

This paper cites Communication-efficient learning of deep networks from decentralized data.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Communication-efficient learning of deep networks from decentralized data

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.830485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:ac206658f359fb4771c302e1b0bc0cf1cde6333dc2a25ca41b60fc4dab964ac0

Observation 5b5e2263-c77a-47ff-91f4-4d3f828cf7f8 · outbound

This paper cites Fedadapt: Adaptive offloading for iot devices in federated learning.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Fedadapt: Adaptive offloading for iot devices in federated learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.866515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:2bec92f401865749de810599dd4821374bdf44756f8ec5f6e0f68ba37fdf31f1

Observation 074ad771-90d2-4000-8bd0-d8228d83cb25 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Qlora: Efficient finetuning of quantized llms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.868958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:5b5d2d4d3fb4c963e3afac96a8d44fc69809cc7c875efcb359034df1889c6a05

Observation e3827efd-b301-4c5a-b170-b681ef926f0d · outbound

This paper cites FedPara: Low-Rank Hadamard Product for Communication-Efficient Federated Learning.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters FedPara: Low-Rank Hadamard Product for Communication-Efficient Federated Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:34:07.204964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:93a7c9c529228e13be0539e052422124e6c44c8839039969498ccdc834a3c817

Observation 993e34c6-12b1-4303-a6ee-487bc8fb2f64 · outbound

This paper cites Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:34:07.212972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:b3ac81248b5e8cba57c4d3062703243f990ee01563bfa4dacd3b955bea1eadeb

Observation 4e16538c-c1d6-4348-9893-df8789664171 · outbound

This paper cites Don’t decay the learning rate, increase the batch size.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Don’t decay the learning rate, increase the batch size

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.829654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:4d1a5e03c351271c58aaf5546642858b86d33354735594f61513bebc0e0b5a83

Observation 856899ea-8e0c-4e58-9267-f87205099c00 · outbound

This paper cites Efficient Coordination of Federated Learning and Inference Offloading at the Edge: A Proactive Optimization Paradigm.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Efficient Coordination of Federated Learning and Inference Offloading at the Edge: A Proactive Optimization Paradigm

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.866062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:619c44711808af9cbc5675cab58b2d3356e67a994b3654591fadef11880fb774

Observation 1c6d0ce5-703b-49f2-bc55-8b14ca1d3a10 · outbound

This paper cites Human-in-the-loop machine learning: a state of the art.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Human-in-the-loop machine learning: a state of the art

Reference 26

Resolution
verified exact
doi, observed 2026-05-21T10:34:06.936589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:c33613277c975ed9fbcfccc4f3ff83925d05f385194c7a5e261cca4373ad12b9

Observation 6392a0aa-4de5-4cb6-8a33-36c724ff6297 · outbound

This paper cites Illustrating reinforcement learning from human feedback (rlhf).

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Illustrating reinforcement learning from human feedback (rlhf)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.819819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:7e454f769f12ef74bddf51411883682e8711b973276162e85fd8fcb4a4a20eab

Observation 5c17ec63-b5f9-4c24-931e-c6dacb33ca8d · outbound

This paper cites manim code.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters manim code

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.812887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:aff1853eb247671035e1a5b5f817e80b336fbc132b46f01bf45bc43aaa1e2e73

Observation ff13a1f7-48aa-4196-b4d1-a180755a37da · outbound

This paper cites Codealpaca-20k.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Codealpaca-20k

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.859220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:2088bf949d36cff7bfd0106db741fc4f9cda5fd007d54f9455363849170b75bf

Observation ffbc9b68-049d-469a-849c-4b136ccf2b18 · outbound

This paper cites code instructions 120k alpaca.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters code instructions 120k alpaca

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.810463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:76a3096afbd56c0f6f8d819b44e29ceb6dcfa62abd884f5f1b11d241465eed90

Observation 2ea815c5-59d5-4377-9774-f45126575679 · outbound

This paper cites an unresolved cited work.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-21T10:34:07.822289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:0fe6d7aef8c97a2ab09b6329aa1c598cca4f2f7f7e3459f5dedaec5530e1047c

Observation 376dbce9-94fa-438d-a184-177b33a9420d · outbound

This paper cites Gpteacher-general-instruct.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Gpteacher-general-instruct

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.806230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:8a8f26e731c0b4e496abb8660e20363d802f1772e99554a6ecbfd4e1ce8962f8

Observation cd0fcd64-4f94-4791-859e-c6fc646776e7 · outbound

This paper cites open-instruct-v1.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters open-instruct-v1

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.863998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:3d0a50b1b60e3c8d343ca1a163a5550b29c326241c3789e25a391b92345040d8

Observation 7ba89d86-4fc0-48c2-a7bd-96e8b5223ab0 · outbound

This paper cites URL https://doi.org/10.1109/HPCA61900.2025.00102.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters URL https://doi.org/10.1109/HPCA61900.2025.00102

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:34:06.923936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:f8bd1e895959e8482baf118ae5d31fc45986399e83993d002f6119d75c73eefb

Observation 691d4f20-924d-4d59-b161-6d5e23210a1a · outbound

This paper cites dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.801803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:92fe46db300cee6752de7c4b57b1ef8f68fc36a4760ae014685f5db64616ee34

Observation 6bb1b16b-5f8c-4308-aec2-d01ccc6cf215 · outbound

This paper cites Peft: State-of-the-art parameter-efficient fine-tuning methods.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Peft: State-of-the-art parameter-efficient fine-tuning methods

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.798265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:ca9a6b82f248322c1d28066412af4188d820497b4968fa391f068f27e9d8b708

Observation 9deabcba-3c8a-4423-9614-3d4b59ffd3a6 · outbound

This paper cites Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Opti- mization.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Opti- mization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.879424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:5ba30023b916dd0e1ce31d15304b3cae3b6091e20e75fe5457919b6545a97200

Observation b9b026c3-cf21-40d7-9e7c-92d9739bbaf2 · outbound

This paper cites Adaptive parameter-efficient federated fine-tuning on heterogeneous devices.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Adaptive parameter-efficient federated fine-tuning on heterogeneous devices

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.856667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:abac72bcaebe884418b6f6850e48662fccab7495c4408cfba14ac64f2c19d1e3

Observation 0ba0af59-037b-42e1-a304-f5b9b825924d · outbound

This paper cites Haflq: Heterogeneous adaptive federated lora fine-tuned llm with quantization.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Haflq: Heterogeneous adaptive federated lora fine-tuned llm with quantization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.795435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:099aa7d4fa0a6fee0084829c29d1d2749eb7e0e9df5d5f7e8957876ddd47ed5f

Observation 7e3457b9-2896-4ac6-94b7-9f907d69f673 · outbound

This paper cites HAFLQ: Heterogeneous Adaptive Federated LoRA Fine-tuned LLM with Quantization.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters HAFLQ: Heterogeneous Adaptive Federated LoRA Fine-tuned LLM with Quantization

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:34:07.201257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:9e7f13d6034c9c6277a1cdfe9a789fecf6bc3d887f850a0cafa6f2dc486796d5

Observation dce4c978-3b52-49dd-ad44-1d2bd507a384 · outbound

This paper cites FwdLLM: Efficient Feder- ated Finetuning of Large Language Models with Perturbed Inferences.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters FwdLLM: Efficient Feder- ated Finetuning of Large Language Models with Perturbed Inferences

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.816008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:ee5a43ae46a3a493325c8ab53f1131c21469c381fef0a4430967c93f10579f8d

Observation eda66662-e4db-46cd-b83d-02ad304d67c3 · outbound

This paper cites Partitioned collaborative inference for on-device models via evolution- ary reinforcement learning.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Partitioned collaborative inference for on-device models via evolution- ary reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.851507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:80a11d41473f6b79eea90518ae9da1666eafb5aef78a0cc6c7195806ab6faa19

Observation a0263514-3ee5-4ba5-8939-f3e0a02b3bae · outbound

This paper cites an unresolved cited work.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-21T10:34:07.863773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:b1ba5a7ae04d638dfb5061d87837858822b69feb4749c5675f897f2d76c892f3

Observation 91d9592c-0a08-4e0b-bf58-a671bfb206fe · outbound

This paper cites Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and Inference.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and Inference

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.791808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:2c550f421016765cf604ca840fe26574f4c7d3a6adfc3017a94e1bd3d17c4b86

Observation 9e09f437-8138-45aa-9a9b-5e8af549157f · outbound

This paper cites LLMStation: Resource Multiplexing in Tuning and Serving Large Language Models.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters LLMStation: Resource Multiplexing in Tuning and Serving Large Language Models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.786283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:a67c3f6a1a06a639d884d17f91d2a8f06ed3bb4ce386f0c3134cbfb601036c76

Observation 5fd616b3-b39c-44b0-af8c-264fbae54aa0 · outbound

This paper cites Flexllm: Token-level co-serving of llm inference and finetuning with slo guarantees.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Flexllm: Token-level co-serving of llm inference and finetuning with slo guarantees

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:34:07.216605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:2e92d20c664464c8a0eb41ec4153394a100431aeb22cfdc0a641ebe958a73b19

Pith citing papers

Observation 20c7d213-03f7-42e7-b8d9-1be3412ad707 · inbound

Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure cites this paper.

Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-12T12:18:48.204349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-07-12T12:16:10.904456Z digest=sha256:ca47074bebf205ea598fe39a6c2235d90fadd592b19c18fc8d513b8d5dfdb657