Pith. sign in

Paper Citation Record · LEDGER

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters

As of 23 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2604.16400.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.16400 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T10:33:00.445749Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T12:16:10.904456Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact7
  • verified fuzzy32
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 775ea295-3671-4c55-a06b-644f12684d2b · outbound

This paper cites (2023) Github copilot.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters (2023) Github copilot

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.876878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:9ed15738e37d9d6abb04f0a37d1d6457848cf066d58270c0c865156758b6775f

Observation 9192094e-7af7-40e9-a642-b62d659bdfcd · outbound

This paper cites an unresolved cited work.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-21T10:34:07.873951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:47ea0c0b24394d543478b75bc4598964e08fbecadda622b93da0b0d9d9f90c4e

Observation ed653408-5948-4088-9d6f-0c33e71ca280 · outbound

This paper cites (2022) Chatgpt: Optimizing language models for dialogue.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters (2022) Chatgpt: Optimizing language models for dialogue

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.871447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:df60a665030620402f27b2801a33aae933a85070b6a7a79d40002e8e39fc0327

Observation 65a5ec62-2864-4bc9-b10d-8c05786e360f · outbound

This paper cites A review on edge large language models: Design, execution, and applications.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters A review on edge large language models: Design, execution, and applications

Reference 4

Resolution
verified exact
doi, observed 2026-05-21T10:34:06.931096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:49b7ab868c94407bf83bdf6b3061cc2fda7221a60d262229fb6ffc3455af29a7

Observation f2113891-e3dd-47f7-8a76-d74edbc4d43e · outbound

This paper cites Mobile edge intelligence for large language models: A contemporary survey.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Mobile edge intelligence for large language models: A contemporary survey

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.861644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:455bb2e6e851bcb0b3868ecb8213f166e25ecb89dcef3bcb227ab47afbfdf979

Observation f6bbf74a-4cbd-4060-bf8d-106b9bad304c · outbound

This paper cites Exploring parameter-efficient fine-tuning to enable foundation models in feder- ated learning.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Exploring parameter-efficient fine-tuning to enable foundation models in feder- ated learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.851757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:15c7bc2c67e755518a3df56a0032d3e3087161e4310a5571eff4af81a2cd2cd8

Observation 96d713be-7495-44c1-8ddb-92f6f4c5a614 · outbound

This paper cites Heterogeneous LoRA for federated fine-tuning of on-device foundation models.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Heterogeneous LoRA for federated fine-tuning of on-device foundation models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.854222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:5222f772801ae7f6532922ee64a65befdccddeec7d1d68e206c9255164e79118

Observation e1195fcc-cfbb-4e19-969c-390963878770 · outbound

This paper cites Federated fine-tuning for pre-trained foundation models over wireless networks.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Federated fine-tuning for pre-trained foundation models over wireless networks

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:34:06.923091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:6ae7fc01ced3018559ce46c2cad4f081eb41646b2e6a738aa52e90c40fe420d8

Observation 2ddfaebd-12b1-497d-8eda-e1f240aff3bf · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters LoRA: Low-rank adaptation of large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.848824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:ca03434d4a28f38a43393af12bcb84a8ee64572d77294dfca0cdca92e212ef06

Observation c3cbe103-acca-4c20-a093-1928c8fbbe31 · outbound

This paper cites Beyond scale: the diversity coefficient as a data quality metric demonstrates llms are pre-trained on formally diverse data.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Beyond scale: the diversity coefficient as a data quality metric demonstrates llms are pre-trained on formally diverse data

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.843236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:e136e4df5201b44703201645fb2398c6ef2f93878e63a70534096830bb7704b4

Observation 49bcd51f-07f6-4fe6-8c07-ffcd311d5c5a · outbound

This paper cites DeepBoot: Dynamic Scheduling System for Training and Inference Deep Learning Tasks in GPU Cluster.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters DeepBoot: Dynamic Scheduling System for Training and Inference Deep Learning Tasks in GPU Cluster

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.845891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:6a2e2cf417fbdb1d4f9c4e9602c6920e085b3b9be0d457b002115fb33a0a6865

Observation 3ddc8af0-704b-4d46-997c-109aed6be3f6 · outbound

This paper cites Multiplexing dynamic deep learning workloads with slo-awareness in gpu clusters.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Multiplexing dynamic deep learning workloads with slo-awareness in gpu clusters

Reference 12

Resolution
verified exact
doi, observed 2026-05-21T10:34:06.934184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:66e74ea6c03cdd565234fa4836db95dd0fbce4da2208dc5f7ea104a16660202d

Observation 4b8318f0-0267-4140-b2ee-3a1106ac983d · outbound

This paper cites Scaling Laws for Neural Language Models.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Scaling Laws for Neural Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T10:34:07.208947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:6b1fbd45e48a159a1c7774ac4d6d6f098c84fd54574085f43033a94781511afa

Observation db49b4a8-191e-4a31-b0ac-7b94630a49fe · outbound

This paper cites Lyra: Elastic scheduling for deep learning clusters.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Lyra: Elastic scheduling for deep learning clusters

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:34:06.930067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:8a83d5e015885d9a4cfe776486c0415bced8ca01d33ef2313fec695364c7c508

Observation 2c5c4f4c-e305-4ee6-9ee7-b5077943bad7 · outbound

This paper cites Serving hetero- geneous machine learning models on Multi-GPU servers with Spatio- Temporal sharing.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Serving hetero- geneous machine learning models on Multi-GPU servers with Spatio- Temporal sharing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.834860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:e2e8b5059166c74eaefc62a0372baa7e23d9264ba705d6b5e60f5374ffdbe7a4

Observation 2a03124e-a738-473a-b4ff-8092765c21fd · outbound

This paper cites (2023) Multi-process service (mps).

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters (2023) Multi-process service (mps)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.837262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:9ed6b24e4099b6a372de688eb0c1d2d7db2325167c8c7989e0b2fbcc06054ff5

Observation 5e589751-e464-4d9e-87e2-0527db0ee71a · outbound

This paper cites Shepherd : Serving DNNs in the Wild.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Shepherd : Serving DNNs in the Wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.831997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:dcf516168a2c5457b4f74d6d161a2ec142d5e2e32b38a9c5e6a4babf2daa8bc1

Observation 77f0ca7c-c4e4-40c6-bed2-4fa9fa9740ae · outbound

This paper cites Federated Learning while Pro- viding Model as a Service: Joint Training and Inference Optimization.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Federated Learning while Pro- viding Model as a Service: Joint Training and Inference Optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.827137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:b3977608e08c830e4e431cf1d382cfd42a27b88cb5877955670c2e158e211e1c

Observation cba83386-c1de-4b08-9c7d-0c460e8cc593 · outbound

This paper cites Communication-efficient learning of deep networks from decentralized data.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Communication-efficient learning of deep networks from decentralized data

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.830485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:610b4343081bd9c1dbcd846f4288b5cb2a17cd090d1870b4cb96916505ab38d8

Observation 5b5e2263-c77a-47ff-91f4-4d3f828cf7f8 · outbound

This paper cites Fedadapt: Adaptive offloading for iot devices in federated learning.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Fedadapt: Adaptive offloading for iot devices in federated learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.866515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:0f9cf6aefedca9468ed47c91271828b25f9e8301ad6100a592c4fd82d4e43e3e

Observation 074ad771-90d2-4000-8bd0-d8228d83cb25 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Qlora: Efficient finetuning of quantized llms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.868958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:3b47f2fcbb4006fd04fefa06d7c0ac9acadd9a5dc6b8e7dd232383af223499ca

Observation e3827efd-b301-4c5a-b170-b681ef926f0d · outbound

This paper cites FedPara: Low-Rank Hadamard Product for Communication-Efficient Federated Learning.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters FedPara: Low-Rank Hadamard Product for Communication-Efficient Federated Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:34:07.204964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:d7164bbd435950b4ecd93f62a520ec5eed1be3b13270488ad56ddeced80e704d

Observation 993e34c6-12b1-4303-a6ee-487bc8fb2f64 · outbound

This paper cites Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:34:07.212972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:7ed59b372027cb8649b97100e258f5c1e7b154548ce1bbe8760440051f305708

Observation 4e16538c-c1d6-4348-9893-df8789664171 · outbound

This paper cites Don’t decay the learning rate, increase the batch size.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Don’t decay the learning rate, increase the batch size

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.829654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:b15675d6e74e458d50cbe46e6cc2a77844e3ecbdc0671e26613e17b927364c61

Observation 856899ea-8e0c-4e58-9267-f87205099c00 · outbound

This paper cites Efficient Coordination of Federated Learning and Inference Offloading at the Edge: A Proactive Optimization Paradigm.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Efficient Coordination of Federated Learning and Inference Offloading at the Edge: A Proactive Optimization Paradigm

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.866062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:e82529fc5babb4e9a0a01471eb313f7b63ca01c38532e7eed7a29880840fc23a

Observation 1c6d0ce5-703b-49f2-bc55-8b14ca1d3a10 · outbound

This paper cites Human-in-the-loop machine learning: a state of the art.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Human-in-the-loop machine learning: a state of the art

Reference 26

Resolution
verified exact
doi, observed 2026-05-21T10:34:06.936589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:4b937b9572b3b27dd86bd11f484d793ff55c23b30bb28e5c17a064885d70b40d

Observation 6392a0aa-4de5-4cb6-8a33-36c724ff6297 · outbound

This paper cites Illustrating reinforcement learning from human feedback (rlhf).

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Illustrating reinforcement learning from human feedback (rlhf)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.819819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:08e7d0c1698f10a810145baaf8e933c54fec1c782ec36fe90e37891a7cf2351e

Observation 5c17ec63-b5f9-4c24-931e-c6dacb33ca8d · outbound

This paper cites manim code.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters manim code

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.812887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:021451a68fcca6b32503d47593d538b9ff8b8b9d530ddf3ecf41ab84256ace1f

Observation ff13a1f7-48aa-4196-b4d1-a180755a37da · outbound

This paper cites Codealpaca-20k.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Codealpaca-20k

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.859220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:6523071d2dcfe5bfdff4c329a0c29b1dde7503162ec29e56f86267e28d448b2f

Observation ffbc9b68-049d-469a-849c-4b136ccf2b18 · outbound

This paper cites code instructions 120k alpaca.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters code instructions 120k alpaca

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.810463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:90ce71e03d72416a32a84f54b913abb2b45f0609b3b348df500560f83648834d

Observation 2ea815c5-59d5-4377-9774-f45126575679 · outbound

This paper cites an unresolved cited work.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-21T10:34:07.822289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:cacd08de62d90ba114bebab1aa4dd249905ce376d6cfc8b2d876b67daf855fdc

Observation 376dbce9-94fa-438d-a184-177b33a9420d · outbound

This paper cites Gpteacher-general-instruct.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Gpteacher-general-instruct

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.806230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:7f74863442e7caa8c2bf202b8883f2727518eac1d7a5f265ab89ff19f8465457

Observation cd0fcd64-4f94-4791-859e-c6fc646776e7 · outbound

This paper cites open-instruct-v1.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters open-instruct-v1

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.863998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:89418932b0aa8e9a6c52227cc27ae9c769fafb2054b9b148f00f7c5623f99465

Observation 7ba89d86-4fc0-48c2-a7bd-96e8b5223ab0 · outbound

This paper cites URL https://doi.org/10.1109/HPCA61900.2025.00102.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters URL https://doi.org/10.1109/HPCA61900.2025.00102

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:34:06.923936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:bd0ad1c2e602df69b9fa4b6704ede586c1af1698a3499564c53c5de65c6026fa

Observation 691d4f20-924d-4d59-b161-6d5e23210a1a · outbound

This paper cites dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.801803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:4f90f39eb63519af6d9ee9794a80e18381b98d6b8d6f9b9339dc1f50b30b3b1f

Observation 6bb1b16b-5f8c-4308-aec2-d01ccc6cf215 · outbound

This paper cites Peft: State-of-the-art parameter-efficient fine-tuning methods.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Peft: State-of-the-art parameter-efficient fine-tuning methods

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.798265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:c88764e970f7248edaf6d34cae8f245076d815cbc076f3433a389e78a126db0c

Observation 9deabcba-3c8a-4423-9614-3d4b59ffd3a6 · outbound

This paper cites Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Opti- mization.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Opti- mization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.879424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:e17fdadb7365f788eb538dca1e7af0971b9c39408dadfc1418b0b89d889f3414

Observation b9b026c3-cf21-40d7-9e7c-92d9739bbaf2 · outbound

This paper cites Adaptive parameter-efficient federated fine-tuning on heterogeneous devices.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Adaptive parameter-efficient federated fine-tuning on heterogeneous devices

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.856667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:310fcca80b17744ab7e7cfb14d9f21ae79b3dd07578d45a1a56574cfa17f3eb3

Observation 0ba0af59-037b-42e1-a304-f5b9b825924d · outbound

This paper cites Haflq: Heterogeneous adaptive federated lora fine-tuned llm with quantization.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Haflq: Heterogeneous adaptive federated lora fine-tuned llm with quantization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.795435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:218ba8b97a4e54eee76e5b31422a2fbb0098a9502ab833e35f0f9e28ad057e00

Observation 7e3457b9-2896-4ac6-94b7-9f907d69f673 · outbound

This paper cites HAFLQ: Heterogeneous Adaptive Federated LoRA Fine-tuned LLM with Quantization.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters HAFLQ: Heterogeneous Adaptive Federated LoRA Fine-tuned LLM with Quantization

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:34:07.201257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:be6c5ca8c5c632bb1662413f406d91b87f69b296c8fd46fca30fea7ba3742ef6

Observation dce4c978-3b52-49dd-ad44-1d2bd507a384 · outbound

This paper cites FwdLLM: Efficient Feder- ated Finetuning of Large Language Models with Perturbed Inferences.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters FwdLLM: Efficient Feder- ated Finetuning of Large Language Models with Perturbed Inferences

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.816008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:6753baa4671d44a1b188eeb73d3c90076f917a39fabb1eba4152421a7c85bf22

Observation eda66662-e4db-46cd-b83d-02ad304d67c3 · outbound

This paper cites Partitioned collaborative inference for on-device models via evolution- ary reinforcement learning.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Partitioned collaborative inference for on-device models via evolution- ary reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.851507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:355e9a229f86b34027b492a717fe2d2a36fde8f7ab17a537b54fa80816ed0fc9

Observation a0263514-3ee5-4ba5-8939-f3e0a02b3bae · outbound

This paper cites an unresolved cited work.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-21T10:34:07.863773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:be073f88cc9f6e4b61b8032d1b83427b2c5d947c908b1054b275d7584d8648c7

Observation 91d9592c-0a08-4e0b-bf58-a671bfb206fe · outbound

This paper cites Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and Inference.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and Inference

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.791808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:e67ced9cac233906783b79775b7d5725f0e61823018969cf064f8c54ac673c62

Observation 9e09f437-8138-45aa-9a9b-5e8af549157f · outbound

This paper cites LLMStation: Resource Multiplexing in Tuning and Serving Large Language Models.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters LLMStation: Resource Multiplexing in Tuning and Serving Large Language Models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T10:34:07.786283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:a74d96fde6e11df0b0d6ff09a7af4891b3af06f147381cd358a172e05335b4b9

Observation 5fd616b3-b39c-44b0-af8c-264fbae54aa0 · outbound

This paper cites Flexllm: Token-level co-serving of llm inference and finetuning with slo guarantees.

CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters Flexllm: Token-level co-serving of llm inference and finetuning with slo guarantees

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:34:07.216605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:33:00.445749Z digest=sha256:d29543b81ad0bb44557f00c9eb5c5aad884c534536715b48a74c17b6ac9d237e

Pith citing papers

Observation 20c7d213-03f7-42e7-b8d9-1be3412ad707 · inbound

Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure cites this paper.

Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-12T12:18:48.204349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-12T12:16:10.904456Z digest=sha256:fd2003367dd4dbc639678b2ad15f8c1e36db2b6a5e6de9eabf045f82b9c7cc43