Pith. sign in

Paper Citation Record · LEDGER

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts

As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2508.00234.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00234 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:22:42.600721Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact4
  • verified fuzzy24
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7efcac1-e101-4a8f-aee0-b89606f8e866 · outbound

This paper cites AIoT smart home via autonomous LLM agents,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts AIoT smart home via autonomous LLM agents,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.090189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.438667Z digest=sha256:3867857dc759ebcd0f01bad6ec36516c1c4e76f8772c041a823ccbc5f27bb48b

Observation 5becd7f5-7a18-4a0b-8b99-0a7ff6399765 · outbound

This paper cites Large language models for human- ai co-creation of robotic dance performances,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Large language models for human- ai co-creation of robotic dance performances,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.079977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.442311Z digest=sha256:1e8cdac92df651f46d3c989e25630361bed5c515da8a5c3b74587d04d473cae4

Observation 66562e3f-b0ec-43cb-95a3-6d327ed826de · outbound

This paper cites EdgeFM: Leveraging foundation model for open-set learning on the edge,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts EdgeFM: Leveraging foundation model for open-set learning on the edge,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.070683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.445752Z digest=sha256:c4d22870b6644eb2176aa0a2b20894d6f737b277b217631ce2a7d1f97a3628c5

Observation fa1a07e2-b807-4ec9-afdd-780f71e7630e · outbound

This paper cites WDMoE: Wireless Distributed Large Language Models with Mixture of Experts.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts WDMoE: Wireless Distributed Large Language Models with Mixture of Experts

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:22:42.849405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.449489Z digest=sha256:a9a77439eb8a6e291fc49cb5df8949cb087ef3f67c87d3bc6392b12945e4b5a2

Observation e29751e0-7f8d-4110-a95a-ee4b27e27192 · outbound

This paper cites On Protecting the Data Privacy of Large Language Models (LLMs): A Survey.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts On Protecting the Data Privacy of Large Language Models (LLMs): A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.453414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.453414Z digest=sha256:34dc7eb775b301923db51817307833ed1f3698eb85627b9a7809b6bcecb038c0

Observation 94d46d84-12c1-4b01-9fc2-9714dd1b58c2 · outbound

This paper cites Edge intelligence: Paving the last mile of artificial intelligence with edge computing,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Edge intelligence: Paving the last mile of artificial intelligence with edge computing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.061734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.457777Z digest=sha256:95e664c249350a28b2173b47b0ba04cbc3bdfb8bf9e4d9e76e85a91842d9e4f5

Observation 057fe7f9-4c30-4db0-bcf3-867bc48a5d07 · outbound

This paper cites Enabling AI-Generated Content (AIGC) Services in Wireless Edge Networks.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Enabling AI-Generated Content (AIGC) Services in Wireless Edge Networks

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:22:42.826912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.461675Z digest=sha256:24aa55a8b77c94f7feae784d3ffb2af0580a911c2b5e3a2efa7bf4f878741902

Observation 178e19ba-07ac-4d24-ae51-acba3c2955d2 · outbound

This paper cites Toward Scalable Generative AI via Mixture of Experts in Mobile Edge Networks.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Toward Scalable Generative AI via Mixture of Experts in Mobile Edge Networks

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:22:42.812607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.465930Z digest=sha256:28a2f23419b7a050c06f1214e9e156905dc2350088c47c0a47a98798ba13bb51

Observation a15f1f24-6d08-4389-8ebc-a4dcd44239c4 · outbound

This paper cites LLM-Blender: Ensembling large language models with pairwise ranking and generative fusion,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts LLM-Blender: Ensembling large language models with pairwise ranking and generative fusion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.052200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.469627Z digest=sha256:5d885d3c66d566505c1a2cbdcc3e7b618b639689385261b2fceb8be235456bfc

Observation c1324dbe-b74d-4485-a844-e9374342ebe8 · outbound

This paper cites Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.473633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.473633Z digest=sha256:0bb62f803778528247e68880609de77192b3de9c2d291ff7955c70660e10f441

Observation 5851814e-fe47-4e49-9f96-335532db538d · outbound

This paper cites Orca: A distributed serving system for Transformer-based generative models,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Orca: A distributed serving system for Transformer-based generative models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.042587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.477410Z digest=sha256:00bb728fa4ed62b2d105e0207a38ad38a7af97f55c02edb5243ff7623ed05d29

Observation f57748b3-8b08-4739-8117-89915ee233ad · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Efficient memory management for large language model serving with PagedAttention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.033027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.480957Z digest=sha256:870dd9edc461a582cac43c049421b81720b7c0d79207c1a91b99d7e5cfedc42d

Observation d6fb5336-512d-48cd-8483-794d817d97f1 · outbound

This paper cites TensorOpera Router: A Multi-Model Router for Efficient LLM Inference.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts TensorOpera Router: A Multi-Model Router for Efficient LLM Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.484238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.484238Z digest=sha256:07a34f9001314810583e4d3f6a76e9edff1fe98109cae5b344efa47b1509b7eb

Observation c3c4b001-6546-45ba-b1e2-86507aa5b3cd · outbound

This paper cites Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.487916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.487916Z digest=sha256:8fc05f05e6aba567b809687df0febdb0d1b5414eaadf46f7acd009204f3aa84a

Observation 5d215f12-f43b-4d2c-be98-bc4d39c890e6 · outbound

This paper cites Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.491345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.491345Z digest=sha256:5067ebe4511adb4404fd43483eac812b847c7f05b358b216e5dbf965c8215b6e

Observation 9e1a3c60-1439-493e-8f61-0320821726da · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts RouteLLM: Learning to Route LLMs with Preference Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.495457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.495457Z digest=sha256:af71c209926a2ac11236269a12977a46cfb0690759712550252cc8860a688714

Observation c8f02fcc-0387-4b39-9431-998ed626a0a2 · outbound

This paper cites BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.499001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.499001Z digest=sha256:1142690fd07cb8119eb2efa6018b884fc481d5cc8f902c8b5beba5bc59657947

Observation 6f8399dc-ff72-4234-bc1e-d302dff8fba4 · outbound

This paper cites Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.502527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.502527Z digest=sha256:514f63e64cddfc5cc3e07a1883b5f8e1caacc34dc69affdf627a65cb2575d06b

Observation 5feb3bae-8f4e-4fc5-a454-b3d2bd10255d · outbound

This paper cites S3: Increasing gpu utiliza- tion during generative inference for higher throughput,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts S3: Increasing gpu utiliza- tion during generative inference for higher throughput,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.022410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.506125Z digest=sha256:c8fcf669155b1e2f6fe2d1fa1ff451c00a23cbf77d6539d71f4ace83e5b66647

Observation ff93db18-1f56-45f4-af19-a093d6a6f59e · outbound

This paper cites FlexGen: High-throughput generative inference of large language models with a single gpu,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts FlexGen: High-throughput generative inference of large language models with a single gpu,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.012284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.509785Z digest=sha256:9585bbe9f50e3e40faf3256201fda8ff756e5b75d1bad9d476e7e0ddf4213ff2

Observation 31825be3-8091-4d3c-942e-620e5bc4d5ab · outbound

This paper cites FlashAttention: Fast and memory-efficient exact attention with io-awareness,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts FlashAttention: Fast and memory-efficient exact attention with io-awareness,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.002330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.512936Z digest=sha256:dcddfce160b822b4f8127e56c450277a61b80fce3b0014d908838dd515abf5b1

Observation 8e83768f-be94-4e1b-9b39-1bfced14ec38 · outbound

This paper cites HeteGen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts HeteGen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.992601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.516096Z digest=sha256:1467fdf910ee6306de4154acaf16baece84272f407a6fd8f9cbaeabfc9ba619f

Observation 6a1aaa9b-af2a-45c9-952d-ed89d8481d74 · outbound

This paper cites ExeGPT: Constraint-aware resource scheduling for LLM inference,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts ExeGPT: Constraint-aware resource scheduling for LLM inference,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.983278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.519086Z digest=sha256:c5f8fe9b610512177c9417b5e41a9ef4d9462b3b4e3f3dfc27571df8a57151b2

Observation 120312c5-8335-48a5-8248-6cb87de67a0f · outbound

This paper cites Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.522976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.522976Z digest=sha256:c9253c4fcce7658c0409255aa4c3b43130cb806d3021c686acd84d5699025dfa

Observation ed792d59-abd2-4839-bdd1-2093adadd75e · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Splitwise: Efficient generative LLM inference using phase splitting,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.971972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.526269Z digest=sha256:9a38d9b2b7fbab408576f211d8ac2ea9f830d7d431b0a66ee7af080d0b1f0924

Observation 99811022-c3ff-47f0-adb3-784c8973e78f · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.529491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.529491Z digest=sha256:0e94da7fd3f33c1bfe568fd071c0416940e79ef03aa96e78caa250045a702a83

Observation 29ca24c1-7612-4688-bc65-4af530a53ced · outbound

This paper cites Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.532900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.532900Z digest=sha256:5ea9d5535c4ba97879d8d9b8ccde4ba683ffed3ba92a51f3e244c1516d08a24a

Observation 972eceb0-f50d-4179-990e-02817071252e · outbound

This paper cites Large Language Model Routing with Benchmark Datasets.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Large Language Model Routing with Benchmark Datasets

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.536535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.536535Z digest=sha256:bbeb01ef6626813421f2c313e7aac26bfbe7946144dc785bd1b7aebf7105a70c

Observation 907191d6-53e4-470b-a88a-964cba56d57b · outbound

This paper cites Octopus v4: Graph of language models.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Octopus v4: Graph of language models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:22:42.699685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.540098Z digest=sha256:98c80074b191374f544ba9289831cb9c409e60b96a159abb6b78b2105b7d5e67

Observation 959d65b7-5550-4a95-9b9c-bbbd9e40ae62 · outbound

This paper cites GraphRouter: A Graph-based Router for LLM Selections.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts GraphRouter: A Graph-based Router for LLM Selections

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.543528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.543528Z digest=sha256:eea7dfa1cfe23d571ffbb6e2e3d8f0e2f691e52872326eebad1b395c8f8a37b5

Observation 6f7900a0-9cbd-40a5-9bfd-a4e84da2a1da · outbound

This paper cites Eagle: Efficient Training-Free Router for Multi-LLM Inference.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Eagle: Efficient Training-Free Router for Multi-LLM Inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.547078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.547078Z digest=sha256:04b24ce5611828b7958add10edba7699799722093f14ccc700283155c58845ff

Observation f118e507-38d5-4f06-9c12-fb9cc61baa17 · outbound

This paper cites RouterBench: A Benchmark for Multi-LLM Routing System.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts RouterBench: A Benchmark for Multi-LLM Routing System

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.551092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.551092Z digest=sha256:0f1dcda2f729e5654f929741d82827d2b07887ad9a5f9306e85639b0c2fcc612

Observation 00649949-c493-4f18-b74b-9b6853a6aac3 · outbound

This paper cites Reinforcement learning in dynamic task scheduling: A review,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Reinforcement learning in dynamic task scheduling: A review,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.962721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.554729Z digest=sha256:16f4da90946ec21952e39bd3562b658b49b111f03d97e455d651fd3642216a22

Observation 8ec0fde6-11b0-46de-aa00-7d53d3a016aa · outbound

This paper cites Collaborative learning-based scheduling for kubernetes-oriented edge-cloud network,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Collaborative learning-based scheduling for kubernetes-oriented edge-cloud network,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.953572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.558272Z digest=sha256:cd78323576f9e96248537264066c2f4366bf046341aa392af624c85e8949f0ac

Observation 8f2438b1-a10b-460d-9795-efc97d7fd20e · outbound

This paper cites Clipper: A low-latency online prediction serving system,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Clipper: A low-latency online prediction serving system,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.943669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.561384Z digest=sha256:256e4f9029c170d138d803804d37674b118c690175ed3b6ab4d307263b885805

Observation 4fc7aa00-d5ce-4947-bc46-6a2ca3ac062c · outbound

This paper cites Tapfinger: Task place- ment and fine-grained resource allocation for edge machine learning,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Tapfinger: Task place- ment and fine-grained resource allocation for edge machine learning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.933924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.564545Z digest=sha256:17ea34880a72332e511b1d7baec0e498250662621c289998b262d01ad9ebe1aa

Observation 805502f2-61a6-4900-91c7-df0b40ef9c12 · outbound

This paper cites The non- stochastic multiarmed bandit problem,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts The non- stochastic multiarmed bandit problem,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.923358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.567752Z digest=sha256:f335dbbef9b301588e6c78e5f2c467c6ca07035e664eaa3690fe77d71c80d424

Observation aeecd7cb-b573-4c6d-86db-3b8f5be55375 · outbound

This paper cites Alpaca: A strong, replicable instruction- following model,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Alpaca: A strong, replicable instruction- following model,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.913571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.570832Z digest=sha256:b8be7689e5c7dbd097233b0b63890cb0c5521e8bad329c2611226b4a307bb0eb

Observation 5ee694a6-7663-4955-b1ba-ba5df4dc6475 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.573898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.573898Z digest=sha256:c05f271f58974db6e46f2c6b03a5291b82b96fe52abcf9334f0c4e1a673c3450

Observation e6446bf0-638f-4d86-8d3b-14769f76bdcb · outbound

This paper cites Introducing Mpt-7b: A new standard for open-source, commercially usable LLMs, 2023,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Introducing Mpt-7b: A new standard for open-source, commercially usable LLMs, 2023,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.903359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.577353Z digest=sha256:7786f9bdf774a34f9b63dcf0c40e7617a00f39090e454539b71d23c059c59565

Observation ccc560dd-3b4a-4490-ae35-6b60998217b6 · outbound

This paper cites BERTScore: Evaluating text generation with BERT,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts BERTScore: Evaluating text generation with BERT,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.893031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.580526Z digest=sha256:7fd482d3ad935eccf7f186f3c6f40bc048a6d22697f3678f129e2b566d8ba73d

Observation 06130ea9-117f-42d2-958d-ebcc97f7adf5 · outbound

This paper cites Soft Actor-Critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Soft Actor-Critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.882580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.583787Z digest=sha256:3cdc8954f99c19b5e174d979b6c2cb4cf502e212d6cb560d2fdc09b68f69bc7b

Observation 218b2815-e419-46b1-8d47-d6dd83cffc48 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.586945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.586945Z digest=sha256:b6b26fd34c09013cd17b53024475926d0cbf114626df00692a9acd6ddd46c8e9

Observation 349b3954-1ab8-4e79-93c3-68053e355c59 · outbound

This paper cites PyTorch: An im- perative style, high-performance deep learning library,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts PyTorch: An im- perative style, high-performance deep learning library,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.872062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.590788Z digest=sha256:7159b6887af72b8238dd655d530c1543858717bac2d477b69ef39d4ed83bdab1

Observation c3e0348b-b300-4653-b8be-43de4843abe8 · outbound

This paper cites TorchRL: A data-driven decision-making library for PyTorch.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts TorchRL: A data-driven decision-making library for PyTorch

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.593937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.593937Z digest=sha256:1149f9e01a7df180238068641c1d6e836b19fda4772dc16b6722db2b57116c50

Observation 744176ad-faf2-4bd9-8275-e68d4e776812 · outbound

This paper cites Fast Graph Representation Learning with PyTorch Geometric.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Fast Graph Representation Learning with PyTorch Geometric

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.597363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.597363Z digest=sha256:6d0962d04e76139d184d81f9476c448776e4b9c34214782253ff3f563ddf275d

Observation 9f7890c3-0601-4b7c-935e-fca2d8f40366 · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language understanding,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts BERT: Pre- training of deep bidirectional transformers for language understanding,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.860805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:22:42.600721Z digest=sha256:55ddce0c9ce67fd8d440efe9095ee03718ddb86bb151355f5029ecff2fcadc12

Pith citing papers

No inbound Pith citation observations are available.