Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T10:22:42.600721Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2508.00234.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T10:22:42.600721Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a7efcac1-e101-4a8f-aee0-b89606f8e866 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts AIoT smart home via autonomous LLM agents,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5becd7f5-7a18-4a0b-8b99-0a7ff6399765 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Large language models for human- ai co-creation of robotic dance performances,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 66562e3f-b0ec-43cb-95a3-6d327ed826de · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts EdgeFM: Leveraging foundation model for open-set learning on the edge,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa1a07e2-b807-4ec9-afdd-780f71e7630e · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts WDMoE: Wireless Distributed Large Language Models with Mixture of Experts
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e29751e0-7f8d-4110-a95a-ee4b27e27192 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts On Protecting the Data Privacy of Large Language Models (LLMs): A Survey
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94d46d84-12c1-4b01-9fc2-9714dd1b58c2 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Edge intelligence: Paving the last mile of artificial intelligence with edge computing,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 057fe7f9-4c30-4db0-bcf3-867bc48a5d07 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Enabling AI-Generated Content (AIGC) Services in Wireless Edge Networks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 178e19ba-07ac-4d24-ae51-acba3c2955d2 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Toward Scalable Generative AI via Mixture of Experts in Mobile Edge Networks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a15f1f24-6d08-4389-8ebc-a4dcd44239c4 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts LLM-Blender: Ensembling large language models with pairwise ranking and generative fusion,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c1324dbe-b74d-4485-a844-e9374342ebe8 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5851814e-fe47-4e49-9f96-335532db538d · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Orca: A distributed serving system for Transformer-based generative models,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f57748b3-8b08-4739-8117-89915ee233ad · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Efficient memory management for large language model serving with PagedAttention,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d6fb5336-512d-48cd-8483-794d817d97f1 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3c4b001-6546-45ba-b1e2-86507aa5b3cd · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d215f12-f43b-4d2c-be98-bc4d39c890e6 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e1a3c60-1439-493e-8f61-0320821726da · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts RouteLLM: Learning to Route LLMs with Preference Data
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f02fcc-0387-4b39-9431-998ed626a0a2 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f8399dc-ff72-4234-bc1e-d302dff8fba4 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5feb3bae-8f4e-4fc5-a454-b3d2bd10255d · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts S3: Increasing gpu utiliza- tion during generative inference for higher throughput,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ff93db18-1f56-45f4-af19-a093d6a6f59e · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts FlexGen: High-throughput generative inference of large language models with a single gpu,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 31825be3-8091-4d3c-942e-620e5bc4d5ab · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts FlashAttention: Fast and memory-efficient exact attention with io-awareness,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8e83768f-be94-4e1b-9b39-1bfced14ec38 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts HeteGen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6a1aaa9b-af2a-45c9-952d-ed89d8481d74 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts ExeGPT: Constraint-aware resource scheduling for LLM inference,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 120312c5-8335-48a5-8248-6cb87de67a0f · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed792d59-abd2-4839-bdd1-2093adadd75e · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Splitwise: Efficient generative LLM inference using phase splitting,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99811022-c3ff-47f0-adb3-784c8973e78f · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29ca24c1-7612-4688-bc65-4af530a53ced · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 972eceb0-f50d-4179-990e-02817071252e · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Large Language Model Routing with Benchmark Datasets
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 907191d6-53e4-470b-a88a-964cba56d57b · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Octopus v4: Graph of language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 959d65b7-5550-4a95-9b9c-bbbd9e40ae62 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts GraphRouter: A Graph-based Router for LLM Selections
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7900a0-9cbd-40a5-9bfd-a4e84da2a1da · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Eagle: Efficient Training-Free Router for Multi-LLM Inference
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f118e507-38d5-4f06-9c12-fb9cc61baa17 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts RouterBench: A Benchmark for Multi-LLM Routing System
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00649949-c493-4f18-b74b-9b6853a6aac3 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Reinforcement learning in dynamic task scheduling: A review,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8ec0fde6-11b0-46de-aa00-7d53d3a016aa · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Collaborative learning-based scheduling for kubernetes-oriented edge-cloud network,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8f2438b1-a10b-460d-9795-efc97d7fd20e · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Clipper: A low-latency online prediction serving system,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4fc7aa00-d5ce-4947-bc46-6a2ca3ac062c · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Tapfinger: Task place- ment and fine-grained resource allocation for edge machine learning,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 805502f2-61a6-4900-91c7-df0b40ef9c12 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts The non- stochastic multiarmed bandit problem,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aeecd7cb-b573-4c6d-86db-3b8f5be55375 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Alpaca: A strong, replicable instruction- following model,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ee694a6-7663-4955-b1ba-ba5df4dc6475 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6446bf0-638f-4d86-8d3b-14769f76bdcb · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Introducing Mpt-7b: A new standard for open-source, commercially usable LLMs, 2023,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ccc560dd-3b4a-4490-ae35-6b60998217b6 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts BERTScore: Evaluating text generation with BERT,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 06130ea9-117f-42d2-958d-ebcc97f7adf5 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Soft Actor-Critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 218b2815-e419-46b1-8d47-d6dd83cffc48 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349b3954-1ab8-4e79-93c3-68053e355c59 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts PyTorch: An im- perative style, high-performance deep learning library,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c3e0348b-b300-4653-b8be-43de4843abe8 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts TorchRL: A data-driven decision-making library for PyTorch
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 744176ad-faf2-4bd9-8275-e68d4e776812 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Fast Graph Representation Learning with PyTorch Geometric
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f7890c3-0601-4b7c-935e-fca2d8f40366 · outbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts BERT: Pre- training of deep bidirectional transformers for language understanding,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.