WISP suppresses wasted drafting time and verification interference in edge-cloud speculative LLM serving through dynamic drafting and SLO-aware batching, delivering up to 2.1x capacity and 1.94x goodput gains over centralized and prior baselines.
Splitllm: Collaborative infer- ence of llms for model placement and throughput optimization
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 4roles
dataset 1polarities
use dataset 1representative citing papers
Green-LLM applies lexicographic multi-objective optimization to distribute LLM inference across renewable-powered heterogeneous data centers, jointly cutting carbon emissions and water consumption while keeping costs near minimum and latency under 2 seconds.
Enriched textual prompts lift local LLM accuracy on binary IoT environmental queries from 50.9-63.7% to 81.7-89.3% while preserving sub-second latency.
A survey synthesizing challenges, system architectures, model optimizations, deployment methods, and resource management techniques for large language model inference at the network edge.
citing papers explorer
-
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
WISP suppresses wasted drafting time and verification interference in edge-cloud speculative LLM serving through dynamic drafting and SLO-aware batching, delivering up to 2.1x capacity and 1.94x goodput gains over centralized and prior baselines.
-
Green-LLM: Optimal Workload Allocation for Environmentally-Aware Distributed Inference
Green-LLM applies lexicographic multi-objective optimization to distribute LLM inference across renewable-powered heterogeneous data centers, jointly cutting carbon emissions and water consumption while keeping costs near minimum and latency under 2 seconds.
-
Enabling Cloud-Level Accuracy in Edge AI through IoT Data Preprocessing
Enriched textual prompts lift local LLM accuracy on binary IoT environmental queries from 50.9-63.7% to 81.7-89.3% while preserving sub-second latency.
-
Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities
A survey synthesizing challenges, system architectures, model optimizations, deployment methods, and resource management techniques for large language model inference at the network edge.