REVIEW 3 cited by
Clipper: A Low-Latency Online Prediction Serving System
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Machine learning is being deployed in a growing number of applications which demand real-time, accurate, and robust predictions under heavy query load. However, most machine learning frameworks and systems only address model training and not deployment. In this paper, we introduce Clipper, a general-purpose low-latency prediction serving system. Interposing between end-user applications and a wide range of machine learning frameworks, Clipper introduces a modular architecture to simplify model deployment across frameworks and applications. Furthermore, by introducing caching, batching, and adaptive model selection techniques, Clipper reduces prediction latency and improves prediction throughput, accuracy, and robustness without modifying the underlying machine learning frameworks. We evaluate Clipper on four common machine learning benchmark datasets and demonstrate its ability to meet the latency, accuracy, and throughput demands of online serving applications. Finally, we compare Clipper to the TensorFlow Serving system and demonstrate that we are able to achieve comparable throughput and latency while enabling model composition and online learning to improve accuracy and render more robust predictions.
Forward citations
Cited by 3 Pith papers
-
Cruise Control: Dynamic Model Selection for ML-Based Network Traffic Analysis
A DPDK-based system dynamically swaps ML models and feature sets for network traffic analysis, using NIC packet-loss counters as a lightweight overload signal, and reports lower loss and comparable or higher median ac...
-
GPUs, CPUs, and... NICs: Rethinking the Network's Role in Serving Complex AI Pipelines
The paper proposes offloading AI pipeline data processing tasks to SmartNICs and sketches designs for normalization, bilinear interpolation, and tokenization, without implementing them.
-
Multi-Layer Perceptron-Based Relay Node Selection for Next-Generation Intelligent Delay-Tolerant Networks
An MLP-enhanced Spray and Wait router improves simulated DTN delivery by 7-8%, but the evaluation uses non-causal features and a self-referential label.
Discussion (0). Continue with ORCID to comment.