LA-IMR reduces P99 inference latency by up to 20.7% versus latency-only autoscaling by combining a fitted power-law latency model with proactive Kubernetes autoscaling and edge-to-cloud offloading.
GrandSLAm: Guaranteeing SLAs for Jobs in Microservices Execution Frameworks,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
LA-IMR reduces P99 inference latency by up to 20.7% versus latency-only autoscaling by combining a fitted power-law latency model with proactive Kubernetes autoscaling and edge-to-cloud offloading.