Pith. sign in

REVIEW 1 cited by

Deep Learning Workload Scheduling in GPU Datacenters: Taxonomy, Challenges and Vision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.11913 v3 pith:TLDF6VSK submitted 2022-05-24 cs.DC cs.AIcs.LG

classification cs.DCcs.AIcs.LG
keywords workloadsdatacenterdatacentersdeepexistinglearningresearchresource
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning (DL) shows its prosperity in a wide variety of fields. The development of a DL model is a time-consuming and resource-intensive procedure. Hence, dedicated GPU accelerators have been collectively constructed into a GPU datacenter. An efficient scheduler design for such GPU datacenter is crucially important to reduce the operational cost and improve resource utilization. However, traditional approaches designed for big data or high performance computing workloads can not support DL workloads to fully utilize the GPU resources. Recently, substantial schedulers are proposed to tailor for DL workloads in GPU datacenters. This paper surveys existing research efforts for both training and inference workloads. We primarily present how existing schedulers facilitate the respective workloads from the scheduling objectives and resource consumption features. Finally, we prospect several promising future research directions. More detailed summary with the surveyed paper and code links can be found at our project website: https://github.com/S-Lab-System-Group/Awesome-DL-Scheduling-Papers

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling

    cs.DC 2025-07 reject novelty 4.0 of 10

    A PPO-based autoscaler for GPU inference in Kubernetes is claimed to cut P95 latency up to 6.7x, but the evidence is weakened by a missing HPA baseline, a spike-traffic slowdown, and a reliance on synthetic feedback.

Pith tools