Pith. sign in

REVIEW 1 cited by

Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.07848 v2 pith:QVNGSZMF submitted 2022-02-16 cs.DC cs.AI

classification cs.DCcs.AI
keywords singularityworkloadsdeeplearningacceleratorsacrossapproachdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Lowering costs by driving high utilization across deep learning workloads is a crucial lever for cloud providers. We present Singularity, Microsoft's globally distributed scheduling service for highly-efficient and reliable execution of deep learning training and inference workloads. At the heart of Singularity is a novel, workload-aware scheduler that can transparently preempt and elastically scale deep learning workloads to drive high utilization without impacting their correctness or performance, across a global fleet of AI accelerators (e.g., GPUs, FPGAs). All jobs in Singularity are preemptable, migratable, and dynamically resizable (elastic) by default: a live job can be dynamically and transparently (a) preempted and migrated to a different set of nodes, cluster, data center or a region and resumed exactly from the point where the execution was preempted, and (b) resized (i.e., elastically scaled-up/down) on a varying set of accelerators of a given type. Our mechanisms are transparent in that they do not require the user to make any changes to their code or require using any custom libraries that may limit flexibility. Additionally, our approach significantly improves the reliability of deep learning workloads. We show that the resulting efficiency and reliability gains with Singularity are achieved with negligible impact on the steady-state performance. Finally, our design approach is agnostic of DNN architectures and handles a variety of parallelism strategies (e.g., data/pipeline/model parallelism).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Green Distributed AI Training: Orchestrating Compute Across Renewable-Powered Micro Datacenters

    cs.NI 2025-11 conditional novelty 4.0 of 10

    Migratory AI training across renewable-powered micro-datacenters is limited by checkpoint transfer time rather than energy cost, and a feasibility-filtered orchestrator reports ~52% non-renewable energy reduction with...

Pith tools