Pith. sign in

REVIEW 2 cited by

Checkmate: Breaking the Memory Wall with Optimal Tensor Rematerialization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.02653 v3 pith:L2ARMLN3 submitted 2019-10-07 cs.LG cs.CVcs.DCstat.ML

classification cs.LGcs.CVcs.DCstat.ML
keywords checkmaterematerializationschedulestrainingcostmemoryoptimalproblem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We formalize the problem of trading-off DNN training time and memory requirements as the tensor rematerialization optimization problem, a generalization of prior checkpointing strategies. We introduce Checkmate, a system that solves for optimal rematerialization schedules in reasonable times (under an hour) using off-the-shelf MILP solvers or near-optimal schedules with an approximation algorithm, then uses these schedules to accelerate millions of training iterations. Our method scales to complex, realistic architectures and is hardware-aware through the use of accelerator-specific, profile-based cost models. In addition to reducing training cost, Checkmate enables real-world networks to be trained with up to 5.1x larger input sizes. Checkmate is an open-source project, available at https://github.com/parasj/checkmate.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration

    cs.DC 2026-07 conditional novelty 6.0 of 10

    Trace-guided fine-grained memory control and offline joint planning raise diffusion serving SLO attainment by up to 3.7× while cutting configuration search from hours to minutes.

  2. DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific Computing

    cs.LG 2025-09 conditional novelty 6.0 of 10

    DaCe AD automatically differentiates scientific Python and Fortran code through an SDFG intermediate representation, beating JAX JIT across NPBench with a 4.1x geometric mean speedup.

Pith tools