REVIEW 2 minor 158 cited by
Training Deep Nets with Sublinear Memory Cost
T0 review · 0 major / 2 minor · reviewed 2026-05-12 · grok-4.3
Pith's one-line read An algorithm trains an n-layer deep network using O(sqrt(n)) memory at the cost of one extra forward pass.
desk verdict Key takeaway: O(sqrt(n)) memory for deep net training via graph segmentation and recomputation, with solid experiments. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The checkpointing strategy that segments the computation graph into sqrt(n) intervals, storing activations only at boundaries and recomputing forwards inside each interval during backpropagation.
What would settle it
Running the algorithm on a 1000-layer residual network and measuring whether peak memory usage scales as O(sqrt(n)), total runtime increases by about 30 percent, and the resulting gradients match those from full-storage training.
Extended reading notes
Core claim
We design an algorithm that costs O(sqrt(n)) memory to train a n layer network, with only the computational cost of an extra forward pass per mini-batch. As many of the state-of-the-art models hit the upper bound of the GPU memory, our algorithm allows deeper and more complex models to be explored. We focus on reducing the memory cost to store the intermediate feature maps and gradients during training. Computation graph analysis is used for automatic in-place operation and memory sharing optimizations. We show that it is possible to trade computation for memory - giving a more memory efficient training algorithm with a little extra computation cost. In the extreme case, our analysis also 7G
Load-bearing premise
The computation graph can be cleanly segmented into sqrt(n) intervals where recomputing forward passes inside each interval is both correct and cheaper than storing all intermediate activations.
Editorial extensions
If this is right
- A 1000-layer residual network trains with memory reduced from 48G to 7G and only 30 percent extra running time on ImageNet.
- Complex recurrent neural networks become trainable on very long sequences with substantially lower memory.
- State-of-the-art models no longer hit GPU memory limits as quickly, enabling exploration of deeper architectures.
- An extreme variant reduces memory to O(log n) at the cost of O(n log n) extra forward computation.
Reading between the lines
- The approach could lower hardware barriers for training large models and make advanced deep learning more accessible on modest GPUs.
- Adaptive checkpoint intervals based on per-layer compute cost might improve the compute-memory trade-off further.
- The method pairs naturally with model parallelism to scale to even larger networks without changing the core algorithm.
- Systems with high compute throughput relative to memory bandwidth would see the smallest effective overhead from the extra forward passes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents an algorithm to train deep neural networks with O(sqrt(n)) memory cost for an n-layer network, incurring only the cost of one extra forward pass per mini-batch. This is achieved through computation graph analysis, segmenting the network into intervals, storing boundary activations, and recomputing forward passes within segments during backpropagation. The approach is extended to O(log n) memory with O(n log n) extra computation, and validated on ImageNet with a 1000-layer ResNet (48G to 7G memory) and long-sequence RNNs.
Significance. If the claims hold, this is a significant contribution to deep learning training efficiency, allowing exploration of deeper models on memory-constrained hardware like GPUs. The systematic use of DAG properties for memory optimization, combined with empirical validation showing memory reduction with modest time overhead and correct gradients, provides a practical tool for advancing DL research. The parameter-free derivation from standard graph segmentation is a strength.
minor comments (2)
- [Abstract] Abstract: the O(sqrt(n)) claim would be clearer if it explicitly stated the segmentation assumption (clean intervals where recomputation is correct and cheaper than storing all activations) that underpins the bound.
- [Experiments] The 30% extra time cost for the 1000-layer ResNet is reported, but a per-component breakdown (recomputation vs. original forward/backward) would make the compute-memory trade-off more transparent.
Simulated Author's Rebuttal
We thank the referee for the positive and accurate summary of our work, the assessment of its significance, and the recommendation to accept the manuscript. No major comments requiring response or revision were raised.
read point-by-point responses
-
Referee: No specific major comments were listed in the report.
Authors: We appreciate the referee's recognition that the algorithm provides a systematic, parameter-free approach to memory reduction via graph segmentation and recomputation, with empirical validation on large models. The description of the O(sqrt(n)) memory bound, the O(log n) extension, and the ImageNet/ResNet and RNN experiments matches our claims exactly. revision: no
Circularity Check
No significant circularity; derivation is self-contained
full rationale
The O(sqrt(n)) memory bound is obtained by partitioning the n-layer computation DAG into sqrt(n) segments, retaining only the sqrt(n) boundary activations, and performing one recomputation of each segment during back-propagation; the total extra work equals one forward pass by direct operation counting on the graph. This counting argument relies only on standard properties of feed-forward and recurrent DAGs plus the in-place/memory-sharing optimizations described in the paper; no parameters are fitted to data, no result is defined in terms of itself, and no load-bearing step reduces to a self-citation. The reported ImageNet and RNN experiments serve as empirical confirmation rather than definitional inputs.
Assumptions & free parameters
assumptions (1)
- standard math The forward computation graph is a directed acyclic graph whose nodes correspond to layer activations.
Cite this review
Pith. "Pith review of Training Deep Nets with Sublinear Memory Cost." pith.science (2026). https://pith.science/paper/OPXDY4I7
@misc{pith2026160406174,
author = {Pith},
title = {Pith review of: Training Deep Nets with Sublinear Memory Cost},
year = {2026},
howpublished = {\url{https://pith.science/paper/OPXDY4I7}},
note = {Machine review of arXiv:1604.06174}
}
read the original abstract
We propose a systematic approach to reduce the memory consumption of deep neural network training. Specifically, we design an algorithm that costs O(sqrt(n)) memory to train a n layer network, with only the computational cost of an extra forward pass per mini-batch. As many of the state-of-the-art models hit the upper bound of the GPU memory, our algorithm allows deeper and more complex models to be explored, and helps advance the innovations in deep learning research. We focus on reducing the memory cost to store the intermediate feature maps and gradients during training. Computation graph analysis is used for automatic in-place operation and memory sharing optimizations. We show that it is possible to trade computation for memory - giving a more memory efficient training algorithm with a little extra computation cost. In the extreme case, our analysis also shows that the memory consumption can be reduced to O(log n) with as little as O(n log n) extra cost for forward computation. Our experiments show that we can reduce the memory cost of a 1,000-layer deep residual network from 48G to 7G with only 30 percent additional running time cost on ImageNet problems. Similarly, significant memory cost reduction is observed in training complex recurrent neural networks on very long sequences.
Forward citations
Showing 60 of 158 Pith papers that cite this
-
UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing
UltraEP is the first exact-load real-time expert balancer for large-EP MoE training and serving on rack-scale nodes, reaching 94.3% of ideal throughput and 1.49x over no-balancing.
-
Systematic Discovery of Semantic Attacks in Online Map Construction through Conditional Diffusion
MIRAGE discovers semantic attacks on online HD map construction via conditional diffusion, enabling boundary removal and injection that degrade AV performance while passing as realistic environmental changes.
-
Differentiate the Solver, Not the Equation: Reverse-Sweep Adjoints for Block Implicit Simulation
A reverse sweep of local 3x3 adjoint solves gives machine-precision gradients through the exact finite-depth Vertex Block Descent solver without forming any global system.
-
D2PO: Optimizing Diffusion Samplers via Dynamic Preference
D2PO learns better low-NFE diffusion timestep schedules and CFG weights via DPO on a score-based energy with a dynamic denser-schedule preference target.
-
A matrix-free, differentiable PyTorch solver for phase-field fracture: Formulation, benchmarks, and inverse analysis
A matrix-free, GPU-compatible PyTorch implementation of phase-field fracture with explicit dynamics, custom differentiable implicit damage solve, benchmarks on dynamic and quasi-static cases, and inverse recovery of f...
-
Diffusion-Driven State Space Models
DDSSM replaces Gaussian transitions in SSMs with diffusion models to jointly train autoencoders and diffusion on sequential data, outperforming standard deep SSMs on simulated multimodal time series.
-
ITNet: A Learnable Integral Transform That Subsumes Convolution, Attention, and Recurrence
ITNet unifies convolution, self-attention, and autoregressive recurrence as special cases of a single learnable integral transform implemented via an MLP kernel and matches specialized models on ImageNet, GLUE, ModelN...
-
RATrain: A Resource-Aware Training Runtime for Large Language Models on Bandwidth-Constrained Heterogeneous Supercomputing Platforms
RATrain introduces a resource-aware scheduler and MT-3000-specific backend for 1F1B LLM training that achieves 1.35x speedup and 97% scaling efficiency while preserving training correctness.
-
Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking
PCM uses success-failure action variance to probabilistically select and mask chunks for gradient updates in GRPO, matching standard success rates with 2.38x wall-clock speedup and 60% lower memory on LIBERO benchmarks.
-
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
An efficiently computable HS-Jacobian acts as a conservative mapping for projections onto polyhedral sets, supporting provably convergent Adam-based end-to-end training of linearly constrained deep neural networks.
-
Locking Pretrained Weights via Deep Low-Rank Residual Distillation
DLR-Lock locks open-weight LLMs against unauthorized fine-tuning by swapping MLPs for deep low-rank residual networks that inflate backprop memory and complicate optimization, yet preserve original capabilities via mo...
-
Finite Volume-Informed Neural Network Framework for 2D Shallow Water Equations: Rugged Loss Landscapes and the Importance of Data Guidance
Data-guided finite-volume PINNs for 2D shallow water equations avoid trivial low-momentum collapse via sparse measurements, achieving up to 22x error reduction on benchmarks and accurate surrogates on real river data.
-
FlatLands: Generative Floormap Completion From a Single Egocentric View
A new multi-source real indoor benchmark shows conditional generative models outperform deterministic and ensemble baselines at single-view BEV floor completion, with uncertainty concentrated at layout boundaries.
-
4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
4D-LRM is a transformer that maps sparse posed frames scattered across time to a cloud of 4D Gaussians and renders any query view at any query time in under 1.5 seconds.
-
DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
A single feed-forward transformer predicts per-pixel deformable 3D Gaussians with dense scene flow from a posed monocular video, enabling real-time dynamic view synthesis and 3D tracking.
-
Training-Free Inference for High-Resolution Sinogram Completion
HRSino is a training-free inference scheme that completes high-resolution sinograms with adaptive patch skipping and step allocation, cutting memory and time without losing accuracy.
-
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
A slice-level pipeline-parallel schedule with attention-work redistribution that cuts activation memory roughly by the pipeline size and reduces pipeline bubbles for long-context LLM training.
-
GME: Improving Universal Multimodal Retrieval by Multimodal LLMs
GME achieves state-of-the-art results in universal multimodal retrieval by training on a balanced synthetic multimodal dataset.
-
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
GaLore performs full-parameter LLM training with up to 65.5% less optimizer memory by projecting gradients onto a low-rank subspace at each step, matching full-rank performance on LLaMA pre-training and RoBERTa fine-tuning.
-
Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations
HSTU-based generative recommenders with 1.5 trillion parameters scale as a power law with compute up to GPT-3 scale, outperform baselines by up to 65.8% NDCG, run 5-15x faster than FlashAttention2 on long sequences, a...
-
Moonwalk: Inverse-Forward Differentiation
Moonwalk enables memory-efficient training of deep networks via mixed-mode gradient computation with vector-inverse-Jacobian products for submersive layers and fragmental checkpointing otherwise, matching backprop run...
-
Ring Attention with Blockwise Transformers for Near-Infinite Context
Ring Attention uses blockwise computation and ring communication to let Transformers process sequences up to device-count times longer than prior memory-efficient methods.
-
Efficient Memory Management for Large Language Model Serving with PagedAttention
PagedAttention achieves near-zero waste in LLM key-value cache memory and enables 2-4x higher serving throughput than prior systems.
-
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
LLM.int8() performs 8-bit inference for transformers up to 175B parameters with no accuracy loss by combining vector-wise quantization for most features with 16-bit mixed-precision handling of systematic outlier dimensions.
-
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
FlashAttention reduces GPU high-bandwidth memory accesses in self-attention via tiling, delivering exact attention with lower IO complexity, 2-3x wall-clock speedups on models like GPT-2, and the ability to train on s...
-
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
ZeRO removes memory redundancies in parallel training to scale deep learning models to over a trillion parameters with high throughput on current hardware.
-
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
ALBERT reduces BERT parameters via embedding factorization and layer sharing, adds inter-sentence coherence pretraining, and reaches SOTA on GLUE, RACE, and SQuAD with fewer parameters than BERT-large.
-
On the Acceleration of Deep Learning Model Parallelism with Staleness
DSP decouples forward and backward passes in model-parallel deep CNN training by giving each layer block a preset staleness, yielding speedups up to 4.8x with comparable or better accuracy.
-
ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization
Reconstruction-guided chunking creates adaptive latent tokens that beat random boundaries on native reconstruction, but downstream readers recover topic information while losing exact lexical details.
-
ElastiCo: Elastic Configuration and Interference-Aware Orchestration for GPU Clusters
A Kubernetes-native scheduler jointly reshapes job configurations, prices cluster resources, and predicts GPU-sharing interference, reporting up to 2.94x lower average JCT and 2.02x higher throughput in testbed and si...
-
LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
LoCA tunes adapters with one calibrated map per layer and closed-form ridge solves, avoiding repeated backpropagation after calibration and roughly matching LoRA quality on tested tasks.
-
TextNCA: Neural Cellular Automata for Language Modeling via Hierarchical Local Attention
A hierarchical local-attention NCA for language modeling shows that the narrow-to-wide window schedule explains most of its behavior, with iteration adding a small bounded benefit, and the model underperforms transformers.
-
Backpropagation-Free Trunk Training via the Split Forward Gradients
Split-FG splits a network into an exactly trained head and a forward-gradient-estimated trunk, reducing variance and reaching 387 perplexity on WikiText-103 with a 16M GPT-2-style model.
-
AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling
An asymmetric geometry-appearance architecture for generalizable 3DGS reallocates computation so smaller models match optimization-based NVS quality at ~800× speedup on 32-view 960P inputs while improving zero-shot results.
-
Architecture Generalization with MetaNCA
A learned local rule (Weight Transformer) iteratively self-organizes task-network weights from local graph neighborhoods and generalizes across unseen MLP, CNN, and ResNet architectures up to ~2M parameters.
-
Weave of Formal Thought
Weave of Formal Thought presents a complete constrained decoder via speculative-lexing GLR and an RWS latent fine-tuning method that reduces per-token cross-entropy by 14.3% on StarCoder2-3B for Python relative to tex...
-
FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs
FoMoE partitions expert layers across workers in MoE LLMs, skips non-resident experts, and reports up to 1.42x lower communication than baselines plus 1.4x throughput gains while maintaining stable routing.
-
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
Qwen-RobotWorld is a language-conditioned video world model using Double-Stream MMDiT, an 8.6M-frame embodied corpus, and progressive curriculum training that ranks first on EWMBench and DreamGen Bench.
-
CD-RCM: Generalizable Continuous-Depth Novel View Synthesis for Reflectance Confocal Microscopy
CD-RCM is a feedforward neural model for novel view synthesis that predicts unseen depths in reflectance confocal microscopy stacks to produce isotropic 3D volumes for arbitrary sectioning.
-
Policy-based Foveated Imaging and Perception
A task-aware policy learned via reinforcement learning allocates high-resolution pixels on dual-stream sensors in real time, outperforming fixed or non-predictive baselines under tight pixel budgets in both simulation...
-
Schedule-Level Shared-Prefix Reuse for LLM RL Training
Schedule-level shared-prefix reuse decouples prefix and suffix passes in GRPO training to compute shared prefixes once, delivering up to 4.395x speedup and 59.1% HBM reduction while preserving numerical equivalence.
-
StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement
StressDream optimizes initial noise in diffusion video world models using VLM semantic and plausibility objectives to steer generations toward specified high-impact outcomes for improved policy evaluation.
-
ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning
ChunkFT enables full-parameter fine-tuning of Llama 3-8B on one 24 GB GPU and Llama 3-70B on two 80 GB GPUs by streaming gradients over dynamically activated sub-tensors.
-
Towards Understanding Self-Pretraining for Sequence Classification
Self-pretraining improves Transformer sequence classification by enabling learning of proximity-biased attention from positional encodings that label supervision alone cannot easily acquire from random starts.
-
STELLAR: Scaling 3D Perception Large Models for Autonomous Driving
STELLAR trains up to 500M-parameter multi-modal models on 50M driving scenes and reports empirical scaling trends plus new state-of-the-art results on the Waymo Open Dataset.
-
Njord: A Probabilistic Graph Neural Network for Ensemble Ocean Forecasting
Njord is a probabilistic GNN model using latent variables and adaptive K-means meshes that produces ensemble forecasts and outperforms deterministic ML baselines on global OceanBench and Baltic Sea domains.
-
AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs
AGoQ delivers up to 52% lower memory use and 1.34x faster training for 8B-32B LLaMA models by using near-4-bit adaptive activations and 8-bit gradients while preserving pretraining convergence and downstream accuracy.
-
MCMit: Hardware-Software Co-Design for Mid-Circuit Measurement Error Mitigation
MCMit proposes a constant-latency multi-control branch instruction, transformer and CNN discriminators, plus static MCM elimination and stochastic branching, evaluated on Qubic with QPU traces to cut latency by 70% an...
-
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
SIEVES improves selective prediction coverage up to 3x on OOD VQA benchmarks by training a selector on visual localization quality, generalizing across datasets and proprietary reasoners without specific adaptation.
-
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
SlideFormer uses layer-sliding async offloading, pre-allocated heterogeneous memory, and fused Triton kernels to fine-tune 123B+ models on one RTX 4090 with 1.4–6.3× higher throughput and roughly half the memory of pr...
-
GeoPT: Scaling Physics Simulation via Lifted Geometric Pre-Training
GeoPT pre-trains on over one million geometry samples augmented with synthetic dynamics to improve neural physics simulators on fluid and solid mechanics benchmarks while reducing labeled data needs by 20-60% and acce...
-
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
A new benchmark and training strategy show LLMs trained to internalize causal reasoning steps are less fooled by semantically similar, label-flipped questions than models using explicit chain-of-thought.
-
Solving Inverse Problems with Flow-based Models via Model Predictive Control
MPC-Flow applies model predictive control to guide pretrained flow models through inverse problems, with a single-step variant that avoids backpropagation and scales to 32B-parameter models on consumer hardware.
-
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
PhyGDPO uses groupwise direct preference optimization with real videos as winners to make text-to-video models generate more physically plausible videos.
-
Multi-view Pyramid Transformer: Look Coarser to See Broader
MVP uses a two-level hierarchy of attention windows and token resolutions to reconstruct large 3D scenes from up to 256 input views in a single feed-forward pass, beating Long-LRM and iLRM on DL3DV and several zero-sh...
-
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
MTraining scales LLM training to 512K-token contexts on 32 A100 GPUs by integrating dynamic sparse training patterns with balanced and hierarchical sparse ring attention, achieving up to 6x throughput gains without ac...
-
OctoPipe: Reducing Pipeline Bubbles for Heterogeneous Models via Co-Optimizing Partitioning, Placement, and Scheduling
Co-optimizing model partition, placement, and workload scheduling for pipeline-parallel LLM training is claimed to improve throughput by 1.15 to 1.44x (abstract) or up to 2.14x (body).
-
CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
CR-Net uses cross-layer low-rank residuals in a dual-path network plus specialized recomputation to outperform prior low-rank methods on 60M-7B model pre-training while using less compute and memory.
-
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
MetaEmbed trains fixed learnable Meta Tokens to produce granularity-organized multi-vector embeddings that support test-time scaling in multimodal retrieval.
-
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences
Chameleon is a swap-based memory optimizer that handles changing operator sequences in eager-mode LLM training, enabling models up to 4x larger than device memory.
Reference graph
Works this paper leans on
-
[1]
Mart ´ın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Good- fellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Man ´e, Rajat Monga, Sherry Moore, Derek Murra...
work page 2015
-
[2]
Amit Agarwal, Eldar Akchurin, Chris Basoglu, Guoguo Chen, Scott Cyphers, Jasha Droppo, Adam Eversole, Brian Guenter, Mark Hillebrand, Ryan Hoens, Xuedong Huang, Zhiheng Huang, Vladimir Ivanov, Alexey Kamenev, Philipp Kranen, Oleksii Kuchaiev, Wolfgang Manousek, Avner May, Bhaskar Mitra, Olivier Nano, Gaizka Navarro, Alexey Orlov, Marko Padmilac, Hari Part...
work page 2014
-
[3]
Aho, Ravi Sethi, and Jeffrey D
Alfred V . Aho, Ravi Sethi, and Jeffrey D. Ullman. Compilers: Principles, Techniques, and Tools. Addison-Wesley Longman Publishing Co., Inc., Boston, MA, USA, 1986
work page 1986
-
[4]
Goodfellow, Arnaud Bergeron, Nicolas Bouchard, and Yoshua Bengio
Fr ´ed´eric Bastien, Pascal Lamblin, Razvan Pascanu, James Bergstra, Ian J. Goodfellow, Arnaud Bergeron, Nicolas Bouchard, and Yoshua Bengio. Theano: new features and speed improve- ments. Deep Learning and Unsupervised Feature Learning NIPS 2012 Workshop, 2012
work page 2012
-
[5]
Theano: a CPU and GPU math expression compiler
James Bergstra, Olivier Breuleux, Fr ´ed´eric Bastien, Pascal Lamblin, Razvan Pascanu, Guil- laume Desjardins, Joseph Turian, David Warde-Farley, and Yoshua Bengio. Theano: a CPU and GPU math expression compiler. In Proceedings of the Python for Scientific Computing Conference (SciPy), June 2010. Oral Presentation
work page 2010
-
[6]
MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, , and Zheng Zhang. MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems. In Neural Information Processing Systems, Workshop on Machine Learning Systems (LearningSys’15), 2015
work page 2015
-
[7]
Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V
Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V . Le, Mark Z. Mao, MarcAurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y . Ng. Large scale distributed deep networks. In NIPS, 2012
work page 2012
-
[8]
Ian Goodfellow, Yoshua Bengio, , and Aaron Courville. Deep learning. Book in preparation for MIT Press, 2016
work page 2016
Show all 19 references
-
[9]
Algorithm 799: Revolve: An implementation of checkpointing for the reverse or adjoint mode of computational differentiation
Andreas Griewank and Andrea Walther. Algorithm 799: Revolve: An implementation of checkpointing for the reverse or adjoint mode of computational differentiation. ACM Trans. Math. Softw., 26(1):19–45, March 2000
2000
-
[10]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. arXiv preprint arXiv:1512.03385, 2015
2015 arXiv
-
[11]
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. arXiv preprint arXiv:1603.05027, 2016
2016
-
[12]
Long short-term memory
Sepp Hochreiter and J ¨urgen Schmidhuber. Long short-term memory. Neural Comput. , 9(8):1735–1780, November 1997. 11
1997
-
[13]
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the 32th International Conference on Machine Learning (ICML’15), 2015
2015
-
[14]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems 25 , pages 1097–1105. 2012
2012
-
[15]
Gradient-based learning applied to document recognition
Yann LeCun, L ´eon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. In S. Haykin and B. Kosko, editors, Intelligent Signal Pro- cessing, pages 306–351. IEEE Press, 2001
2001
-
[16]
Virtualizing deep neural networks for memory-efficient neural network design.arXiv preprint arXiv:1602.08124, 2016
Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W Keckler. Virtualizing deep neural networks for memory-efficient neural network design.arXiv preprint arXiv:1602.08124, 2016
2016
-
[17]
Senior, and Franc ¸oise Beaufays
Hasim Sak, Andrew W. Senior, and Franc ¸oise Beaufays. Long short-term memory recur- rent neural network architectures for large scale acoustic modeling. In INTERSPEECH 2014, 15th Annual Conference of the International Speech Communication Association, Singapore, September 14-...
2014
-
[18]
Training very deep networks
Rupesh Kumar Srivastava, Klaus Greff, and J¨urgen Schmidhuber. Training very deep networks. arXiv preprint arXiv:1507.06228, 2015
2015
-
[19]
Highway long short-term memory rnns for distant speech recognition
Yu Zhang, Guoguo Chen, Dong Yu, Kaisheng Yao, Sanjeev Khudanpur, and James Glass. Highway long short-term memory rnns for distant speech recognition. arXiv preprint arXiv:1510.08983, 2015. A Search over Budget B Alg. 3 allows us to generate an optimized memory plan given a sin...
2015
Reviewed May 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.