Pith. sign in

REVIEW 12 cited by

The Computational Limits of Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.05558 v2 pith:DVYHNZXG submitted 2020-07-10 cs.LG stat.ML

classification cs.LGstat.ML
keywords learningprogressdeepapplicationscomecomputingmethodsother
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning's recent history has been one of achievement: from triumphing over humans in the game of Go to world-leading performance in image classification, voice recognition, translation, and other tasks. But this progress has come with a voracious appetite for computing power. This article catalogs the extent of this dependency, showing that progress across a wide variety of applications is strongly reliant on increases in computing power. Extrapolating forward this reliance reveals that progress along current lines is rapidly becoming economically, technically, and environmentally unsustainable. Thus, continued progress in these applications will require dramatically more computationally-efficient methods, which will either have to come from changes to deep learning or from moving to other machine learning methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Koopman Model Dimension Reduction via Variational Bayesian Inference and Graph Search

    eess.SY 2026-01 conditional novelty 6.0 of 10

    Variational Bayesian inclusion-flag estimates, thresholded into a directed graph, select a smaller Koopman dictionary while leaving output influence paths intact.

  2. Quantum optical neural networks using atom-cavity interactions to provide all-optical nonlinearity

    quant-ph 2025-11 conditional novelty 6.0 of 10

    A simulated neural network uses atom-cavity two-level neurons as all-optical nonlinear activations and reports ~95% accuracy on MNIST and SAT-6.

  3. Progressive Depth Up-scaling via Optimal Transport

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Optimal-transport alignment of adjacent layers gives a cheap and slightly better initialization for progressive depth up-scaling than copying, averaging, or a learned predictor.

  4. MOSAIC-FL, a micro-service based privacy-preserving framework with application to genomics

    cs.CR 2026-07 conditional novelty 5.0 of 10

    A gRPC micro-service FL stack with t-out-of-N CKKS secure aggregation matches cleartext accuracy on EMNIST and TCGA BRCA subtyping at modest extra cost for large models.

  5. Physical Analogue Kolmogorov-Arnold Networks based on Reconfigurable Nonlinear-Processing Units

    cs.ET 2026-02 conditional novelty 5.0 of 10

    A proposed analog KAN chip uses silicon RNPUs as physically programmable nonlinear edges, with estimated ~250 pJ per inference and ~10x smaller area than a digital MLP.

  6. Real-Time Analysis of Unstructured Data with Machine Learning on Heterogeneous Architectures

    physics.data-an 2025-08 conditional novelty 5.0 of 10

    A graph neural network (ETX4VELO) reconstructs LHCb VELO tracks with performance comparable to the production 'search by triplet' algorithm while running end to end in the GPU-based first-level trigger, with additiona...

  7. PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs

    cs.LG 2025-06 conditional novelty 5.0 of 10

    PC-MoE shards the expert layers of an MoE LLM across parties and routes only sparse top-k activations between them, achieving near-centralized accuracy with about 70% memory savings and resistance to one partial-gradi...

  8. From Propagator to Oscillator: The Dual Role of Symmetric Differential Equations in Neural Systems

    cs.NE 2025-07 reject novelty 4.0 of 10

    The same symmetric differential equation system can act as a stable signal propagator or as a self-oscillating signal generator, with the mode controlled by a parameter or by inhibitory-loop topology.

  9. The Generalist Brain Module: Module Repetition in Neural Networks in Light of the Minicolumn Hypothesis

    q-bio.NC 2025-07 conditional novelty 4.0 of 10

    A review arguing that repeating a single generalist neural module, inspired by cortical minicolumns, yields robustness, scalability, and generalization benefits compared to monolithic networks.

  10. What Makes Local Updates Effective: The Role of Data Heterogeneity and Smoothness

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Under bounded second-order heterogeneity, local updates are shown to achieve faster convergence than mini-batch SGD in several convex and non-convex regimes, with matching lower bounds.

  11. Improve Underwater Object Detection through YOLOv12 Architecture and Physics-informed Augmentation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Applying YOLOv12 with physics-flavored augmentations yields high reported mAP on four underwater detection benchmarks, but the claims are weakened by missing code, variance, and inconsistent speed numbers.

  12. A Layered Self-Supervised Knowledge Distillation Framework for Efficient Multimodal Learning on the Edge

    cs.CV 2025-06 reject novelty 3.0 of 10

    LSSKD trains compact classifiers with auxiliary self-supervised branches at each stage and past-epoch soft labels as targets, claiming teacher-free accuracy gains on classification benchmarks.

Pith tools