Pith. sign in

REVIEW 24 cited by

The Computational Limits of Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.05558 v2 pith:DVYHNZXG submitted 2020-07-10 cs.LG stat.ML

classification cs.LGstat.ML
keywords learningprogressdeepapplicationscomecomputingmethodsother
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep learning's recent history has been one of achievement: from triumphing over humans in the game of Go to world-leading performance in image classification, voice recognition, translation, and other tasks. But this progress has come with a voracious appetite for computing power. This article catalogs the extent of this dependency, showing that progress across a wide variety of applications is strongly reliant on increases in computing power. Extrapolating forward this reliance reveals that progress along current lines is rapidly becoming economically, technically, and environmentally unsustainable. Thus, continued progress in these applications will require dramatically more computationally-efficient methods, which will either have to come from changes to deep learning or from moving to other machine learning methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Koopman Model Dimension Reduction via Variational Bayesian Inference and Graph Search

    eess.SY 2026-01 conditional novelty 6.0 of 10

    Variational Bayesian inclusion-flag estimates, thresholded into a directed graph, select a smaller Koopman dictionary while leaving output influence paths intact.

  2. Quantum optical neural networks using atom-cavity interactions to provide all-optical nonlinearity

    quant-ph 2025-11 conditional novelty 6.0 of 10

    A simulated neural network uses atom-cavity two-level neurons as all-optical nonlinear activations and reports ~95% accuracy on MNIST and SAT-6.

  3. Progressive Depth Up-scaling via Optimal Transport

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Optimal-transport alignment of adjacent layers gives a cheap and slightly better initialization for progressive depth up-scaling than copying, averaging, or a learned predictor.

  4. Beyond Low-rank Decomposition: A Shortcut Approach for Efficient On-Device Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    ASI compresses training activations with a single warm-started subspace iteration and a once-per-model rank selection, cutting on-device training memory by up to 120x and FLOPs by up to 1.86x on standard benchmarks.

  5. Attention Is All You Need For Mixture-of-Depths Routing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A-MoD routes tokens in Mixture-of-Depths vision transformers using attention maps from the prior layer, improving accuracy and convergence without extra trainable parameters.

  6. tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI

    cs.AR 2024-12 conditional novelty 6.0 of 10

    A temporal unary GEMM architecture performs exact low-precision matrix multiply with post-synthesis area and power reductions of up to 15x and 11x versus stochastic uGEMM, at the cost of data-dependent latency.

  7. Language Models in Software Development Tasks: An Experimental Analysis of Energy and Accuracy

    cs.SE 2024-11 conditional novelty 6.0 of 10

    Larger language models do not reliably deliver higher accuracy on software tasks, and quantized large models can outperform full-precision medium models on energy and accuracy together.

  8. A Self-Supervised Robotic System for Autonomous Contact-Based Spatial Mapping of Semiconductor Properties

    cs.RO 2024-11 conditional novelty 6.0 of 10

    A self-supervised, spatially differentiable CNN chooses probe poses on images of perovskite films, a noisy Dijkstra planner routes the robot, and the system autonomously maps photoconductivity across 35 film compositi...

  9. MOSAIC-FL, a micro-service based privacy-preserving framework with application to genomics

    cs.CR 2026-07 conditional novelty 5.0 of 10

    A gRPC micro-service FL stack with t-out-of-N CKKS secure aggregation matches cleartext accuracy on EMNIST and TCGA BRCA subtyping at modest extra cost for large models.

  10. Physical Analogue Kolmogorov-Arnold Networks based on Reconfigurable Nonlinear-Processing Units

    cs.ET 2026-02 conditional novelty 5.0 of 10

    A proposed analog KAN chip uses silicon RNPUs as physically programmable nonlinear edges, with estimated ~250 pJ per inference and ~10x smaller area than a digital MLP.

  11. Real-Time Analysis of Unstructured Data with Machine Learning on Heterogeneous Architectures

    physics.data-an 2025-08 conditional novelty 5.0 of 10

    A graph neural network (ETX4VELO) reconstructs LHCb VELO tracks with performance comparable to the production 'search by triplet' algorithm while running end to end in the GPU-based first-level trigger, with additiona...

  12. PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs

    cs.LG 2025-06 conditional novelty 5.0 of 10

    PC-MoE shards the expert layers of an MoE LLM across parties and routes only sparse top-k activations between them, achieving near-centralized accuracy with about 70% memory savings and resistance to one partial-gradi...

  13. Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Inference energy of seven text-to-audio diffusion models grows linearly with denoising steps, while quality saturates, so the best quality-per-energy settings use 10 to 50 steps.

  14. Single-shot Star-convex Polygon-based Instance Segmentation for Spatially-correlated Biomedical Objects

    cs.CV 2025-04 conditional novelty 5.0 of 10

    A branched StarDist architecture with a within-boundary penalty segments nested biomedical objects such as nuclei in cells and plaques in wells in one shot, with a new joint true-positive metric.

  15. TNNGen: Automated Design of Neuromorphic Sensory Processing Units for Time-Series Clustering

    cs.AR 2024-12 conditional novelty 5.0 of 10

    A new toolchain automatically converts PyTorch temporal neural networks into post-layout chip designs, with forecasting equations for die area and leakage power.

  16. From Propagator to Oscillator: The Dual Role of Symmetric Differential Equations in Neural Systems

    cs.NE 2025-07 reject novelty 4.0 of 10

    The same symmetric differential equation system can act as a stable signal propagator or as a self-oscillating signal generator, with the mode controlled by a parameter or by inhibitory-loop topology.

  17. The Generalist Brain Module: Module Repetition in Neural Networks in Light of the Minicolumn Hypothesis

    q-bio.NC 2025-07 conditional novelty 4.0 of 10

    A review arguing that repeating a single generalist neural module, inspired by cortical minicolumns, yields robustness, scalability, and generalization benefits compared to monolithic networks.

  18. What Makes Local Updates Effective: The Role of Data Heterogeneity and Smoothness

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Under bounded second-order heterogeneity, local updates are shown to achieve faster convergence than mini-batch SGD in several convex and non-convex regimes, with matching lower bounds.

  19. Improve Underwater Object Detection through YOLOv12 Architecture and Physics-informed Augmentation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Applying YOLOv12 with physics-flavored augmentations yields high reported mAP on four underwater detection benchmarks, but the claims are weakened by missing code, variance, and inconsistent speed numbers.

  20. Can a Quantum Support Vector Machine algorithm be utilized to identify Key Biomarkers from Multi-Omics data of COVID19 patients?

    quant-ph 2025-04 reject novelty 4.0 of 10

    Applying a simulated quantum support vector machine to proteomic and metabolomic COVID-19 data gives classification performance comparable to a classical SVM in selected settings, but the paper's biomarker-identificat...

  21. Bayesian Data Augmentation and Training for Perception DNN in Autonomous Aerial Vehicles

    cs.RO 2024-12 conditional novelty 4.0 of 10

    A closed-loop simulation and Bayesian optimization framework selects data augmentation hyperparameters for a YOLO landing-pad detector, improving simulated VTOL landing success from 50% to 70%.

  22. LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models

    cs.CL 2024-11 reject novelty 4.0 of 10

    This paper presents a CPU-only data curation pipeline for LLMs, but the central claim of high-quality output is not supported by any training or quality evaluation.

  23. A Layered Self-Supervised Knowledge Distillation Framework for Efficient Multimodal Learning on the Edge

    cs.CV 2025-06 reject novelty 3.0 of 10

    LSSKD trains compact classifiers with auxiliary self-supervised branches at each stage and past-epoch soft labels as targets, claiming teacher-free accuracy gains on classification benchmarks.

  24. YOLOv11 Optimization for Efficient Resource Utilization

    cs.CV 2024-12 conditional novelty 3.0 of 10

    Pruning YOLOv11's detection heads yields six size-specialized variants that match original accuracy on targeted datasets while cutting model size, FLOPs, and inference time.

Pith tools