{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2016:OPXDY4I77QZCOJIAYK7MZBA62N","short_pith_number":"pith:OPXDY4I7","schema_version":"1.0","canonical_sha256":"73ee3c711ffc32272500c2becc841ed34b1fd3e6c2e152b19cec9bc8ffd27284","source":{"kind":"arxiv","id":"1604.06174","version":2},"attestation_state":"computed","paper":{"title":"Training Deep Nets with Sublinear Memory Cost","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"An algorithm trains an n-layer deep network using O(sqrt(n)) memory at the cost of one extra forward pass.","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Bing Xu, Carlos Guestrin, Chiyuan Zhang, Tianqi Chen","submitted_at":"2016-04-21T04:15:27Z","abstract_excerpt":"We propose a systematic approach to reduce the memory consumption of deep neural network training. Specifically, we design an algorithm that costs O(sqrt(n)) memory to train a n layer network, with only the computational cost of an extra forward pass per mini-batch. As many of the state-of-the-art models hit the upper bound of the GPU memory, our algorithm allows deeper and more complex models to be explored, and helps advance the innovations in deep learning research. We focus on reducing the memory cost to store the intermediate feature maps and gradients during training. Computation graph a"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"1604.06174","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2016-04-21T04:15:27Z","cross_cats_sorted":[],"title_canon_sha256":"09e011e37ca446694c16a9daae35de00bbd8cefc18f9a0e6f527cd1274393225","abstract_canon_sha256":"c787b38ba02754d98018eff0fb4fc2ae10b4a8b51e8bf91a93b9912c20a48fb4"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-04T20:58:46.588195Z","signature_b64":"1acO1ffphUZGym1d38/nWMZLRyulfz9VduveIB+J3NEU2y4D4w0khVJgjAJqFwb0uq38PZHNFQu2oTSAlu7ACw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"73ee3c711ffc32272500c2becc841ed34b1fd3e6c2e152b19cec9bc8ffd27284","last_reissued_at":"2026-07-04T20:58:46.587642Z","signature_status":"signed_v1","first_computed_at":"2026-07-04T20:58:46.587642Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Training Deep Nets with Sublinear Memory Cost","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"An algorithm trains an n-layer deep network using O(sqrt(n)) memory at the cost of one extra forward pass.","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Bing Xu, Carlos Guestrin, Chiyuan Zhang, Tianqi Chen","submitted_at":"2016-04-21T04:15:27Z","abstract_excerpt":"We propose a systematic approach to reduce the memory consumption of deep neural network training. Specifically, we design an algorithm that costs O(sqrt(n)) memory to train a n layer network, with only the computational cost of an extra forward pass per mini-batch. As many of the state-of-the-art models hit the upper bound of the GPU memory, our algorithm allows deeper and more complex models to be explored, and helps advance the innovations in deep learning research. We focus on reducing the memory cost to store the intermediate feature maps and gradients during training. Computation graph a"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"we design an algorithm that costs O(sqrt(n)) memory to train a n layer network, with only the computational cost of an extra forward pass per mini-batch.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"The computation graph can be cleanly segmented into sqrt(n) intervals where recomputing forward passes inside each interval is both correct and cheaper than storing all intermediate activations.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"An algorithm trains n-layer networks with O(sqrt(n)) memory via selective recomputation of activations, at the cost of one extra forward pass.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"An algorithm trains an n-layer deep network using O(sqrt(n)) memory at the cost of one extra forward pass.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"98958f2c83cbc1aca57783a43ec86b0e032b3df334bca54776e672a99b7bdeb5"},"source":{"id":"1604.06174","kind":"arxiv","version":2},"verdict":{"id":"54fd2481-e864-4ac0-b13a-338671298d71","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-12T03:37:54.105843Z","strongest_claim":"we design an algorithm that costs O(sqrt(n)) memory to train a n layer network, with only the computational cost of an extra forward pass per mini-batch.","one_line_summary":"An algorithm trains n-layer networks with O(sqrt(n)) memory via selective recomputation of activations, at the cost of one extra forward pass.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"The computation graph can be cleanly segmented into sqrt(n) intervals where recomputing forward passes inside each interval is both correct and cheaper than storing all intermediate activations.","pith_extraction_headline":"An algorithm trains an n-layer deep network using O(sqrt(n)) memory at the cost of one extra forward pass."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/1604.06174/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":19,"sample":[{"doi":"","year":2015,"title":"Mart ´ın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Good- fellow, Andrew Harp, Geoffr","work_id":"88a19a08-50d4-45ca-9c7a-9ac7a1dcb85f","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2014,"title":"Seltzer, Malcolm Slaney, Andreas Stolcke, Yongqiang Wang, Huaming Wang, Kaisheng Yao, Dong Yu, Yu Zhang, and Geoffrey Zweig","work_id":"5f78b504-4380-4587-bb13-6cef5cc780a9","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":1986,"title":"Aho, Ravi Sethi, and Jeffrey D","work_id":"f7065209-a388-4299-9529-cdf786afe897","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2012,"title":"Goodfellow, Arnaud Bergeron, Nicolas Bouchard, and Yoshua Bengio","work_id":"27872beb-9cb7-4732-8a63-c4c093f69c16","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2010,"title":"Theano: a CPU and GPU math expression compiler","work_id":"f53ea14a-639c-4146-bfc1-9285b7c68e27","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":19,"snapshot_sha256":"8a616d8dad6b69a1d8b987f1f5920ef0dac33caabb71233394c28710fc4cec05","internal_anchors":2},"formal_canon":{"evidence_count":2,"snapshot_sha256":"c30732137fe599a7ae3e14eed67371afb49fe612ce9703b3de6cb41f429383bd"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"1604.06174","created_at":"2026-07-04T20:58:46.587711+00:00"},{"alias_kind":"arxiv_version","alias_value":"1604.06174v2","created_at":"2026-07-04T20:58:46.587711+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.1604.06174","created_at":"2026-07-04T20:58:46.587711+00:00"},{"alias_kind":"pith_short_12","alias_value":"OPXDY4I77QZC","created_at":"2026-07-04T20:58:46.587711+00:00"},{"alias_kind":"pith_short_16","alias_value":"OPXDY4I77QZCOJIA","created_at":"2026-07-04T20:58:46.587711+00:00"},{"alias_kind":"pith_short_8","alias_value":"OPXDY4I7","created_at":"2026-07-04T20:58:46.587711+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":103,"internal_anchor_count":103,"sample":[{"citing_arxiv_id":"2607.06609","citing_title":"D2PO: Optimizing Diffusion Samplers via Dynamic Preference","ref_index":6,"is_internal_anchor":true},{"citing_arxiv_id":"2607.07743","citing_title":"Architecture Generalization with MetaNCA","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2607.07675","citing_title":"Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence","ref_index":14,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25987","citing_title":"Weave of Formal Thought","ref_index":42,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24780","citing_title":"BluTrain: A C++/CUDA Framework for AI Systems","ref_index":32,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26361","citing_title":"Does Aurora Encode Atmospheric Structure? Latent Regime Analysis and Attribution","ref_index":8,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24937","citing_title":"The Hitchhiker's Guide to Agentic AI: From Foundations to Systems","ref_index":232,"is_internal_anchor":true},{"citing_arxiv_id":"2606.23546","citing_title":"The Energy Consumption of Transformer Fine-Tuning: A Roofline-Inspired Scaling Model","ref_index":34,"is_internal_anchor":true},{"citing_arxiv_id":"2606.21036","citing_title":"Diffusion-Driven State Space Models","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19538","citing_title":"ITNet: A Learnable Integral Transform That Subsumes Convolution, Attention, and Recurrence","ref_index":10,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19528","citing_title":"Techniques for Peak Memory Reduction for LoRA Fine-tuning of LLMs on Edge Devices","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19025","citing_title":"FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs","ref_index":90,"is_internal_anchor":true},{"citing_arxiv_id":"2607.01394","citing_title":"The Wiola Architecture for Efficient Small Language Models","ref_index":22,"is_internal_anchor":true},{"citing_arxiv_id":"2606.17030","citing_title":"Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation","ref_index":210,"is_internal_anchor":true},{"citing_arxiv_id":"2607.02158","citing_title":"Efficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUs","ref_index":15,"is_internal_anchor":true},{"citing_arxiv_id":"2606.12635","citing_title":"CD-RCM: Generalizable Continuous-Depth Novel View Synthesis for Reflectance Confocal Microscopy","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2606.11682","citing_title":"Parameter-Efficient Adapter Tuning for Tabular-Image Multimodal Learning","ref_index":35,"is_internal_anchor":true},{"citing_arxiv_id":"2606.10415","citing_title":"RATrain: A Resource-Aware Training Runtime for Large Language Models on Bandwidth-Constrained Heterogeneous Supercomputing Platforms","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2607.01844","citing_title":"Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03209","citing_title":"DECA: Decentralizing Block-Wise Adam for Efficient LLM Full-Parameter Fine-Tuning on Non-IID Data","ref_index":9,"is_internal_anchor":true},{"citing_arxiv_id":"2606.02565","citing_title":"Policy-based Foveated Imaging and Perception","ref_index":111,"is_internal_anchor":true},{"citing_arxiv_id":"2606.01143","citing_title":"Schedule-Level Shared-Prefix Reuse for LLM RL Training","ref_index":2,"is_internal_anchor":true},{"citing_arxiv_id":"2606.00267","citing_title":"StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement","ref_index":121,"is_internal_anchor":true},{"citing_arxiv_id":"2604.25863","citing_title":"MCMit: Hardware-Software Co-Design for Mid-Circuit Measurement Error Mitigation","ref_index":6,"is_internal_anchor":true},{"citing_arxiv_id":"2606.31813","citing_title":"Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in RLVR","ref_index":75,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/OPXDY4I77QZCOJIAYK7MZBA62N","json":"https://pith.science/pith/OPXDY4I77QZCOJIAYK7MZBA62N.json","graph_json":"https://pith.science/api/pith-number/OPXDY4I77QZCOJIAYK7MZBA62N/graph.json","events_json":"https://pith.science/api/pith-number/OPXDY4I77QZCOJIAYK7MZBA62N/events.json","paper":"https://pith.science/paper/OPXDY4I7"},"agent_actions":{"view_html":"https://pith.science/pith/OPXDY4I77QZCOJIAYK7MZBA62N","download_json":"https://pith.science/pith/OPXDY4I77QZCOJIAYK7MZBA62N.json","view_paper":"https://pith.science/paper/OPXDY4I7","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=1604.06174&json=true","fetch_graph":"https://pith.science/api/pith-number/OPXDY4I77QZCOJIAYK7MZBA62N/graph.json","fetch_events":"https://pith.science/api/pith-number/OPXDY4I77QZCOJIAYK7MZBA62N/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/OPXDY4I77QZCOJIAYK7MZBA62N/action/timestamp_anchor","attest_storage":"https://pith.science/pith/OPXDY4I77QZCOJIAYK7MZBA62N/action/storage_attestation","attest_author":"https://pith.science/pith/OPXDY4I77QZCOJIAYK7MZBA62N/action/author_attestation","sign_citation":"https://pith.science/pith/OPXDY4I77QZCOJIAYK7MZBA62N/action/citation_signature","submit_replication":"https://pith.science/pith/OPXDY4I77QZCOJIAYK7MZBA62N/action/replication_record"}},"created_at":"2026-07-04T20:58:46.587711+00:00","updated_at":"2026-07-04T20:58:46.587711+00:00"}