Pith. sign in

REVIEW 2 cited by

ZiCo: Zero-shot NAS via Inverse Coefficient of Variation on Gradients

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.11300 v3 pith:ND6K7ZEC submitted 2023-01-26 cs.LG

classification cs.LG
keywords zicozero-shotarchitecturesbetterneuralproxiesproxysearch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural Architecture Search (NAS) is widely used to automatically obtain the neural network with the best performance among a large number of candidate architectures. To reduce the search time, zero-shot NAS aims at designing training-free proxies that can predict the test performance of a given architecture. However, as shown recently, none of the zero-shot proxies proposed to date can actually work consistently better than a naive proxy, namely, the number of network parameters (#Params). To improve this state of affairs, as the main theoretical contribution, we first reveal how some specific gradient properties across different samples impact the convergence rate and generalization capacity of neural networks. Based on this theoretical analysis, we propose a new zero-shot proxy, ZiCo, the first proxy that works consistently better than #Params. We demonstrate that ZiCo works better than State-Of-The-Art (SOTA) proxies on several popular NAS-Benchmarks (NASBench101, NATSBench-SSS/TSS, TransNASBench-101) for multiple applications (e.g., image classification/reconstruction and pixel-level prediction). Finally, we demonstrate that the optimal architectures found via ZiCo are as competitive as the ones found by one-shot and multi-shot NAS methods, but with much less search time. For example, ZiCo-based NAS can find optimal architectures with 78.1%, 79.4%, and 80.4% test accuracy under inference budgets of 450M, 600M, and 1000M FLOPs, respectively, on ImageNet within 0.4 GPU days. Our code is available at https://github.com/SLDGroup/ZiCo.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PrototypeNAS: Rapid Design of Deep Neural Networks for Microcontroller Units

    cs.AI 2026-03 conditional novelty 6.0 of 10

    A three-step zero-shot NAS pipeline with multi-architecture search, proxy ensembles as MOO objectives, and Hypervolume subset selection yields MCU-deployable DNNs competitive with larger models across 12 datasets.

  2. Coflex: Enhancing HW-NAS with Sparse Gaussian Processes for Efficient and Scalable DNN Accelerator Design

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Coflex applies sparse Gaussian processes to multi-objective hardware-aware NAS, claiming near-linear scaling and superior Pareto fronts for DNN accelerator co-design.

Pith tools