Pith. sign in

REVIEW 1 cited by

Towards Task and Architecture-Independent Generalization Gap Predictors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.01550 v1 pith:DNGA7CSQ submitted 2019-06-04 stat.ML cs.LG

classification stat.MLcs.LG
keywords architecture-independentdifferentgeneralizationrnnsdatasetdeepdnnslearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Can we use deep learning to predict when deep learning works? Our results suggest the affirmative. We created a dataset by training 13,500 neural networks with different architectures, on different variations of spiral datasets, and using different optimization parameters. We used this dataset to train task-independent and architecture-independent generalization gap predictors for those neural networks. We extend Jiang et al. (2018) to also use DNNs and RNNs and show that they outperform the linear model, obtaining $R^2=0.965$. We also show results for architecture-independent, task-independent, and out-of-distribution generalization gap prediction tasks. Both DNNs and RNNs consistently and significantly outperform linear models, with RNNs obtaining $R^2=0.584$.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DRIFT: Data Reduction via Informative Feature Transformation- Generalization Begins Before Deep Learning starts

    cs.LG 2025-06 conditional novelty 2.0 of 10

    DRIFT preprocesses images by projecting them onto fixed sine mode shapes, which lets small networks train with tens of features while showing smoother convergence than PCA or pixel inputs.

Pith tools