Pith. sign in

REVIEW 2 cited by

Coding schemes in neural networks learning classification tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16689 v1 pith:OC2WHAMU submitted 2024-06-24 cs.LG cond-mat.dis-nncond-mat.stat-mechstat.ML

classification cs.LGcond-mat.dis-nncond-mat.stat-mechstat.ML
keywords networkslearningcodingrepresentationsneuralemergentschemestrong
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural networks posses the crucial ability to generate meaningful representations of task-dependent features. Indeed, with appropriate scaling, supervised learning in neural networks can result in strong, task-dependent feature learning. However, the nature of the emergent representations, which we call the `coding scheme', is still unclear. To understand the emergent coding scheme, we investigate fully-connected, wide neural networks learning classification tasks using the Bayesian framework where learning shapes the posterior distribution of the network weights. Consistent with previous findings, our analysis of the feature learning regime (also known as `non-lazy', `rich', or `mean-field' regime) shows that the networks acquire strong, data-dependent features. Surprisingly, the nature of the internal representations depends crucially on the neuronal nonlinearity. In linear networks, an analog coding scheme of the task emerges. Despite the strong representations, the mean predictor is identical to the lazy case. In nonlinear networks, spontaneous symmetry breaking leads to either redundant or sparse coding schemes. Our findings highlight how network properties such as scaling of weights and neuronal nonlinearity can profoundly influence the emergent representations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptive kernel predictors from feature-learning infinite limits of neural networks

    cs.LG 2025-02 conditional novelty 7.0 of 10

    Feature-learning infinite-width neural networks are kernel machines with data-dependent kernels, defined by a min-max saddle point (Bayesian/Langevin) or a DMFT fixed point (gradient flow with weight decay).

  2. From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning

    cond-mat.dis-nn 2025-02 conditional novelty 7.0 of 10

    A multi-scale adaptive theory shows that kernel rescaling and directional feature adaptation are two approximations of the same posterior distribution, with differences appearing in output covariances and in non-linea...

Pith tools