Pith. sign in

REVIEW 2 cited by

Online Learning of Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.09167 v1 pith:62WSKAAN submitted 2025-05-14 stat.ML cs.LG

classification stat.MLcs.LG
keywords gammaboundapproximatelymarginmistakenetworkeveryfirst
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study online learning of feedforward neural networks with the sign activation function that implement functions from the unit ball in $\mathbb{R}^d$ to a finite label set $\{1, \ldots, Y\}$. First, we characterize a margin condition that is sufficient and in some cases necessary for online learnability of a neural network: Every neuron in the first hidden layer classifies all instances with some margin $\gamma$ bounded away from zero. Quantitatively, we prove that for any net, the optimal mistake bound is at most approximately $\mathtt{TS}(d,\gamma)$, which is the $(d,\gamma)$-totally-separable-packing number, a more restricted variation of the standard $(d,\gamma)$-packing number. We complement this result by constructing a net on which any learner makes $\mathtt{TS}(d,\gamma)$ many mistakes. We also give a quantitative lower bound of approximately $\mathtt{TS}(d,\gamma) \geq \max\{1/(\gamma \sqrt{d})^d, d\}$ when $\gamma \geq 1/2$, implying that for some nets and input sequences every learner will err for $\exp(d)$ many times, and that a dimension-free mistake bound is almost always impossible. To remedy this inevitable dependence on $d$, it is natural to seek additional natural restrictions to be placed on the network, so that the dependence on $d$ is removed. We study two such restrictions. The first is the multi-index model, in which the function computed by the net depends only on $k \ll d$ orthonormal directions. We prove a mistake bound of approximately $(1.5/\gamma)^{k + 2}$ in this model. The second is the extended margin assumption. In this setting, we assume that all neurons (in all layers) in the network classify every ingoing input from previous layer with margin $\gamma$ bounded away from zero. In this model, we prove a mistake bound of approximately $(\log Y)/ \gamma^{O(L)}$, where L is the depth of the network.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mistake-bounded online learning with operation caps

    cs.LG 2025-09 conditional novelty 7.0 of 10

    The paper proves opt_ag,weak(F, eta) = O((k ln k) opt_std(F) + k eta), matching the known lower bound, and develops a new operation-caps model for mistake-bounded online learning.

  2. The Generative Leap: Sharp Sample Complexity for Efficiently Learning Gaussian Multi-Index Models

    cs.LG 2025-06 conditional novelty 7.0 of 10

    For any Gaussian multi-index model, the generative leap exponent k⋆ sharply characterizes the sample complexity of efficient subspace recovery as Θ(d^(1∨k⋆/2)).

Pith tools