Pith. sign in

REVIEW 3 cited by

Few-Shot Learning via Learning the Representation, Provably

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.09434 v2 pith:PNBI7VW7 submitted 2020-02-21 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords representationlearningfracleftrightratesourcecommon
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This paper studies few-shot learning via representation learning, where one uses $T$ source tasks with $n_1$ data per task to learn a representation in order to reduce the sample complexity of a target task for which there is only $n_2 (\ll n_1)$ data. Specifically, we focus on the setting where there exists a good \emph{common representation} between source and target, and our goal is to understand how much of a sample size reduction is possible. First, we study the setting where this common representation is low-dimensional and provide a fast rate of $O\left(\frac{\mathcal{C}\left(\Phi\right)}{n_1T} + \frac{k}{n_2}\right)$; here, $\Phi$ is the representation function class, $\mathcal{C}\left(\Phi\right)$ is its complexity measure, and $k$ is the dimension of the representation. When specialized to linear representation functions, this rate becomes $O\left(\frac{dk}{n_1T} + \frac{k}{n_2}\right)$ where $d (\gg k)$ is the ambient input dimension, which is a substantial improvement over the rate without using representation learning, i.e. over the rate of $O\left(\frac{d}{n_2}\right)$. This result bypasses the $\Omega(\frac{1}{T})$ barrier under the i.i.d. task assumption, and can capture the desired property that all $n_1T$ samples from source tasks can be \emph{pooled} together for representation learning. Next, we consider the setting where the common representation may be high-dimensional but is capacity-constrained (say in norm); here, we again demonstrate the advantage of representation learning in both high-dimensional linear regression and neural network learning. Our results demonstrate representation learning can fully utilize all $n_1T$ samples from source tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Attention-based representations for multi-task computation

    cs.LG 2026-08 accept novelty 7.0 of 10

    For min/max readout, two attention heads beat one head by an exponential resource gap, and for n-bit parity and symmetric Boolean functions, heads times polynomial degree must reach the threshold degree, with matching...

  2. Joint estimation of smooth graph signals from partial linear measurements

    math.ST 2025-05 conditional novelty 6.0 of 10

    A smoothness-penalized least squares estimator jointly recovers graph-indexed signals from partial noisy measurements, and is weakly consistent as the graph grows even when per-vertex measurements are rank-one and mos...

  3. Minimax Optimal Two-Stage Algorithm For Moment Estimation Under Covariate Shift

    stat.ML 2025-06 reject novelty 5.0 of 10

    A two-stage importance-weighted estimator is claimed to achieve the minimax rate for moment estimation under covariate shift, but the lower-bound proof is flawed.

Pith tools