Pith. sign in

REVIEW 2 cited by

Shift-Invariance Sparse Coding for Audio Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1206.5241 v1 pith:SI5SSJ6S submitted 2012-06-20 cs.LG stat.ML

classification cs.LGstat.ML
keywords sparsecodingclassificationlearninglinearproblemvariablesfeatures
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Sparse coding is an unsupervised learning algorithm that learns a succinct high-level representation of the inputs given only unlabeled data; it represents each input as a sparse linear combination of a set of basis functions. Originally applied to modeling the human visual cortex, sparse coding has also been shown to be useful for self-taught learning, in which the goal is to solve a supervised classification task given access to additional unlabeled data drawn from different classes than that in the supervised learning problem. Shift-invariant sparse coding (SISC) is an extension of sparse coding which reconstructs a (usually time-series) input using all of the basis functions in all possible shifts. In this paper, we present an efficient algorithm for learning SISC bases. Our method is based on iteratively solving two large convex optimization problems: The first, which computes the linear coefficients, is an L1-regularized linear least squares problem with potentially hundreds of thousands of variables. Existing methods typically use a heuristic to select a small subset of the variables to optimize, but we present a way to efficiently compute the exact solution. The second, which solves for bases, is a constrained linear least squares problem. By optimizing over complex-valued variables in the Fourier domain, we reduce the coupling between the different variables, allowing the problem to be solved efficiently. We show that SISC's learned high-level representations of speech and music provide useful features for classification tasks within those domains. When applied to classification, under certain conditions the learned features outperform state of the art spectral and cepstral features.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sparse Generative Adversarial Network

    cs.CV 2019-08 conditional novelty 5.0 of 10

    A sparse-patch GAN with an encoder reconstructor reports improved Inception scores on CIFAR-10 and CelebA, but its mode-collapse guarantee is asserted, not proven.

  2. Modeling Musical Genre Trajectories through Pathlet Learning

    cs.IR 2025-05 conditional novelty 4.0 of 10

    Pathlet learning on genre-ranked listening trajectories yields interpretable embeddings that slightly improve prediction of genre appearance and disappearance on Deezer, but not consistently on Last.fm.

Pith tools