Pith. sign in

REVIEW 9 cited by

Audio Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.00335 v2 pith:V6BSAQE7 submitted 2021-05-01 cs.SD cs.AIcs.LGcs.MMeess.AS

classification cs.SDcs.AIcs.LGcs.MMeess.AS
keywords architecturesaudioconvolutionalmodelstransformercomputerdesignedimprove
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Over the past two decades, CNN architectures have produced compelling models of sound perception and cognition, learning hierarchical organizations of features. Analogous to successes in computer vision, audio feature classification can be optimized for a particular task of interest, over a wide variety of datasets and labels. In fact similar architectures designed for image understanding have proven effective for acoustic scene analysis. Here we propose applying Transformer based architectures without convolutional layers to raw audio signals. On a standard dataset of Free Sound 50K,comprising of 200 categories, our model outperforms convolutional models to produce state of the art results. This is significant as unlike in natural language processing and computer vision, we do not perform unsupervised pre-training for outperforming convolutional architectures. On the same training set, with respect mean aver-age precision benchmarks, we show a significant improvement. We further improve the performance of Transformer architectures by using techniques such as pooling inspired from convolutional net-work designed in the past few years. In addition, we also show how multi-rate signal processing ideas inspired from wavelets, can be applied to the Transformer embeddings to improve the results. We also show how our models learns a non-linear non constant band-width filter-bank, which shows an adaptable time frequency front end representation for the task of audio understanding, different from other tasks e.g. pitch estimation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes

    cs.GR 2026-07 conditional novelty 6.0 of 10

    Mass-weighted FEM attention on intrinsic mesh features is triangulation-agnostic and beats current mesh and point-cloud baselines on several geometry-learning benchmarks.

  2. The iNaturalist Sounds Dataset

    cs.SD 2025-05 accept novelty 6.0 of 10

    A new large-scale, weakly labeled audio dataset of 230K recordings across 5,569 species, with benchmarks showing that models trained on it transfer to downstream bioacoustic classification.

  3. Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

    cs.SD 2024-12 conditional novelty 6.0 of 10

    A hybrid causal transformer that combines mel-spectrogram frames with EnCodec acoustic tokens matches or beats a 10-times larger token-only GPT on next-token likelihood for speech and music.

  4. Homeostasis and Sparsity in Transformer

    cs.LG 2024-11 reject novelty 6.0 of 10

    RFB-kWTA and Smart Inhibition, two activation-statistics-based sparsity mechanisms, are reported to improve transformer BLEU on Multi30K from 0.2768 to 0.3062, but without error bars and with best-of-grid selection.

  5. SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A small convolutional adapter plus a frozen patch embedding lets SAM segment depth, thermal, polarization, HHA, and NIR images far better than training from scratch, with parameter-efficient fine-tuning matching full ...

  6. LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging

    cs.SD 2025-01 conditional novelty 5.0 of 10

    A graph neural network that mixes k-nearest-neighbor and fuzzy C-means cluster features outperforms transformer baselines on AudioSet, FSD50K, and ESC-50.

  7. Probing Audio-Generation Capabilities of Text-Based Language Models

    cs.SD 2025-05 conditional novelty 4.0 of 10

    Text-only LLMs can synthesize simple musical notes via generated Python code, but their environmental sound outputs score near chance and speech generation fails entirely.

  8. Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers

    cs.LG 2025-02 conditional novelty 4.0 of 10

    Hamming Attention Distillation binarizes transformer keys and queries to +1/-1 and prunes attention to the top N links, reporting single-point accuracy losses and large simulated hardware savings.

  9. Advances in Transformers for Robotic Applications: A Review

    cs.RO 2024-12 unverdicted

    A survey of Transformer applications in robotic perception, planning, control, human-robot interaction, and reinforcement learning.

Pith tools