REVIEW 9 cited by
Audio Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Over the past two decades, CNN architectures have produced compelling models of sound perception and cognition, learning hierarchical organizations of features. Analogous to successes in computer vision, audio feature classification can be optimized for a particular task of interest, over a wide variety of datasets and labels. In fact similar architectures designed for image understanding have proven effective for acoustic scene analysis. Here we propose applying Transformer based architectures without convolutional layers to raw audio signals. On a standard dataset of Free Sound 50K,comprising of 200 categories, our model outperforms convolutional models to produce state of the art results. This is significant as unlike in natural language processing and computer vision, we do not perform unsupervised pre-training for outperforming convolutional architectures. On the same training set, with respect mean aver-age precision benchmarks, we show a significant improvement. We further improve the performance of Transformer architectures by using techniques such as pooling inspired from convolutional net-work designed in the past few years. In addition, we also show how multi-rate signal processing ideas inspired from wavelets, can be applied to the Transformer embeddings to improve the results. We also show how our models learns a non-linear non constant band-width filter-bank, which shows an adaptable time frequency front end representation for the task of audio understanding, different from other tasks e.g. pitch estimation.
Forward citations
Cited by 9 Pith papers
-
Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes
Mass-weighted FEM attention on intrinsic mesh features is triangulation-agnostic and beats current mesh and point-cloud baselines on several geometry-learning benchmarks.
-
The iNaturalist Sounds Dataset
A new large-scale, weakly labeled audio dataset of 230K recordings across 5,569 species, with benchmarks showing that models trained on it transfer to downstream bioacoustic classification.
-
Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music
A hybrid causal transformer that combines mel-spectrogram frames with EnCodec acoustic tokens matches or beats a 10-times larger token-only GPT on next-token likelihood for speech and music.
-
Homeostasis and Sparsity in Transformer
RFB-kWTA and Smart Inhibition, two activation-statistics-based sparsity mechanisms, are reported to improve transformer BLEU on Multi30K from 0.2768 to 0.3062, but without error bars and with best-of-grid selection.
-
SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality
A small convolutional adapter plus a frozen patch embedding lets SAM segment depth, thermal, polarization, HHA, and NIR images far better than training from scratch, with parameter-efficient fine-tuning matching full ...
-
LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging
A graph neural network that mixes k-nearest-neighbor and fuzzy C-means cluster features outperforms transformer baselines on AudioSet, FSD50K, and ESC-50.
-
Probing Audio-Generation Capabilities of Text-Based Language Models
Text-only LLMs can synthesize simple musical notes via generated Python code, but their environmental sound outputs score near chance and speech generation fails entirely.
-
Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers
Hamming Attention Distillation binarizes transformer keys and queries to +1/-1 and prunes attention to the top N links, reporting single-point accuracy losses and large simulated hardware savings.
-
Advances in Transformers for Robotic Applications: A Review
A survey of Transformer applications in robotic perception, planning, control, human-robot interaction, and reinforcement learning.
Discussion (0). Continue with ORCID to comment.