REVIEW 31 cited by
AST: Audio Spectrogram Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
AST: Audio Spectrogram Transformer
read the original abstract
In the past decade, convolutional neural networks (CNNs) have been widely adopted as the main building block for end-to-end audio classification models, which aim to learn a direct mapping from audio spectrograms to corresponding labels. To better capture long-range global context, a recent trend is to add a self-attention mechanism on top of the CNN, forming a CNN-attention hybrid model. However, it is unclear whether the reliance on a CNN is necessary, and if neural networks purely based on attention are sufficient to obtain good performance in audio classification. In this paper, we answer the question by introducing the Audio Spectrogram Transformer (AST), the first convolution-free, purely attention-based model for audio classification. We evaluate AST on various audio classification benchmarks, where it achieves new state-of-the-art results of 0.485 mAP on AudioSet, 95.6% accuracy on ESC-50, and 98.1% accuracy on Speech Commands V2.
Forward citations
Cited by 31 Pith papers
-
VZCrash: A Large-Scale IMU Dataset of Ego-Vehicle Crashes
Introduces VZCrash, the largest public IMU dataset for ego-vehicle crashes, and shows through benchmarks that larger data scale improves crash detection models especially for real-world deployment.
-
SEABAD: A Tropical Bird Activity Detection Dataset for Passive Acoustic Monitoring
SEABAD is a publicly released, balanced dataset of 50,000 curated 16 kHz audio clips spanning 1,677 tropical bird species, with a dual-branch curation pipeline and MobileNetV3-Small baseline reaching 99.57% accuracy.
-
SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification
SpurAudio benchmark shows state-of-the-art few-shot audio classifiers suffer large performance drops when background correlations are disrupted, even in large pretrained models.
-
LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations
LIMSSR reformulates incomplete multimodal learning as LLM-driven sequence-to-score reasoning with prompt-guided imputation and mask-aware aggregation, outperforming baselines on action quality assessment without compl...
-
SAND: The Challenge on Speech Analysis for Neurodegenerative Disease Assessment
SAND creates a clinically annotated speech dataset and associated challenge to enable AI models for automatic early detection and progression prediction of ALS from voice signals.
-
AILive Mixer: A Deep Learning based Zero Latency Automatic Music Mixer for Live Music Performances
AILive Mixer uses deep learning to predict mono gains for multitrack live audio inputs, handling bleeds and achieving zero latency as the first such end-to-end system for live performances.
-
M2R2: MultiModal Robotic Representation for Temporal Action Segmentation
M2R2 proposes a multimodal robotic representation for temporal action segmentation that combines proprioceptive and exteroceptive sensors with a novel training strategy enabling feature reuse across models, achieving ...
-
Unified Audio Intelligence Without Regressing on Text Intelligence
Audex unifies audio understanding and generation on a strong text MoE backbone with multi-stage SFT plus text-only Cascade RL, matching open SOTA audio scores while mostly retaining text capability.
-
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
Automatically constructed synthetic exact-GT data plus multi-model pseudo-labels with interval-aware GRPO rewards improve LALM open-vocabulary audio event grounding on AEGBench and DESED.
-
Missingness as Signal: Channel-Independent Spectrogram Learning for Clinical Time Series Prediction
A channel-independent spectrogram model with an aligned missingness stream improves ICU mortality prediction by treating which measurements were taken as structured 2D signal.
-
Hierarchical Policy Learning via Spectral Decomposition
Causal Spectral Policy decomposes actions spectrally into coarse motion from obs/language and conditional fine corrections, outperforming baselines on precision manipulation tasks.
-
Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks
A conditional generator operating in neural audio codec latent space produces targeted adversarial audio examples in one forward pass, reaching up to 99% success rate at sub-7 ms inference.
-
Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding
Spatial-Omni introduces an SO-Encoder and new datasets to integrate FOA spatial audio into Omni LLMs, improving results on 16 spatial subtasks while preserving general audio performance.
-
Executable Boundary Contracts for Sound Event Traces
Defines executable boundary contracts for sound event traces using an STL-embeddable Boolean fragment plus interval and duration clauses, then evaluates them on speech and soundscape data where they disagree with stan...
-
Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology
Meow-Omni 1 is a quad-modal MLLM that fuses video, audio, physiological time-series, and text to achieve 71.16% accuracy on feline intent recognition in the new MeowBench benchmark.
-
MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment
MLCR organizes quality cues at intra-modal, cross-modal, and stage-wise levels to improve long-term multimodal action quality assessment, achieving top results on gymnastics datasets.
-
MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment
PIDNet uses progressive implicit decoupling with iMambaWave and Group3M blocks to fuse multimodal cues for improved action quality assessment on gymnastics datasets.
-
SAND: The Challenge on Speech Analysis for Neurodegenerative Disease Assessment
SAND creates a new annotated speech dataset and open challenge to benchmark AI models for automatic early identification and progression prediction of ALS using voice signals.
-
VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection
VAN-AD adapts a pretrained visual MAE with distribution mapping and normalizing flow modules to detect anomalies in time series data more effectively across different datasets.
-
Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment
Acoustic scattering signals fed into fine-tuned self-supervised deep learning models classify hair type and moisture at nearly 90% accuracy as a non-invasive alternative to visual methods.
-
Histogram-based Parameter-efficient Tuning for Passive and Active Sonar Classification
HPT uses histograms of feature embeddings to modulate pre-trained models for sonar classification, achieving higher accuracy than standard adapters on passive sonar datasets like VTUAD.
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
VaViT adapts vanilla ViT for point cloud semantic segmentation on nuScenes, SemanticKITTI, and Waymo, matching or exceeding SOTA performance with a tokenizer, lightweight decoder, and augmentations.
-
ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals
ULTRAS unifies audio and speech representation learning in a single transformer by applying patch masking to log-mel spectrograms and using a joint spectral-temporal prediction loss.
-
You're Pushing My Buttons: Instrumented Learning of Gentle Button Presses
Training-time instrumentation with audio and privileged button-state signals produces contact policies that match success rates but apply lower forces using only vision and audio at inference.
-
Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types
On WiWa, factorised multi-task bird species and call-type classification with adaptive loss balancing improves call-type recognition most consistently, while preferred weighting and adaptation depth depend on backbone...
-
From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping
A user study found that 71% of 24 participants preferred an improved multimodal HRI grasping system over baseline, with significantly higher ratings on three perceptual scales after statistical correction.
-
Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks
MLLMs achieve only 42% accuracy on a new audio-visual task requiring second-order spatial ToM under perceptual limits, while a proposed sensory-bounded CoT outperforms egocentric and allocentric baselines.
-
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning
A survey that organizes audio SSL into five objective paradigms, relates their demands to architectural biases, and interprets downstream applications as tests of generalization.
-
Transformer Based Machine Fault Detection From Audio Input
Transformer architectures are applied to machine fault detection from audio spectrograms and compared to CNN embeddings.
-
Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task
An ablation study isolates the contributions of LLM choice, visual perception configuration, and motion controller to success rate and execution time in a human-robot grasping task.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.