Pith. sign in

REVIEW 31 cited by

AST: Audio Spectrogram Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.01778 v3 pith:UTUOOLIP submitted 2021-04-05 cs.SD cs.AI

AST: Audio Spectrogram Transformer

classification cs.SD cs.AI
keywords audioclassificationaccuracymodelnetworksneuralpurelyspectrogram
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In the past decade, convolutional neural networks (CNNs) have been widely adopted as the main building block for end-to-end audio classification models, which aim to learn a direct mapping from audio spectrograms to corresponding labels. To better capture long-range global context, a recent trend is to add a self-attention mechanism on top of the CNN, forming a CNN-attention hybrid model. However, it is unclear whether the reliance on a CNN is necessary, and if neural networks purely based on attention are sufficient to obtain good performance in audio classification. In this paper, we answer the question by introducing the Audio Spectrogram Transformer (AST), the first convolution-free, purely attention-based model for audio classification. We evaluate AST on various audio classification benchmarks, where it achieves new state-of-the-art results of 0.485 mAP on AudioSet, 95.6% accuracy on ESC-50, and 98.1% accuracy on Speech Commands V2.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VZCrash: A Large-Scale IMU Dataset of Ego-Vehicle Crashes

    cs.CV 2026-06 unverdicted novelty 7.0

    Introduces VZCrash, the largest public IMU dataset for ego-vehicle crashes, and shows through benchmarks that larger data scale improves crash detection models especially for real-world deployment.

  2. SEABAD: A Tropical Bird Activity Detection Dataset for Passive Acoustic Monitoring

    cs.SD 2026-05 accept novelty 7.0

    SEABAD is a publicly released, balanced dataset of 50,000 curated 16 kHz audio clips spanning 1,677 tropical bird species, with a dual-branch curation pipeline and MobileNetV3-Small baseline reaching 99.57% accuracy.

  3. SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

    cs.CV 2026-05 unverdicted novelty 7.0

    SpurAudio benchmark shows state-of-the-art few-shot audio classifiers suffer large performance drops when background correlations are disrupted, even in large pretrained models.

  4. LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations

    cs.CV 2026-05 unverdicted novelty 7.0

    LIMSSR reformulates incomplete multimodal learning as LLM-driven sequence-to-score reasoning with prompt-guided imputation and mask-aware aggregation, outperforming baselines on action quality assessment without compl...

  5. SAND: The Challenge on Speech Analysis for Neurodegenerative Disease Assessment

    eess.AS 2026-04 unverdicted novelty 7.0

    SAND creates a clinically annotated speech dataset and associated challenge to enable AI models for automatic early detection and progression prediction of ALS from voice signals.

  6. AILive Mixer: A Deep Learning based Zero Latency Automatic Music Mixer for Live Music Performances

    eess.AS 2026-03 unverdicted novelty 7.0

    AILive Mixer uses deep learning to predict mono gains for multitrack live audio inputs, handling bleeds and achieving zero latency as the first such end-to-end system for live performances.

  7. M2R2: MultiModal Robotic Representation for Temporal Action Segmentation

    cs.RO 2025-04 unverdicted novelty 7.0

    M2R2 proposes a multimodal robotic representation for temporal action segmentation that combines proprioceptive and exteroceptive sensors with a novel training strategy enabling feature reuse across models, achieving ...

  8. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0

    Audex unifies audio understanding and generation on a strong text MoE backbone with multi-stage SFT plus text-only Cascade RL, matching open SOTA audio scores while mostly retaining text capability.

  9. Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

    cs.SD 2026-07 conditional novelty 6.0

    Automatically constructed synthetic exact-GT data plus multi-model pseudo-labels with interval-aware GRPO rewards improve LALM open-vocabulary audio event grounding on AEGBench and DESED.

  10. Missingness as Signal: Channel-Independent Spectrogram Learning for Clinical Time Series Prediction

    cs.LG 2026-07 conditional novelty 6.0

    A channel-independent spectrogram model with an aligned missingness stream improves ICU mortality prediction by treating which measurements were taken as structured 2D signal.

  11. Hierarchical Policy Learning via Spectral Decomposition

    cs.RO 2026-06 unverdicted novelty 6.0

    Causal Spectral Policy decomposes actions spectrally into coarse motion from obs/language and conditional fine corrections, outperforming baselines on precision manipulation tasks.

  12. Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks

    cs.SD 2026-06 unverdicted novelty 6.0

    A conditional generator operating in neural audio codec latent space produces targeted adversarial audio examples in one forward pass, reaching up to 99% success rate at sub-7 ms inference.

  13. Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

    eess.AS 2026-06 unverdicted novelty 6.0

    Spatial-Omni introduces an SO-Encoder and new datasets to integrate FOA spatial audio into Omni LLMs, improving results on 16 spatial subtasks while preserving general audio performance.

  14. Executable Boundary Contracts for Sound Event Traces

    cs.LO 2026-05 unverdicted novelty 6.0 partial

    Defines executable boundary contracts for sound event traces using an STL-embeddable Boolean fragment plus interval and duration clauses, then evaluates them on speech and soundscape data where they disagree with stan...

  15. Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology

    cs.CL 2026-05 unverdicted novelty 6.0

    Meow-Omni 1 is a quad-modal MLLM that fuses video, audio, physiological time-series, and text to achieve 71.16% accuracy on feline intent recognition in the new MeowBench benchmark.

  16. MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment

    cs.CV 2026-05 unverdicted novelty 6.0

    MLCR organizes quality cues at intra-modal, cross-modal, and stage-wise levels to improve long-term multimodal action quality assessment, achieving top results on gymnastics datasets.

  17. MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment

    cs.CV 2026-05 unverdicted novelty 6.0

    PIDNet uses progressive implicit decoupling with iMambaWave and Group3M blocks to fuse multimodal cues for improved action quality assessment on gymnastics datasets.

  18. SAND: The Challenge on Speech Analysis for Neurodegenerative Disease Assessment

    eess.AS 2026-04 unverdicted novelty 6.0

    SAND creates a new annotated speech dataset and open challenge to benchmark AI models for automatic early identification and progression prediction of ALS using voice signals.

  19. VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection

    cs.LG 2026-03 unverdicted novelty 6.0

    VAN-AD adapts a pretrained visual MAE with distribution mapping and normalizing flow modules to detect anomalies in time series data more effectively across different datasets.

  20. Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment

    cs.SD 2025-06 unverdicted novelty 6.0

    Acoustic scattering signals fed into fine-tuned self-supervised deep learning models classify hair type and moisture at nearly 90% accuracy as a non-invasive alternative to visual methods.

  21. Histogram-based Parameter-efficient Tuning for Passive and Active Sonar Classification

    cs.LG 2025-04 unverdicted novelty 6.0

    HPT uses histograms of feature embeddings to modulate pre-trained models for sonar classification, achieving higher accuracy than standard adapters on passive sonar datasets like VTUAD.

  22. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 5.0

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  23. Vanilla ViT for Automotive Point Cloud Semantic Segmentation

    cs.CV 2026-05 unverdicted novelty 5.0

    VaViT adapts vanilla ViT for point cloud semantic segmentation on nuScenes, SemanticKITTI, and Waymo, matching or exceeding SOTA performance with a tokenizer, lightweight decoder, and augmentations.

  24. ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals

    eess.AS 2026-04 unverdicted novelty 5.0

    ULTRAS unifies audio and speech representation learning in a single transformer by applying patch masking to log-mel spectrograms and using a joint spectral-temporal prediction loss.

  25. You're Pushing My Buttons: Instrumented Learning of Gentle Button Presses

    cs.RO 2026-04 unverdicted novelty 5.0

    Training-time instrumentation with audio and privileged button-state signals produces contact policies that match success rates but apply lower forces using only vision and audio at inference.

  26. Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

    cs.SD 2026-07 conditional novelty 4.5

    On WiWa, factorised multi-task bird species and call-type classification with adaptive loss balancing improves call-type recognition most consistently, while preferred weighting and adaptation depth depend on backbone...

  27. From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping

    cs.RO 2026-07 unverdicted novelty 4.0

    A user study found that 71% of 24 participants preferred an improved multimodal HRI grasping system over baseline, with significantly higher ratings on three perceptual scales after statistical correction.

  28. Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks

    cs.AI 2026-05 unverdicted novelty 4.0

    MLLMs achieve only 42% accuracy on a new audio-visual task requiring second-order spatial ToM under perceptual limits, while a proposed sensory-bounded CoT outperforms egocentric and allocentric baselines.

  29. From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning

    eess.AS 2026-07 unverdicted novelty 3.0

    A survey that organizes audio SSL into five objective paradigms, relates their demands to architectural biases, and interprets downstream applications as tests of generalization.

  30. Transformer Based Machine Fault Detection From Audio Input

    cs.SD 2026-04 unverdicted novelty 3.0

    Transformer architectures are applied to machine fault detection from audio spectrograms and compared to CNN embeddings.

  31. Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task

    cs.RO 2026-05 unverdicted novelty 2.0

    An ablation study isolates the contributions of LLM choice, visual perception configuration, and motion controller to success rate and execution time in a human-robot grasping task.