Pith. sign in

REVIEW 12 cited by

Efficient Training of Audio Transformers with Patchout

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.05069 v3 pith:ODYMEM4B submitted 2021-10-11 cs.SD cs.LGeess.AS

Efficient Training of Audio Transformers with Patchout

classification cs.SD cs.LGeess.AS
keywords transformersaudiocnnsmodelsperformanceworkcomplexitypropose
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The great success of transformer-based models in natural language processing (NLP) has led to various attempts at adapting these architectures to other domains such as vision and audio. Recent work has shown that transformers can outperform Convolutional Neural Networks (CNNs) on vision and audio tasks. However, one of the main shortcomings of transformer models, compared to the well-established CNNs, is the computational complexity. In transformers, the compute and memory complexity is known to grow quadratically with the input length. Therefore, there has been extensive work on optimizing transformers, but often at the cost of degrading predictive performance. In this work, we propose a novel method to optimize and regularize transformers on audio spectrograms. Our proposed models achieve a new state-of-the-art performance on Audioset and can be trained on a single consumer-grade GPU. Furthermore, we propose a transformer model that outperforms CNNs in terms of both performance and training speed. Source code: https://github.com/kkoutini/PaSST

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators

    cs.SD 2026-05 unverdicted novelty 7.0

    Live Music Diffusion Models adapt bidirectional diffusion for interactive music generation via KV caching and ARC-Forcing, recovering and exceeding discrete autoregressive efficiency while enabling post-training align...

  2. SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

    cs.CV 2026-05 unverdicted novelty 7.0

    SpurAudio benchmark shows state-of-the-art few-shot audio classifiers suffer large performance drops when background correlations are disrupted, even in large pretrained models.

  3. Omni2Sound: Towards Unified Video-Text-to-Audio Generation

    cs.SD 2026-01 unverdicted novelty 7.0

    A single DiT-based diffusion model unifies video-to-audio, text-to-audio, and joint video-text-to-audio generation, supported by a new 470k-pair dataset and three-stage progressive training that resolves task competition.

  4. Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation

    cs.SD 2026-06 unverdicted novelty 6.0

    A data-free streaming consistency distillation framework enables single-step autoregressive generation from text-to-music models for real-time interactive use while preserving timbre and rhythm via latent, spectral, a...

  5. ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling

    cs.MM 2026-04 unverdicted novelty 6.0

    ControlFoley introduces a unified framework for controllable video-to-audio generation using joint visual encoding, temporal-timbre decoupling, and robust multimodal training to handle cross-modal conflicts.

  6. Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

    cs.CV 2026-02 unverdicted novelty 6.0

    MMHNet enables video-to-audio models trained on short clips to generalize and generate audio for videos over 5 minutes long.

  7. Conditional Flow Matching for Visually-Guided Acoustic Highlighting

    eess.AS 2026-02 conditional novelty 6.0

    Conditional flow matching with a rollout loss and early audio-visual fusion achieves state-of-the-art results on visually-guided acoustic highlighting.

  8. FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation

    eess.AS 2026-07 conditional novelty 5.5

    MeanFlow-anchored multi-representation FD post-training improves one-step text-to-audio quality without collapsing multi-step sampling.

  9. BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing

    cs.SD 2026-05 unverdicted novelty 5.0

    Unsupervised VQ-VAE training on PaSST embeddings discovers repeatable discrete acoustic tokens in honey bee buzzing that separate queenright from queenless conditions and identify three stable sub-states in queenless hives.

  10. JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

    cs.CV 2025-12 conditional novelty 5.0

    A dual-branch diffusion transformer with joint video-audio self-attention and a keypoint-based mouth-area loss reports top lip-sync and speech metrics on two benchmarks.

  11. STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

    eess.AS 2026-06 unverdicted novelty 4.0

    STAR-VAE introduces topology-aware regularization to reshape VAE latent geometry for audio, claiming to resolve the Rate-Distortion-Regularity Trilemma and achieve SOTA reconstruction.

  12. MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

    cs.SD 2025-10 unverdicted novelty 4.0

    MMAudioSep adapts a pretrained video-to-audio model via fine-tuning for video/text-queried sound separation, outperforming baselines while preserving generation ability.