Pith. sign in

REVIEW 27 cited by

BEATs: Audio Pre-Training with Acoustic Tokenizers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.09058 v1 pith:5B3FNTLT submitted 2022-12-18 eess.AS cs.AIcs.CLcs.LGcs.SD

classification eess.AScs.AIcs.CLcs.LGcs.SD
keywords audioacousticmodelmodelstokenizerpre-trainingbeatsdiscrete
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The massive growth of self-supervised learning (SSL) has been witnessed in language, vision, speech, and audio domains over the past few years. While discrete label prediction is widely adopted for other modalities, the state-of-the-art audio SSL models still employ reconstruction loss for pre-training. Compared with reconstruction loss, semantic-rich discrete label prediction encourages the SSL model to abstract the high-level audio semantics and discard the redundant details as in human perception. However, a semantic-rich acoustic tokenizer for general audio pre-training is usually not straightforward to obtain, due to the continuous property of audio and unavailable phoneme sequences like speech. To tackle this challenge, we propose BEATs, an iterative audio pre-training framework to learn Bidirectional Encoder representation from Audio Transformers, where an acoustic tokenizer and an audio SSL model are optimized by iterations. In the first iteration, we use random projection as the acoustic tokenizer to train an audio SSL model in a mask and label prediction manner. Then, we train an acoustic tokenizer for the next iteration by distilling the semantic knowledge from the pre-trained or fine-tuned audio SSL model. The iteration is repeated with the hope of mutual promotion of the acoustic tokenizer and audio SSL model. The experimental results demonstrate our acoustic tokenizers can generate discrete labels with rich audio semantics and our audio SSL models achieve state-of-the-art results across various audio classification benchmarks, even outperforming previous models that use more training data and model parameters significantly. Specifically, we set a new state-of-the-art mAP 50.6% on AudioSet-2M for audio-only models without using any external data, and 98.1% accuracy on ESC-50. The code and pre-trained models are available at https://aka.ms/beats.

Discussion (0). Sign in to comment.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Jointly training audio and video VAEs with segment contrastive loss and semantic distillation yields more learnable, cross-aligned latents that improve downstream joint generation quality and sync.

  2. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  3. Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection

    eess.AS 2026-01 conditional novelty 6.0 of 10

    Better proxy-task performance does not generally improve anomalous sound detection; only source separation showed a strong, consistent positive correlation.

  4. EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

    cs.CV 2025-12 conditional novelty 6.0 of 10

    EchoingPixels prunes audio-visual LLM input tokens jointly across modalities and re-tunes RoPE frequencies so that 5–20% of tokens retain roughly full-model performance.

  5. Assessing Factual Music Comprehension in Large Audio Language Models

    cs.SD 2025-11 conditional novelty 6.0 of 10

    Standard NLP metrics fail to capture factual music understanding in audio-language models; a CLAP-based metric and an LLM-parsed factual QA protocol measure it more directly.

  6. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  7. NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

    cs.SD 2025-08 unverdicted novelty 6.0 of 10

    A Mandarin speech pipeline that tags non-word vocalizations (laughter, breath, filled pauses) for ASR and controls their generation in TTS, backed by a claimed 573-hour, 174,179-utterance word-level annotated corpus.

  8. Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Benchmarks seven open-source audio-video-text MLLMs on six affective datasets and shows a generative-knowledge prompting step improves fine-tuned emotion recognition.

  9. DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

    cs.MM 2025-07 conditional novelty 6.0 of 10

    A single multimodal language model can generate intelligible speech and synchronized background audio jointly from a silent video, transcript, and reference voice, outperforming separately concatenated speech and audi...

  10. Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Adversarial audio noise, optimized with audio augmentations, can trigger and distort the behavior of audio-based LLMs both digitally and when played through the air.

  11. Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A zero-shot AudioLLM can classify speech as cognitively normal or impaired with accuracy close to supervised baselines, although prompt selection and small margins temper the claim.

  12. RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language

    cs.CL 2025-05 conditional novelty 6.0 of 10

    RAVEN uses query-conditioned token gating plus a new audio-video-sensor QA dataset to improve multimodal question answering, with reported gains of up to 14.5% over prior models.

  13. FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation

    eess.AS 2026-07 conditional novelty 5.5 of 10

    MeanFlow-anchored multi-representation FD post-training improves one-step text-to-audio quality without collapsing multi-step sampling.

  14. Hidden-Domain Routing for All-Type Audio Deepfake Detection

    cs.SD 2026-08 accept novelty 5.0 of 10

    A router-then-specialist audio deepfake detector, which classifies audio type first and then applies type-specific models and thresholds, achieved 96.10% Macro-F1 and first place on AT-ADD Track2.

  15. Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026

    cs.SD 2026-07 conditional novelty 5.0 of 10

    For BirdCLEF+ 2026, a frozen Perch-v2 probe plus a trained HGNetV2-B0 SED net and non-bird prototype heads reach private LB 0.936, while WavTokenizer codec tokens collapse and four general audio transformers lag under...

  16. SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM

    cs.CV 2025-11 conditional novelty 5.0 of 10

    SMART, an audio-enhanced MLLM with shot-aware token compression, reports new state-of-the-art moment retrieval accuracy on Charades-STA and QVHighlights.

  17. AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation

    cs.SD 2025-08 conditional novelty 5.0 of 10

    AudioSet-R reannotates AudioSet with a three-stage Qwen-Audio, Mistral, and DeepSeek R1 pipeline followed by CLAP filtering, and shows consistent mAP gains on AST, PANNs, SSAST, and AudioMAE.

  18. Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation

    cs.MM 2025-08 conditional novelty 5.0 of 10

    A reasoning-based agent (multimodal LLM, then Grounding-DINO, then SAM2) achieves state-of-the-art Referring Audio-Visual Segmentation without pixel-level supervision.

  19. Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    An agentic video-reasoning framework, VITAL, uses tool-based frame sampling, multimodal chain-of-thought, new datasets, and a difficulty-aware RL algorithm to improve long-video QA and temporal grounding.

  20. Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Quantizing an intermediate layer of a pretrained audio model with residual vector quantization, and finetuning the model with task and codebook losses, preserves ASR and audio classification accuracy at bitrates near ...

  21. Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization

    cs.SD 2025-05 reject novelty 5.0 of 10

    A masked-autoencoder SSL model with sparse cross-attention and pretrained audio embeddings reports the best localization error on LuViRA music3 and speech3 while also estimating faulty microphone positions.

  22. LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs

    cs.AI 2025-05 conditional novelty 5.0 of 10

    LiSTEN shows that dynamically selecting a few learnable prompt tokens from a shared pool can replace LoRA fine-tuning for audio-language models, matching or beating it with less training data.

  23. X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance

    cs.SD 2025-05 conditional novelty 5.0 of 10

    X-ARES evaluates 13 audio encoders on 22 speech, sound, and music tasks using linear probing and nearest-neighbor classifiers, revealing strong domain-dependent performance differences.

  24. Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Lower speech-tokenizer frame rates degrade Mandarin ASR far more than English ASR, and signal padding can partly realign the lost tonal information.

  25. Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

    cs.SD 2026-07 conditional novelty 4.5 of 10

    On WiWa, factorised multi-task bird species and call-type classification with adaptive loss balancing improves call-type recognition most consistently, while preferred weighting and adaptation depth depend on backbone...

  26. VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A challenge report showing multi-modal models reach SROCC 0.710 in predicting short-video engagement continuation rate, beating a 0.660 baseline.

  27. A Survey on Video Temporal Grounding with Multimodal Large Language Model

    cs.CV 2025-08 unverdicted novelty 3.0 of 10

    A taxonomized review of video temporal grounding with multimodal large language models, covering model roles, training paradigms, feature processing, benchmarks, and open problems.

Pith tools