REVIEW 27 cited by
BEATs: Audio Pre-Training with Acoustic Tokenizers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The massive growth of self-supervised learning (SSL) has been witnessed in language, vision, speech, and audio domains over the past few years. While discrete label prediction is widely adopted for other modalities, the state-of-the-art audio SSL models still employ reconstruction loss for pre-training. Compared with reconstruction loss, semantic-rich discrete label prediction encourages the SSL model to abstract the high-level audio semantics and discard the redundant details as in human perception. However, a semantic-rich acoustic tokenizer for general audio pre-training is usually not straightforward to obtain, due to the continuous property of audio and unavailable phoneme sequences like speech. To tackle this challenge, we propose BEATs, an iterative audio pre-training framework to learn Bidirectional Encoder representation from Audio Transformers, where an acoustic tokenizer and an audio SSL model are optimized by iterations. In the first iteration, we use random projection as the acoustic tokenizer to train an audio SSL model in a mask and label prediction manner. Then, we train an acoustic tokenizer for the next iteration by distilling the semantic knowledge from the pre-trained or fine-tuned audio SSL model. The iteration is repeated with the hope of mutual promotion of the acoustic tokenizer and audio SSL model. The experimental results demonstrate our acoustic tokenizers can generate discrete labels with rich audio semantics and our audio SSL models achieve state-of-the-art results across various audio classification benchmarks, even outperforming previous models that use more training data and model parameters significantly. Specifically, we set a new state-of-the-art mAP 50.6% on AudioSet-2M for audio-only models without using any external data, and 98.1% accuracy on ESC-50. The code and pre-trained models are available at https://aka.ms/beats.
Forward citations
Cited by 27 Pith papers
-
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
Jointly training audio and video VAEs with segment contrastive loss and semantic distillation yields more learnable, cross-aligned latents that improve downstream joint generation quality and sync.
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection
Better proxy-task performance does not generally improve anomalous sound detection; only source separation showed a strong, consistent positive correlation.
-
EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
EchoingPixels prunes audio-visual LLM input tokens jointly across modalities and re-tunes RoPE frequencies so that 5–20% of tokens retain roughly full-model performance.
-
Assessing Factual Music Comprehension in Large Audio Language Models
Standard NLP metrics fail to capture factual music understanding in audio-language models; a CLAP-based metric and an LLM-parsed factual QA protocol measure it more directly.
-
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.
-
NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations
A Mandarin speech pipeline that tags non-word vocalizations (laughter, breath, filled pauses) for ASR and controls their generation in TTS, backed by a claimed 573-hour, 174,179-utterance word-level annotated corpus.
-
Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting
Benchmarks seven open-source audio-video-text MLLMs on six affective datasets and shows a generative-knowledge prompting step improves fine-tuned emotion recognition.
-
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
A single multimodal language model can generate intelligible speech and synchronized background audio jointly from a silent video, transcript, and reference voice, outperforming separately concatenated speech and audi...
-
Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World
Adversarial audio noise, optimized with audio augmentations, can trigger and distort the behavior of audio-based LLMs both digitally and when played through the air.
-
Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
A zero-shot AudioLLM can classify speech as cognitively normal or impaired with accuracy close to supervised baselines, although prompt selection and small margins temper the claim.
-
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
RAVEN uses query-conditioned token gating plus a new audio-video-sensor QA dataset to improve multimodal question answering, with reported gains of up to 14.5% over prior models.
-
FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation
MeanFlow-anchored multi-representation FD post-training improves one-step text-to-audio quality without collapsing multi-step sampling.
-
Hidden-Domain Routing for All-Type Audio Deepfake Detection
A router-then-specialist audio deepfake detector, which classifies audio type first and then applies type-specific models and thresholds, achieved 96.10% Macro-F1 and first place on AT-ADD Track2.
-
Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026
For BirdCLEF+ 2026, a frozen Perch-v2 probe plus a trained HGNetV2-B0 SED net and non-bird prototype heads reach private LB 0.936, while WavTokenizer codec tokens collapse and four general audio transformers lag under...
-
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
SMART, an audio-enhanced MLLM with shot-aware token compression, reports new state-of-the-art moment retrieval accuracy on Charades-STA and QVHighlights.
-
AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation
AudioSet-R reannotates AudioSet with a three-stage Qwen-Audio, Mistral, and DeepSeek R1 pipeline followed by CLAP filtering, and shows consistent mAP gains on AST, PANNs, SSAST, and AudioMAE.
-
Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation
A reasoning-based agent (multimodal LLM, then Grounding-DINO, then SAM2) achieves state-of-the-art Referring Audio-Visual Segmentation without pixel-level supervision.
-
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
An agentic video-reasoning framework, VITAL, uses tool-based frame sampling, multimodal chain-of-thought, new datasets, and a difficulty-aware RL algorithm to improve long-video QA and temporal grounding.
-
Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
Quantizing an intermediate layer of a pretrained audio model with residual vector quantization, and finetuning the model with task and codebook losses, preserves ASR and audio classification accuracy at bitrates near ...
-
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
A masked-autoencoder SSL model with sparse cross-attention and pretrained audio embeddings reports the best localization error on LuViRA music3 and speech3 while also estimating faulty microphone positions.
-
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
LiSTEN shows that dynamically selecting a few learnable prompt tokens from a shared pool can replace LoRA fine-tuning for audio-language models, matching or beating it with less training data.
-
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
X-ARES evaluates 13 audio encoders on 22 speech, sound, and music tasks using linear probing and nearest-neighbor classifiers, revealing strong domain-dependent performance differences.
-
Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English
Lower speech-tokenizer frame rates degrade Mandarin ASR far more than English ASR, and signal padding can partly realign the lost tonal information.
-
Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types
On WiWa, factorised multi-task bird species and call-type classification with adaptive loss balancing improves call-type recognition most consistently, while preferred weighting and adaptation depth depend on backbone...
-
VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
A challenge report showing multi-modal models reach SROCC 0.710 in predicting short-video engagement continuation rate, beating a 0.660 baseline.
-
A Survey on Video Temporal Grounding with Multimodal Large Language Model
A taxonomized review of video temporal grounding with multimodal large language models, covering model roles, training paradigms, feature processing, benchmarks, and open problems.
Discussion (0). Sign in to comment.