ProVoice-Bench is the first framework to evaluate proactive voice agents, revealing that state-of-the-art multimodal LLMs struggle with over-triggering and context-aware reasoning.
ESC: Dataset for Environmental Sound Classification
7 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
A model-free diffusion test for discrete time series that uses the scaling of excursion counts with quadratic variation to classify signals as stochastic or deterministic.
ZEBRA reduces the base-to-novel generalization gap in audio-language models by fusing zero-shot and prompt-learning logits with entropy regularization.
FoleySet is a Creative Commons dataset of 10,000 audio clips with two-level human annotations for Foley sound classification, retrieval, and generation.
DeePen demonstrates that both production and academic audio deepfake detectors can be reliably deceived by simple signal processing attacks such as time-stretching or echo addition, with some attacks resistible via retraining and others remaining effective.
MLAAD provides a large-scale multi-language synthetic audio dataset for training and evaluating audio anti-spoofing models, showing better training performance than InTheWild and FakeOrReal and alternating superiority with ASVspoof 2019 across eight test sets.
An automatic audio annotation pipeline using BEATs and CLAP filtering produces a 2130-hour dataset that yields 3.97% average accuracy gains on three domestic audio classification tasks.
citing papers explorer
-
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
ProVoice-Bench is the first framework to evaluate proactive voice agents, revealing that state-of-the-art multimodal LLMs struggle with over-triggering and context-aware reasoning.
-
Detecting Stochasticity in Discrete Signals via Nonparametric Excursion Theorem
A model-free diffusion test for discrete time series that uses the scaling of excursion counts with quadratic variation to classify signals as stochastic or deterministic.
-
ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models
ZEBRA reduces the base-to-novel generalization gap in audio-language models by fusing zero-shot and prompt-learning logits with entropy regularization.
-
FoleySet: A Multi-Level Human-Annotated Foley Sound Dataset
FoleySet is a Creative Commons dataset of 10,000 audio clips with two-level human annotations for Foley sound classification, retrieval, and generation.
-
DeePen: Penetration Testing for Audio Deepfake Detection
DeePen demonstrates that both production and academic audio deepfake detectors can be reliably deceived by simple signal processing attacks such as time-stretching or echo addition, with some attacks resistible via retraining and others remaining effective.
-
MLAAD: The Multi-Language Audio Anti-Spoofing Dataset
MLAAD provides a large-scale multi-language synthetic audio dataset for training and evaluating audio anti-spoofing models, showing better training performance than InTheWild and FakeOrReal and alternating superiority with ASVspoof 2019 across eight test sets.
-
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios
An automatic audio annotation pipeline using BEATs and CLAP filtering produces a 2130-hour dataset that yields 3.97% average accuracy gains on three domestic audio classification tasks.