One stem at a time beats single-pass mixing
Each stem is blended into a fixed submix, beating parallel baselines and enabling interactive control.
· “Rethinking Automatic Music Mixing as Sequential Stem Blending”
Audio and Speech Processing
Theory and methods for processing signals representing audio, speech, and language, and their applications. This includes analysis, synthesis, enhancement, transformation, classification and interpretation of such signals as well as the design, development, and evaluation of associated signal processing systems. Machine learning and pattern analysis applied to any of the above areas is also welcome. Specific topics of interest include: auditory modeling and hearing aids; acoustic beamforming and source localization; classification of acoustic scenes; speaker separation; active noise control and echo cancellation; enhancement; de-reverberation; bioacoustics; music signals analysis, synthesis and modification; music information retrieval; audio for multimedia and joint audio-video processing; spoken and written language modeling, segmentation, tagging, parsing, understanding, and translation; text mining; speech production, perception, and psychoacoustics; speech analysis, synthesis, and perceptual modeling and coding; robust speech recognition; speaker recognition and characterization; deep learning, online learning, and graphical models applied to speech, audio, and language signals; and implementation aspects ranging from system architecture to fast algorithms.
sort pith recommended most recent
Each stem is blended into a fixed submix, beating parallel baselines and enabling interactive control.
· “Rethinking Automatic Music Mixing as Sequential Stem Blending”
Noise-invariant acoustic-semantic representations from clean targets improve linguistic accuracy under severe noise and reverberation.
Scaled to hundreds of billions of parameters with over 100 million hours of audio-visual training data for long-context and multilingual use
Sinusoidal position encoding plus separate early-late decoding yields lower error than baselines in simulations across missing rates and lay
· “RIR-Former: Coordinate-Guided Transformer for Continuous Reconstruction of Room Impulse Responses”
Worst-case log-likelihood leakage is the metric that actually tracks Shannon perfect secrecy for anonymized speech.
A fixed speaker-independent split on two public datasets gives all models the same 31 early-stage speakers to classify.
· “A Benchmark for Early-stage Parkinson's Disease Detection from Speech”
Dual-brain retrieval hides inside the silence gap, adding facts and empathy without slowing the conversation down.
· “VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction”
The same monitor flags the deviations behind two fatal accidents, seconds before impact.
External AUCs often drop below 0.6 even when internal scores hit 0.75, pointing to device bias.
Server and client agents jointly learn aggregation weights and personalization without pretraining.
VAD-guided emotional TTS beats four baselines and two commercial systems on blind preference, with no added latency.
· “EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis”
Natural mixtures, synthetic remixes, and degraded video are scored by quality, intelligibility, and speaker identity.
· “The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge”
No G2P dictionaries, no GANs — syllable units beat phone-level baselines across languages.
The clean-text champion loses the most under accented speech; entity corruption drives the failures.
· “Better Retrieval, Worse Robustness: How Multi-hop RAG Amplifies Upstream ASR Errors”
DiaScriber reports lower diarization and word errors than two end-to-end baselines across diverse test sets.
· “DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios”
A diarization model flags leaked words and deletes them; the gains are biggest on meeting audio.
Extracts 28 expression-text types from MusicXML; joint model scores 2.64 vs 2.80 note-only baseline.
· “MusPyExpress: Extending MusPy with Enhanced Expression Text Support”
Age-free acoustic models keep 0.717 AUC on cohorts where raw age and gender are at chance.
Per-stream phrase lists run inside one GPU batch with no slowdown, enabling personalized speech recognition.
· “TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems”
Adaptive network width across denoising steps keeps PESQ within 0.08 of the full model at one-eighth the cost.
· “SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks”
Integer-only at 28 MMACs, and a match for low-complexity rivals; hearable-grade hardware suffices.
On real measured rooms, hybrid-simulated training data cut word error rates more than image-source data.
Rejecting just 8% of least-confident predictions lifts retained macro-F1 on rāga ID from 0.89 to 0.98.
· “TCP_α: Margin-Controlled Confidence estimation for reliable Music Information Retrieval”
The same network ranks which synthesizer components drove each attribution, with no post-hoc explainer.
Even warned IT professionals fall to 0.48 F1 on 2024 clones and 9 percent on partial spoofs.
· “Tracking the Trend in How Speech Synthesizers Deceive People”
Distilled from CAV-MAE, the 14M-parameter DAVSS matches or tops the 165M teacher by using finer patches and deeper joint layers.
A fine-tuned multilingual model reaches 18.08% word error, and 7.22% when Mizo's spacing quirks are forgiven.
Trained on 35 European bats, ChiroEcho covers 41 of 48 native species by resolving genus predictions against location.
· “ChiroEcho: extending automated bat vocalisation classification beyond the learned taxonomy”
iRDT uses only adds and shifts, runs about ten times faster, and hits 94.7% accuracy on 12 speech commands.
· “A Multiplication-Free Feature Extractor for Signal Classification: Keyword Spotting Case Study”
Tuning neurons found from voices shifts facial emotion recognition too — the same sparse units work both ways.
· “Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models”
3D-scan head-related transfer functions perform as well as anechoic measurements, while a generic mannequin HRTF lags.
· “Numerical and perceptual validity of synthetic Head-Related Transfer Functions at scale”
Post-training with GRPO lets an 8B audio model find creator-authored boundaries across speech, music, and gaming.
· “Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization”
Clock and data in separate channels modestly improve error-class separation, a first step toward audio fault diagnosis.
· “Sonifying I2S Transport Signals to Detect Transmission Faults”
Across five rtMRI corpora, Denoiser raised word-error rates in 13 of 15 comparisons while RE-USE lowered them in 11.
A modified FxLMS update uses past, current, and predicted future samples; neural forecasts match perfect ones in tests.
Training-free cached LLM probabilities cut Whisper-small WER/CER by 8.13% absolute, with no gains past a 32-token context.
In low-data settings, iterative self-labeling improves emotion and prominence control and beats single-pass labeling.
· “Iterative Self-Learning for Expressive Text-to-Speech Synthesis”
Stroke counts match the 62–67% band seen for word lengths; the redundancy buys a readable 2D script.
On the small Elephant Voices dataset, one or two labelled calls per type beat models trained on all its labels.
· “A Parameter-Free Few-Shot Evaluation for Elephant Vocalisation Classification”
Across 12 models and 15 tasks, geometric layer ranking matches or beats multi-layer fusion, especially with few labels.
· “What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models”
Separating audio first, then matching faces to voices with CLIP, beats the baseline on every reported metric.
A brief audio sample steers a neural mask to isolate one voice in a duet.
· “Singer-Informed Vocal Source Separation for Multi-Singer Music Mixtures”
A device-agnostic prior plus measurement-consistent sampling beats linear and neural encoders in metrics and listening tests.
· “Ambisonics Encoding of Room Impulse Responses using a Device-Agnostic Diffusion Model”
Six hours of duo improvisation, annotated from both sides, show perception lags intention.
· “H2H Music Improv: A Communication Model and Audio-Visual Dataset for Music Improvisation”
Ten iconic composers, 17,053 chords, 23,447 notes—enough to test what defines MPB's shared style.
Always-on TTS decoder keeps near-offline quality and stops within 320 ms of barge-in.
· “VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents”
A causal LSTM trained only on real speech cuts the equal error rate from 22.9% to 5.7% on MLAAD-EN.
· “Trajectory Dynamics in Self-Supervised Learning Latent Space for Audio Deepfake Detection”
A fixed offline teacher's labels lift cache-aware streaming WER on four financial and call-center datasets.
· “StreamHear: Domain-Adapted Pseudo-Labeling for Semi-Supervised Streaming Speech Recognition”
Models trained for Parkinson's separate patients from healthy speakers but cannot pick out Parkinson's from dementia.
CASA splits delivery and content, hitting RMSE 0.358 on Speak & Improve with half the inference parameters.
· “CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model”
Across 22 Indic languages, training on four TTS systems raises out-of-domain synthetic recall from 7% to 51%.
A flow-matching transformer with per-token time conditioning reports the best prompt alignment and lowest word error among compared systems.
A non-autoregressive model over continuous audio-codec features tops six paradigms; fine-tuning lifts all of them.
106-hour egocentric Mandarin corpus pairs who-said-what with understanding; overlap and audio-only questions stay hard.
A student ASR model learns from its own decoded prefixes while a frozen teacher keeps Mandarin error rates flat.
· “On-Policy Self-Distillation for Multi-Dialect ASR: Mastering Dialects, Retaining Mandarin”
Unified speech, music, and effects generation now approaches dedicated TTS intelligibility.
· “MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching”
Supervised models replace covariance for multi-source rooms and improve speech-enhancement scores.
Diffusion TTS matches autoregressive quality; streaming variant opens audio in 41.6 ms.
A family-aware speaker token plus smoothing helps one model label daylong recordings across four vocalization tiers.
Frozen embeddings from score-conditioned JEPA training beat separate baselines on quality, ranking, technique, and mistake tasks.
· “MAJEPPA: Morphing and Assessing in a Unified Piano Performance Space”
Live performance is compiled into prompts that steer a frozen generator, measurably raising quarter-tone content.
A binaural model that uses target direction and voice-activity gating outperforms MVDR and BCCTN on SPEAR glass-array recordings.
· “BiTSE: Binaural Target Speaker Extraction in Noisy Multi-Talker Environments for AR Glass Arrays”
Fundamental-frequency tracking error drops from 1.4 to 0.8 kHz; classifier F1 rises from 83 to 89 percent.
· “Training Set Synthesis for Bioacoustic Denoising: A Case Study With Mice”
Adding PhonoQ to WavLM or HuBERT improves macro-F1 for phonemes and phonological features across unseen speech and speakers.
· “Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification”
A plug-and-play post-processing step keeps separated speakers consistent, reaching near-oracle SI-SDR in dense and sparse mixtures.
· “Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation”
Trained once with simulated array transfer functions, it stays accurate on phone-style devices in reverberant rooms.
· “Neural Array-Generic Direction-of-Arrival Estimation Exploiting Array Transfer Functions”
How much respondents talk and how often they use the intended language predict proficiency ratings even when labels are automatic.
Supervising loudness cues instead of waveforms beats zero-shot transfer on guitar recordings.
· “Beyond Piano: Cross-Instrument MIDI Velocity Estimation via Differentiable SoundFont Proxies”
Error-matched corruption during training teaches the music renderer to complete, not parrot, the language model's plan.
· “Beyond Reconstruction: Full-Context Generative DiT for Music Generation”
Deactivating or steering these shared neurons shifts emotion recognition in languages never used for identification.
· “Multilingual Emotion Neurons in Large Audio-Language Models”
Concept edits like 'more piano' or 'no guitar' work better when supports are recovered by geometry, not wording.
· “Steering dense music retrieval with open-vocabulary concept discovery”
A new design framework matches each audio latent's dependency horizon and ambiguity to its generator instead.
On EnCodec, PESQ gains 0.088–0.151 across rates; MOS rises from 3.449 to 3.780.
· “BAMU: Bitstream-Aware Marginal-Utility Allocation for Frozen Pretrained Neural Speech Codecs”
Phone-aligned prosody cues cut pitch RMSE ~44% and duration MAE by more than half.
· “CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis”
A single-codebook codec keeps self-supervised phoneme structure, adds acoustic detail, and moves the 650/800 bps frontier.
· “ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure”
Tribal and low-resource Indic languages get an open transcription system for the first time.
· “SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages”
It beats the previous neural baseline and approaches a top commercial virtual instrument in clarity and dynamics.
· “VIOLET: High-Fidelity Violin Synthesis with Techniques and Dynamics”
Per-segment mixing, guided by quality and speech-preservation scores, lifts PESQ and STOI and cuts ASR word errors.
Semantic-token supervision during training lowers WER/CER in TTS and singing while inference stays fully continuous.
· “SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation”
Delayed server outputs and fused spatial statistics give a tiny edge model most of a large server's benefit.
· “Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement”