Pith. sign in

eess.AS

Audio and Speech Processing

Theory and methods for processing signals representing audio, speech, and language, and their applications. This includes analysis, synthesis, enhancement, transformation, classification and interpretation of such signals as well as the design, development, and evaluation of associated signal processing systems. Machine learning and pattern analysis applied to any of the above areas is also welcome. Specific topics of interest include: auditory modeling and hearing aids; acoustic beamforming and source localization; classification of acoustic scenes; speaker separation; active noise control and echo cancellation; enhancement; de-reverberation; bioacoustics; music signals analysis, synthesis and modification; music information retrieval; audio for multimedia and joint audio-video processing; spoken and written language modeling, segmentation, tagging, parsing, understanding, and translation; text mining; speech production, perception, and psychoacoustics; speech analysis, synthesis, and perceptual modeling and coding; robust speech recognition; speaker recognition and characterization; deep learning, online learning, and graphical models applied to speech, audio, and language signals; and implementation aspects ranging from system architecture to fast algorithms.

Papers reviewed in the last 7 days lead, then the papers readers actually read. Ranking is not a quality score.

sort pith recommended most recent

Distillation cuts hallucinations in LM-based speech enhancement

Noise-invariant acoustic-semantic representations from clean targets improve linguistic accuracy under severe noise and reverberation.

· “Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation”

open re-runnable review →
Figure from the paper

Transformer reconstructs room impulse responses from sparse mics

Sinusoidal position encoding plus separate early-late decoding yields lower error than baselines in simulations across missing rates and lay

· “RIR-Former: Coordinate-Guided Transformer for Continuous Reconstruction of Room Impulse Responses”

open re-runnable review →

Makes early-Parkinson's speech detection comparable via fixed protocol

A fixed speaker-independent split on two public datasets gives all models the same 31 early-stage speakers to classify.

· “A Benchmark for Early-stage Parkinson's Disease Detection from Speech”

open re-runnable review →

Slimmable diffusion cuts speech-enhancement compute by 87.5%

Adaptive network width across denoising steps keeps PESQ within 0.08 of the full model at one-eighth the cost.

· “SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks”

open re-runnable review →
Figure from the paper

Higher-fidelity room acoustics improve speech enhancement training

On real measured rooms, hybrid-simulated training data cut word error rates more than image-source data.

· “Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement”

open re-runnable review →

Deepfake speech traced to its generator with 99.64% accuracy

The same network ranks which synthesizer components drove each attribution, with no post-hoc explainer.

· “Explainability by Design: Structured Kolmogorov-Arnold Networks over Probabilistic Attributes for Speech Deepfake Source Tracing”

open re-runnable review →
Figure from the paper

Whisper learns Mizo in under 18 hours of speech data

A fine-tuned multilingual model reaches 18.08% word error, and 7.22% when Mizo's spacing quirks are forgiven.

· “A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation”

open re-runnable review →
Figure from the paper

Higher quality scores don't mean better ASR on rtMRI audio

Across five rtMRI corpora, Denoiser raised word-error rates in 13 of 15 comparisons while RE-USE lowered them in 11.

· “Navigating Speech Enhancement for Real-Time MRI: A Systematic Assessment of Signal Quality, Source Preservation, and Downstream Tasks”

open re-runnable review →
Figure from the paper

Label-free metrics beat trained fusion at picking music AI layers

Across 12 models and 15 tasks, geometric layer ranking matches or beats multi-layer fusion, especially with few labels.

· “What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models”

open re-runnable review →

Deepfake recall tracks proximity to unseen TTS

Across 22 Indic languages, training on four TTS systems raises out-of-domain synthetic recall from 7% to 51%.

· “Evaluating Pre-trained Speech Encoders for Spontaneous Speech Detection and Out of Domain Synthetic Speech Generalisation in Indic Languages”

open re-runnable review →
Figure from the paper

Smart-glasses speech test: overlap and acoustic cues stump AI

106-hour egocentric Mandarin corpus pairs who-said-what with understanding; overlap and audio-only questions stay hard.

· “The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models”

open re-runnable review →
Figure from the paper

Prompt compiler lifts a music AI into Arabic quarter tones

Live performance is compiled into prompts that steer a frozen generator, measurably raising quarter-tone content.

· “MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model”

open re-runnable review →
Figure from the paper

Training-free speaker alignment for long recordings: 98.5%

A plug-and-play post-processing step keeps separated speakers consistent, reaching near-oracle SI-SDR in dense and sparse mixtures.

· “Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation”

open re-runnable review →

One network locates sound on microphone arrays it has never seen

Trained once with simulated array transfer functions, it stays accurate on phone-style devices in reverberant rooms.

· “Neural Array-Generic Direction-of-Arrival Estimation Exploiting Array Transfer Functions”

open re-runnable review →

Automatic diarization preserves proficiency signals in interviews

How much respondents talk and how often they use the intended language predict proficiency ratings even when labels are automatic.

· “Speaker Role and Language Diarization for Analyzing Multilingual Interviews for Language Proficiency of Older Adults”

open re-runnable review →
Figure from the paper

browse all of eess.AS → full archive · search · sub-categories