An open-source benchmark for speech-to-speech models shows that current systems produce intelligible audio but diverge from human conversational behavior in latency, dialect consistency, emotional entrainment, and prosody.
Robust speech recognition via large-scale weak supervi- sion
7 Pith papers cite this work. Polarity classification is still indexing.
years
2026 7representative citing papers
PnP reformulates adversarial purification as learning positive-incentive noise to defend speaker verification against attacks with high efficiency and limited impact on genuine utterances.
LuxEmo is a new 21-hour conversational expressive speech corpus for Luxembourgish with 4 emotion categories, created via semi-automatic curation from RTL broadcasts and used to benchmark five TTS systems.
S2ST-Omni 2 uses typology-informed hierarchical encoding, gated Dual-CTC, and typology-aware prompting to improve multilingual S2ST over flat-label baselines on CVSS-C, with gains in low-data regimes.
UniPASE extends the low-hallucination PASE framework to universal speech enhancement, restoring seven distortion types at flexible sampling rates with better word-error and speaker-similarity scores than prior generative systems.
A survey proposing a three-pillar framework to evaluate LLMs as tools for measuring latent psychological constructs and reviewing applications in personality and mental health.
HQTN-SER combines a low-parameter quantum tensor network module with classical latent embeddings to reach 73-80% accuracy on three speech emotion datasets while using few qubits and showing stable training.
citing papers explorer
-
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
An open-source benchmark for speech-to-speech models shows that current systems produce intelligible audio but diverge from human conversational behavior in latency, dialect consistency, emotional entrainment, and prosody.
-
Positive-Incentive Noise Predictor for Adversarial Purification in Speaker Verification
PnP reformulates adversarial purification as learning positive-incentive noise to defend speaker verification against attacks with high efficiency and limited impact on genuine utterances.
-
LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish
LuxEmo is a new 21-hour conversational expressive speech corpus for Luxembourgish with 4 emotion categories, created via semi-automatic curation from RTL broadcasts and used to benchmark five TTS systems.
-
From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation
S2ST-Omni 2 uses typology-informed hierarchical encoding, gated Dual-CTC, and typology-aware prompting to improve multilingual S2ST over flat-label baselines on CVSS-C, with gains in low-data regimes.
-
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
UniPASE extends the low-hallucination PASE framework to universal speech enhancement, restoring seven distortion types at flexible sampling rates with better word-error and speaker-similarity scores than prior generative systems.
-
A Survey of Large Language Models for Perception and Measurement of Human Psychology
A survey proposing a three-pillar framework to evaluate LLMs as tools for measuring latent psychological constructs and reviewing applications in personality and mental health.
-
HQTN-SER: Speech Emotion Recognition with Hybrid Quantum Tensor Networks
HQTN-SER combines a low-parameter quantum tensor network module with classical latent embeddings to reach 73-80% accuracy on three speech emotion datasets while using few qubits and showing stable training.