NVBench provides a standardized bilingual benchmark and evaluation protocol for assessing non-verbal vocalization generation, placement, and salience in text-to-speech systems.
of of the idea that has been the same idea for a thousand years that they believe that—
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6representative citing papers
FCaps supplies 19M fine-grained speech style captions on 47k hours of audio via direct grounding, enabling the CLSP model to produce multi-granular representations that improve retrieval, zero-shot classification, and style scoring aligned with human judgments.
TRACE detects synthetic emotional entrainment disruption in dyadic speech at up to 93.47% accuracy when conditioned on relationship, using windowed emotion-Whisper sequences on the new DyadEE dataset.
Foley-Omni extends isolated audio synthesis to joint generation of full video soundtracks across speech, effects, and music, with a new V2ST-Bench for evaluation showing competitive single-task results and gains in mixed-track consistency.
ProPS uses a mixture density network conditioned on SBERT text embeddings to generate Gaussian mixture models over speaker x-vectors from natural language profile descriptions.
VoxCPM2 scales hierarchical continuous-latent speech modeling to 2B parameters and over 2M hours of multilingual data, unifying voice cloning, style control, and continuation in one backbone with open release.
citing papers explorer
-
NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations
NVBench provides a standardized bilingual benchmark and evaluation protocol for assessing non-verbal vocalization generation, placement, and salience in text-to-speech systems.
-
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
FCaps supplies 19M fine-grained speech style captions on 47k hours of audio via direct grounding, enabling the CLSP model to produce multi-granular representations that improve retrieval, zero-shot classification, and style scoring aligned with human judgments.
-
TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech
TRACE detects synthetic emotional entrainment disruption in dyadic speech at up to 93.47% accuracy when conditioned on relationship, using windowed emotion-Whisper sequences on the new DyadEE dataset.
-
Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation
Foley-Omni extends isolated audio synthesis to joint generation of full video soundtracks across speech, effects, and music, with a new V2ST-Bench for evaluation showing competitive single-task results and gains in mixed-track consistency.
-
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
ProPS uses a mixture density network conditioned on SBERT text embeddings to generate Gaussian mixture models over speaker x-vectors from natural language profile descriptions.
-
VoxCPM2 Technical Report
VoxCPM2 scales hierarchical continuous-latent speech modeling to 2B parameters and over 2M hours of multilingual data, unifying voice cloning, style control, and continuation in one backbone with open release.