An open-source benchmark for speech-to-speech models shows that current systems produce intelligible audio but diverge from human conversational behavior in latency, dialect consistency, emotional entrainment, and prosody.
Gpt-4 technical report,
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
CONDITIONAL 2representative citing papers
ProPS uses a mixture density network conditioned on SBERT text embeddings to generate Gaussian mixture models over speaker x-vectors from natural language profile descriptions.
citing papers explorer
-
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
An open-source benchmark for speech-to-speech models shows that current systems produce intelligible audio but diverge from human conversational behavior in latency, dialect consistency, emotional entrainment, and prosody.
-
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
ProPS uses a mixture density network conditioned on SBERT text embeddings to generate Gaussian mixture models over speaker x-vectors from natural language profile descriptions.