REVIEW 3 cited by
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis. We apply this to two tasks: speech resynthesis and speech-to-speech translation. In both cases, the benchmark evaluates the ability of the model to encode emphasis in the speech input and accurately reproduce it in the output, potentially across a change of speaker and language. As part of the evaluation pipeline, we introduce EmphaClass, a new model that classifies emphasis at the frame or word level.
Forward citations
Cited by 3 Pith papers
-
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
A speech language model trained only on audiobooks with explicit word-level prosody tokens displays emerging abilities in prosody-controlled generation, emphasis and emotion understanding, and prosodic consistency acr...
-
WHISTRESS: Enriching Transcriptions with Sentence Stress Detection
WHISTRESS extends Whisper with a token-level stress classifier trained on a new synthetic dataset, and shows zero-shot transfer to natural speech benchmarks.
-
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
Quantizing speech at frame, phone, word, and utterance levels preserves emotion and prominence better than frame-only discrete units at similar bitrates.
Discussion (0). Continue with ORCID to comment.