REVIEW 2 cited by
How Much Does Prosody Help Turn-taking? Investigations using Voice Activity Projection Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Turn-taking is a fundamental aspect of human communication and can be described as the ability to take turns, project upcoming turn shifts, and supply backchannels at appropriate locations throughout a conversation. In this work, we investigate the role of prosody in turn-taking using the recently proposed Voice Activity Projection model, which incrementally models the upcoming speech activity of the interlocutors in a self-supervised manner, without relying on explicit annotation of turn-taking events, or the explicit modeling of prosodic features. Through manipulation of the speech signal, we investigate how these models implicitly utilize prosodic information. We show that these systems learn to utilize various prosodic aspects of speech both on aggregate quantitative metrics of long-form conversations and on single utterances specifically designed to depend on prosody.
Forward citations
Cited by 2 Pith papers
-
Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals
A multi-modal (text, audio, video) model trained on a newly collected 210-hour conversation dataset predicts turn-taking and backchannel actions with F1 about 0.81 and 0.91.
-
Real-Time Textless Dialogue Generation
A streaming, textless dialogue model predicts turn-taking actions every 160 ms and generates speech units, improving naturalness over cascaded systems at the cost of lower semantic coherence.
Discussion (0). Continue with ORCID to comment.