REVIEW 6 cited by
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
read the original abstract
A demonstration of a real-time and continuous turn-taking prediction system is presented. The system is based on a voice activity projection (VAP) model, which directly maps dialogue stereo audio to future voice activities. The VAP model includes contrastive predictive coding (CPC) and self-attention transformers, followed by a cross-attention transformer. We examine the effect of the input context audio length and demonstrate that the proposed system can operate in real-time with CPU settings, with minimal performance degradation.
Forward citations
Cited by 6 Pith papers
-
Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot
Context-aware prefaces generated before a user finishes speaking reduce the delay to the main answer, but make the first response slightly later than a fixed filler.
-
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue
TurnNat introduces a likelihood-based automatic evaluation method for turn-taking naturalness in dyadic spoken dialogues using a causal prediction model and a human-validated perturbation benchmark.
-
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue
A causal future voice-activity model’s negative log-likelihood, pooled over turn-taking boundary units, ranks natural dyadic clips above five classes of controlled timing perturbations.
-
Toward Signing Activity Projection in Sign Language Interaction
Initial adaptation of Voice Activity Projection to dyadic sign language interaction on the Public DGS Corpus shows SHIFT/HOLD prediction is feasible with hand cues while SHIFT prediction remains difficult.
-
Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders
Pretrained audio-visual speech encoders adapted with LoRA improve multimodal voice activity projection for turn-taking prediction across multiple languages and a robot mediation corpus.
-
Endpoint Anticipation for Low-Latency Spoken Dialogue
A speech-based model forecasts conversation turn endpoints up to 2.56 seconds ahead to enable lower-latency spoken dialogue via speculative LLM and TTS execution.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.