REVIEW 4 cited by
INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Imagine having a conversation with a socially intelligent agent. It can attentively listen to your words and offer visual and linguistic feedback promptly. This seamless interaction allows for multiple rounds of conversation to flow smoothly and naturally. In pursuit of actualizing it, we propose INFP, a novel audio-driven head generation framework for dyadic interaction. Unlike previous head generation works that only focus on single-sided communication, or require manual role assignment and explicit role switching, our model drives the agent portrait dynamically alternates between speaking and listening state, guided by the input dyadic audio. Specifically, INFP comprises a Motion-Based Head Imitation stage and an Audio-Guided Motion Generation stage. The first stage learns to project facial communicative behaviors from real-life conversation videos into a low-dimensional motion latent space, and use the motion latent codes to animate a static image. The second stage learns the mapping from the input dyadic audio to motion latent codes through denoising, leading to the audio-driven head generation in interactive scenarios. To facilitate this line of research, we introduce DyConv, a large scale dataset of rich dyadic conversations collected from the Internet. Extensive experiments and visualizations demonstrate superior performance and effectiveness of our method. Project Page: https://grisoon.github.io/INFP/.
Forward citations
Cited by 4 Pith papers
-
Multi-human Interactive Talking Dataset
The paper contributes a 12-hour multi-person conversational video dataset with pose and speaking annotations, plus a baseline model for generating full-body talking videos of two to four people.
-
Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System
A VAP-based model predicts the timing and type (short, long, long_p) of listener nodding in real time, and multi-task learning with backchannel prediction plus pretraining improves type prediction.
-
ARIG: Autoregressive Interactive Head Generation for Real-time Conversations
ARIG introduces a real-time, frame-wise autoregressive head generation framework with diffusion-based continuous motion prediction, improving interactive realism over clip-wise methods.
-
LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
Using consistency distillation, INT8 quantization, and pipeline parallelism, the LLIA system generates portrait video from audio at 78 FPS, with 140 ms initial latency.
Discussion (0). Continue with ORCID to comment.