Pith. sign in

REVIEW 2 cited by

InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15758 v1 pith:PQEYZWEC submitted 2024-05-24 cs.CV cs.AI

classification cs.CVcs.AI
keywords controlinstructavataravataravatarsemotionaudiofine-grainedgeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the avatar, making the generated video less vivid and controllable. In this paper, we propose a novel text-guided approach for generating emotionally expressive 2D avatars, offering fine-grained control, improved interactivity, and generalizability to the resulting video. Our framework, named InstructAvatar, leverages a natural language interface to control the emotion as well as the facial motion of avatars. Technically, we design an automatic annotation pipeline to construct an instruction-video paired training dataset, equipped with a novel two-branch diffusion-based generator to predict avatars with audio and text instructions at the same time. Experimental results demonstrate that InstructAvatar produces results that align well with both conditions, and outperforms existing methods in fine-grained emotion control, lip-sync quality, and naturalness. Our project page is https://wangyuchi369.github.io/InstructAvatar/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single-DiT portrait animation framework with style and emotion branches plus parallel audio-style cross-attention claims faster, controllable talking-head generation without a Reference Net.

  2. Exploring Timeline Control for Facial Motion Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A diffusion model generates natural facial motions from user-specified multi-track timelines, using TICC-based frame-level action interval annotation for training and evaluation.

Pith tools