Pith. sign in

REVIEW 1 cited by

MPE4G: Multimodal Pretrained Encoder for Co-Speech Gesture Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.15740 v1 pith:OU7X4XNF submitted 2023-05-25 cs.CV cs.AI

classification cs.CVcs.AI
keywords gesturesco-speechmodalitiesencodergenerationgestureinputmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some input modalities are removed or contain noise, the model may not generate the gestures properly. To acquire robust and generalized encodings, we propose a novel framework with a multimodal pre-trained encoder for co-speech gesture generation. In the proposed method, the multi-head-attention-based encoder is trained with self-supervised learning to contain the information on each modality. Moreover, we collect full-body gestures that consist of 3D joint rotations to improve visualization and apply gestures to the extensible body model. Through the series of experiments and human evaluation, the proposed method renders realistic co-speech gestures not only when all input modalities are given but also when the input modalities are missing or noisy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Navigation Framework Utilizing Vision-Language Models

    cs.RO 2025-06 reject novelty 3.0 of 10

    A frozen Qwen2.5-VL model with prompt engineering and a two-frame egocentric view achieves 5% success on 20 R2R VLN-CE episodes, near the zero-movement baseline.

Pith tools