Pith. sign in

REVIEW 1 cited by

Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.06786 v3 pith:E4XW6S3Y submitted 2024-12-09 cs.CV

classification cs.CV
keywords gesturesgenerationgestureapproachproduceretrievalsemanticsemantically
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Non-verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems that can generate rhythmic beat gestures, but struggle to produce semantically meaningful gestures. Therefore, we present RAG-Gesture, a diffusion-based gesture generation approach that leverages Retrieval Augmented Generation (RAG) to produce natural-looking and semantically rich gestures. Our neuro-explicit gesture generation approach is designed to produce semantic gestures grounded in interpretable linguistic knowledge. We achieve this by using explicit domain knowledge to retrieve exemplar motions from a database of co-speech gestures. Once retrieved, we then inject these semantic exemplar gestures into our diffusion-based gesture generation pipeline using DDIM inversion and retrieval guidance at the inference time without any need of training. Further, we propose a control paradigm for guidance, that allows the users to modulate the amount of influence each retrieval insertion has over the generated sequence. Our comparative evaluations demonstrate the validity of our approach against recent gesture generation approaches. The reader is urged to explore the results on our project page.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EVA: Expressive Virtual Avatars from Multi-view Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EVA reconstructs an actor-specific avatar from multi-view video that renders in real time and allows independent control of body motion, hand gestures, and facial expressions through a disentangled two-layer Gaussian ...

Pith tools