Pith. sign in

REVIEW 2 cited by

Reversible Graph Neural Network-based Reaction Distribution Learning for Multiple Appropriate Facial Reactions Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.15270 v3 pith:YJMEEQI7 submitted 2023-05-24 cs.CV

classification cs.CV
keywords facialappropriatereactiondistributionreactionsprocessorgenerationmultiple
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating facial reactions in a human-human dyadic interaction is complex and highly dependent on the context since more than one facial reactions can be appropriate for the speaker's behaviour. This has challenged existing machine learning (ML) methods, whose training strategies enforce models to reproduce a specific (not multiple) facial reaction from each input speaker behaviour. This paper proposes the first multiple appropriate facial reaction generation framework that re-formulates the one-to-many mapping facial reaction generation problem as a one-to-one mapping problem. This means that we approach this problem by considering the generation of a distribution of the listener's appropriate facial reactions instead of multiple different appropriate facial reactions, i.e., 'many' appropriate facial reaction labels are summarised as 'one' distribution label during training. Our model consists of a perceptual processor, a cognitive processor, and a motor processor. The motor processor is implemented with a novel Reversible Multi-dimensional Edge Graph Neural Network (REGNN). This allows us to obtain a distribution of appropriate real facial reactions during the training process, enabling the cognitive processor to be trained to predict the appropriate facial reaction distribution. At the inference stage, the REGNN decodes an appropriate facial reaction by using this distribution as input. Experimental results demonstrate that our approach outperforms existing models in generating more appropriate, realistic, and synchronized facial reactions. The improved performance is largely attributed to the proposed appropriate facial reaction distribution learning strategy and the use of a REGNN. The code is available at https://github.com/TongXu-05/REGNN-Multiple-Appropriate-Facial-Reaction-Generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. REACT 2025: the Third Multiple Appropriate Facial Reaction Generation Challenge

    cs.CV 2025-05 conditional novelty 6.0 of 10

    REACT 2025 presents the MARS dataset of dyadic conversations and benchmark results for multiple appropriate facial reaction generation.

  2. ReactDiff: Latent Diffusion for Facial Reaction Generation

    cs.CV 2025-05 reject novelty 4.0 of 10

    ReactDiff generates multiple listener facial reactions from a speaker's audio and video using a multi-modality transformer with latent diffusion, but its reported benchmark superiority conflicts with its own tables.

Pith tools