A practical multi-camera and multi-microphone system with unified timing architecture for synchronized audio-visual recordings suitable for fine-grained conversation analysis.
Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
dataset 1
citation-polarity summary
years
2026 2roles
dataset 1polarities
use dataset 1representative citing papers
A lightweight transformer predicts iconic gesture placement and intensity from text and emotion alone, outperforming GPT-4o on the BEAT2 dataset for real-time robot deployment.
citing papers explorer
-
A Synchronized Audio-Visual Multi-View Capture System
A practical multi-camera and multi-microphone system with unified timing architecture for synchronized audio-visual recordings suitable for fine-grained conversation analysis.
-
Efficient Emotion-Aware Iconic Gesture Prediction for Robot Co-Speech
A lightweight transformer predicts iconic gesture placement and intensity from text and emotion alone, outperforming GPT-4o on the BEAT2 dataset for real-time robot deployment.