OPEN releases landmark and feature data from 35 hours of older adult virtual rehab sessions with engagement, affect, behavior, and context annotations, plus baselines reaching up to 81% accuracy.
Class-attention Video Transformer for Engagement Intensity Prediction
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In order to deal with variant-length long videos, prior works extract multi-modal features and fuse them to predict students' engagement intensity. In this paper, we present a new end-to-end method Class Attention in Video Transformer (CavT), which involves a single vector to process class embedding and to uniformly perform end-to-end learning on variant-length long videos and fixed-length short videos. Furthermore, to address the lack of sufficient samples, we propose a binary-order representatives sampling method (BorS) to add multiple video sequences of each video to augment the training set. BorS+CavT not only achieves the state-of-the-art MSE (0.0495) on the EmotiW-EP dataset, but also obtains the state-of-the-art MSE (0.0377) on the DAiSEE dataset. The code and models have been made publicly available at https://github.com/mountainai/cavt.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
OPEN: A Benchmark Dataset and Baseline for Older Adult Patient Engagement Recognition in Virtual Rehabilitation Learning Environments
OPEN releases landmark and feature data from 35 hours of older adult virtual rehab sessions with engagement, affect, behavior, and context annotations, plus baselines reaching up to 81% accuracy.