Pith. sign in

REVIEW 2 cited by

Class-attention Video Transformer for Engagement Intensity Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.07216 v2 pith:WU6LWCTI submitted 2022-08-12 cs.CV cs.LG

classification cs.CVcs.LG
keywords videocavtvideosborsclassdatasetend-to-endengagement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In order to deal with variant-length long videos, prior works extract multi-modal features and fuse them to predict students' engagement intensity. In this paper, we present a new end-to-end method Class Attention in Video Transformer (CavT), which involves a single vector to process class embedding and to uniformly perform end-to-end learning on variant-length long videos and fixed-length short videos. Furthermore, to address the lack of sufficient samples, we propose a binary-order representatives sampling method (BorS) to add multiple video sequences of each video to augment the training set. BorS+CavT not only achieves the state-of-the-art MSE (0.0495) on the EmotiW-EP dataset, but also obtains the state-of-the-art MSE (0.0377) on the DAiSEE dataset. The code and models have been made publicly available at https://github.com/mountainai/cavt.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OPEN: A Benchmark Dataset and Baseline for Older Adult Patient Engagement Recognition in Virtual Rehabilitation Learning Environments

    cs.CV 2025-07 conditional novelty 6.0 of 10

    OPEN releases landmark and feature data from 35 hours of older adult virtual rehab sessions with engagement, affect, behavior, and context annotations, plus baselines reaching up to 81% accuracy.

  2. Supervised Contrastive Learning for Ordinal Engagement Measurement

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A supervised contrastive ordinal classifier with time-series augmentation improves minority-class recall on DAiSEE, but not overall accuracy, and the best non-contrastive baseline nearly matches it.

Pith tools