Pith. sign in

REVIEW 2 cited by

Multimodal Stress Detection Using Facial Landmarks and Biometric Signals

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.03606 v1 pith:ZEBPRZMA submitted 2023-11-06 cs.CV cs.AI

classification cs.CVcs.AI
keywords stressmulti-modaldetectionfacialresearchsignalsapproachbiometric
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The development of various sensing technologies is improving measurements of stress and the well-being of individuals. Although progress has been made with single signal modalities like wearables and facial emotion recognition, integrating multiple modalities provides a more comprehensive understanding of stress, given that stress manifests differently across different people. Multi-modal learning aims to capitalize on the strength of each modality rather than relying on a single signal. Given the complexity of processing and integrating high-dimensional data from limited subjects, more research is needed. Numerous research efforts have been focused on fusing stress and emotion signals at an early stage, e.g., feature-level fusion using basic machine learning methods and 1D-CNN Methods. This paper proposes a multi-modal learning approach for stress detection that integrates facial landmarks and biometric signals. We test this multi-modal integration with various early-fusion and late-fusion techniques to integrate the 1D-CNN model from biometric signals and 2-D CNN using facial landmarks. We evaluate these architectures using a rigorous test of models' generalizability using the leave-one-subject-out mechanism, i.e., all samples related to a single subject are left out to train the model. Our findings show that late-fusion achieved 94.39\% accuracy, and early-fusion surpassed it with a 98.38\% accuracy rate. This research contributes valuable insights into enhancing stress detection through a multi-modal approach. The proposed research offers important knowledge in improving stress detection using a multi-modal approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UL-DD: A Multimodal Drowsiness Dataset Using Video, Biometric Signals, and Behavioral Data

    cs.CV 2025-07 conditional novelty 6.0 of 10

    UL-DD is a 1,400-minute multimodal driver drowsiness dataset with video, biometric, behavioral, and telemetry streams labeled with KSS every four minutes.

  2. Blood Glucose Level Prediction in Type 1 Diabetes Using Machine Learning

    q-bio.QM 2025-01 conditional novelty 4.0 of 10

    A benchmark of 15 machine learning models for 30-minute-ahead glucose prediction on the DiaTrend dataset finds a voting ensemble of MLP, LSTM, and GRU marginally best with an RMSE of 22.50 mg/dL.

Pith tools