Pith. sign in

REVIEW 1 cited by

On the use of Self-supervised Pre-trained Acoustic and Linguistic Features for Continuous Speech Emotion Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.09212 v1 pith:SXQNWI4R submitted 2020-11-18 cs.CL

classification cs.CL
keywords continuousallosatdataemotionfeaturespre-trainedrecognitionself-supervised
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-training for feature extraction is an increasingly studied approach to get better continuous representations of audio and text content. In the present work, we use wav2vec and camemBERT as self-supervised learned models to represent our data in order to perform continuous emotion recognition from speech (SER) on AlloSat, a large French emotional database describing the satisfaction dimension, and on the state of the art corpus SEWA focusing on valence, arousal and liking dimensions. To the authors' knowledge, this paper presents the first study showing that the joint use of wav2vec and BERT-like pre-trained features is very relevant to deal with continuous SER task, usually characterized by a small amount of labeled training data. Evaluated by the well-known concordance correlation coefficient (CCC), our experiments show that we can reach a CCC value of 0.825 instead of 0.592 when using MFCC in conjunction with word2vec word embedding on the AlloSat dataset.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Dataset for Automatic Assessment of TTS Quality in Spanish

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A new Spanish-language dataset of 4,326 MOS-rated TTS audio clips enables automated naturalness prediction with a mean absolute error around 0.8 on a five-point scale.

Pith tools