REVIEW 4 major objections 5 minor 28 references
CG-MER: A Card Game-based Multimodal dataset for Emotion Recognition
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper introduces CG-MER, a multimodal French emotion-recognition dataset built from ten dyadic card-game sessions, and argues it fills a gap in spontaneous, French-language affective computing benchmarks.
desk verdict CG-MER is a genuinely new French multimodal collection with an interesting card-game protocol, but the paper as written undermines its own ground-truth claim and ships no data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the card-game elicitation setup itself: two participants facing each other across a table, two Kinect v1 cameras capturing RGB and depth, audio capture per participant, and Blazepose-generated skeleton tracks. The three card categories (green, yellow, red; scores +1, +3, +5) are the intensity-control mechanism meant to produce a spread of genuine emotional responses. The labeling machinery is three-layer: immediate self-annotation, immediate partner-annotation, and delayed external-observer annotation over time ranges. The paper states that the external raters used a data-driven approach relying on the self and partner annotations, so the third layer is built on the two in-game label sets rather than being fully separate from them.
What would settle it
One could test the independence of the external annotations by asking fresh raters to label the same videos without ever seeing the self or partner annotations and measuring agreement; if those independent ratings are no closer to the recorded behavior than chance, or if the published external labels simply reproduce the self/partner majority, the dataset's claim of providing a meaningful observer perspective would be refuted. A more direct check would be to run a replication of the card game, record it, and compare participants' self-labels with blind viewers' ratings of the same moments.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is a collection protocol and dataset: CG-MER captures genuinely spontaneous emotional expressions by having two people play a card game whose red, yellow, and green cards vary the intimacy of the questions. The resulting ten hours of interaction are stored per participant as RGB video, depth video, audio, and a skeleton video, with external-observer annotations in JSON and CSV plus the in-game self and partner labels. The paper presents this as a new benchmark for emotion recognition in French, one that offers synchronized multimodal evidence of how emotions appear in real social interaction rather than in posed recordings. It also positions the card game's intensity gradient as a way to study emotional dynamics, since questions with scores +1, +3, and +5 were designed to elicit progressively stronger responses.
Load-bearing premise
The load-bearing premise is that the card game made participants genuinely feel and naturally display the emotions they report, and that the self- and partner-based labels—which the external observers were told to use as their basis—therefore constitute reliable ground truth for emotion recognition.
Editorial extensions
If this is right
- A model trained on CG-MER can be evaluated separately on facial video, depth, audio, or skeleton streams, and then on their fusion, giving a common benchmark for comparing unimodal and multimodal emotion recognition.
- Because the corpus is French and spontaneous, it offers a testbed for whether emotion recognition systems built on English or acted data transfer to conversational French dyads.
- The card color and score design makes it possible to relate question intensity to the strength or clarity of the emotion expressed, something most existing corpora do not encode.
- The time-range external annotations plus per-question self/partner labels allow both discrete event-level and continuous segment-level emotion modeling.
- If transcripts are added, the same recordings could support text-based and audio-text fusion models, an extension the authors explicitly mention.
Reading between the lines
- Because the external raters were told to base their judgments on the participants' self and partner annotations, the three 'perspectives' are not statistically independent; the external labels are best treated as a derived or smoothed variant of the two in-game labels, not as a third source of ground truth.
- The intensity gradient of the cards (green/yellow/red) is a natural independent variable for testing whether emotion-recognition accuracy rises with elicitation intensity, a comparison the paper does not itself run.
- The dataset is not publicly available at the time of writing, so its benchmark value depends on release; a useful next step would be a public subset with fixed train/test splits and baseline unimodal results.
- A future version aimed at neurodegenerative populations, as the conclusion proposes, would need to re-examine whether the same card questions are emotionally appropriate and whether the annotation scheme captures atypical expression.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces CG-MER, a French multimodal emotion recognition dataset collected from 20 participants in 10 dyadic emotional card game sessions. The dataset is described as containing RGB video, depth video, audio, and skeleton data, together with annotations from three perspectives: self, partner, and external observers. The paper details the acquisition setup, the card-game protocol, the annotation procedure, and the final dataset organization, but it reports no baseline experiments, no inter-rater reliability metrics, and states that the dataset is not publicly available at this time. The central claim is that CG-MER is a comprehensive multimodal resource for emotion recognition research in French.
Significance. If the dataset were publicly accessible and the annotation protocol were genuinely independent across perspectives, CG-MER would fill a practical gap: a French spontaneous dyadic emotion corpus with RGB, depth, audio, and skeleton modalities. The game-based elicitation with varying question intensity is an original protocol, and the GDPR-compliant consent procedure is a clear strength. However, the contribution is currently unverifiable: the dataset is not released, the external-observer annotations are explicitly based on the self and partner annotations, and the final dataset listing omits self and partner annotation files. The absence of inter-rater agreement and any baseline evaluation further weakens the paper's ability to substantiate its central claim.
major comments (4)
- [3.4] The external-observer annotation procedure is circular as described: external raters "adopt a data-driven approach that relies on both the self and partner annotations" and these annotations "served as the basis" for the external ratings. Because the third perspective is therefore not independent of the self and partner reports, the claim in Section 3.3 of continuous annotations from three distinct perspectives is not supported. The authors must either conduct a fully blind re-annotation in which external raters see only the audiovisual data, or provide evidence such as inter-rater agreement against independent naive ratings showing that the external labels contain information beyond the self/partner labels.
- [3.5] The final dataset listing includes only External Observers' Annotations among the annotation files; the self and partner annotations described in Section 3.4 are not listed as part of the released dataset. If CG-MER is to support the claimed three-perspective ground truth, all three annotation sets must be included and aligned with the per-question and per-time-range segmentation.
- [3.3, Table 3] The reported audio duration (5h3m45s) is roughly half the video duration (10h1m24s) for the same sessions and participants. Since each participant is recorded with a separate audio track, the audio total should be comparable to the video total unless audio files are missing or the session time is double-counted. This inconsistency affects the dataset's completeness claim for the speech modality and must be explained with per-participant duration statistics.
- [Overall] No inter-rater reliability, label distribution statistics, or baseline emotion recognition experiments are reported. As a dataset paper, the core contribution is the annotations and the multimodal recordings, and without any agreement measure, label statistics, or a data-access mechanism, the reader cannot assess the quality or utility of the resource. The conclusion's statement that the dataset is "not publicly available at this time" further prevents verification of the dataset's contents.
minor comments (5)
- [Table 1] The MELD row lists 7 participants but then gives "3f, 3m +o", and the MSP-IMPROV citation [23] points to Caridakis et al. rather than to the MSP-IMPROV corpus; these entries need correction.
- [Abstract] The abstract mentions the potential to incorporate NLP, but no text or transcript modality is described in Section 3; please remove this claim or clarify how NLP data would be added.
- [3.2] The instruction that participants select two cards from each of three categories implies six questions per participant, but Table 3 reports 120 questions for 10 sessions, which is 12 per session; please clarify the number of questions per session.
- [Figure 2] Figure 2 is referenced as showing emotion distribution across participants in selected sessions, but the caption lacks axis labels and normalization details, making the plot difficult to interpret.
- [3.4] The phrase "Using "label-studio", the open-source data labeling tool" should be grammatically integrated, and the description should specify whether the label-studio output was post-processed or filtered.
Circularity Check
External-observer annotations are based on self/partner labels, so the claimed third perspective collapses into the other two by construction.
-
self definitional
[Section 3.4, Data annotation, External Observers Annotation; Tables 3 and 4]
"For the external observers’ annotation, we adopted a data-driven approach that relies on both the self and partner annotations provided during the game sessions. These annotations served as the basis for external observers to assess and annotate the participants’ emotions."
The paper presents self, partner, and external observer annotations as three independent perspectives ('three distinct perspectives... comprehensive ground truth'). However, the external raters' labels are explicitly based on the self and partner annotations: the third perspective is constructed from the first two. Any external label is therefore a transformed copy of participant/partner judgments, not an independent observation of the audiovisual signal. Models trained on these labels are indirectly fitting self-report, and the external perspective cannot validate or triangulate the other two. This is a by-construction reduction of the claimed third annotation channel to its inputs, not an empirical finding.
full rationale
The only load-bearing circular step is in Section 3.4: external observer annotations are said to rely on both the self and partner annotations and to use them as the basis for rating. Since the dataset's contribution includes continuous emotion annotations from three distinct perspectives, this makes the third perspective a derived artifact rather than an independent ground-truth source. The raw multimodal recordings and the self/partner annotations are still independent contributions, and the paper makes no equation-level derivations or fitted predictions. Citations [24, 27, 28] are background context, not load-bearing self-citation. The circularity is localized but affects the comprehensive three-perspective ground truth claim, so the score is moderate rather than zero.
Assumptions & free parameters
assumptions (4)
- domain assumption The seven basic emotions (happiness, anger, neutral, fear, sadness, surprise, disgust) are a sufficient categorical scheme for the elicited emotions.
- domain assumption The emotional card game elicits spontaneous and genuine emotional expressions representative of real affective states.
- domain assumption The recording setup (two Kinect cameras and microphones) captures facial, gesture, and speech modalities with sufficient quality for emotion recognition.
- domain assumption External observer annotations informed by self and partner annotations provide a valid ground truth.
Cite this review
Pith. "Pith review of CG-MER: A Card Game-based Multimodal dataset for Emotion Recognition." pith.science (2026). https://pith.science/paper/ATBQZU4M
@misc{pith2026250108182,
author = {Pith},
title = {Pith review of: CG-MER: A Card Game-based Multimodal dataset for Emotion Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/ATBQZU4M}},
note = {Machine review of arXiv:2501.08182}
}
read the original abstract
The field of affective computing has seen significant advancements in exploring the relationship between emotions and emerging technologies. This paper presents a novel and valuable contribution to this field with the introduction of a comprehensive French multimodal dataset designed specifically for emotion recognition. The dataset encompasses three primary modalities: facial expressions, speech, and gestures, providing a holistic perspective on emotions. Moreover, the dataset has the potential to incorporate additional modalities, such as Natural Language Processing (NLP) to expand the scope of emotion recognition research. The dataset was curated through engaging participants in card game sessions, where they were prompted to express a range of emotions while responding to diverse questions. The study included 10 sessions with 20 participants (9 females and 11 males). The dataset serves as a valuable resource for furthering research in emotion recognition and provides an avenue for exploring the intricate connections between human emotions and digital technologies.
Figures
Reference graph
Works this paper leans on
-
[1]
Deep learning -based facial emotion recognition for human –computer inter- action applications,
M. KalpanaN. T. N. H. D . J. Chowdary, “Deep learning -based facial emotion recognition for human –computer inter- action applications,” Neural Computing and Applications (2021)
work page 2021
-
[2]
E. E. Lee, J. Torous, M. De Choudhury, C. A. Depp, S. A. Graham, H. -C. Kim, M. P. Paulus, J. H. Krystal , and D. V. Jeste, “Artificial intelligenc e for mental health care: Clinical applications, barriers, facilitators, and artificial wisdom,” Biological Psychiatry: Cognitive Neuroscience and Neuroimaging 6(9), 856 –864 (2021)
work page 2021
-
[3]
Hello marketing, what can artificial intelligence help you with ?,
N. Wirth, “Hello marketing, what can artificial intelligence help you with ?,” International Journal of Market Re - search 60(5), 435–438 (2018)
work page 2018
-
[4]
Artificial intelligence in education: A review,
L. Chen, P. Chen, and Z. Lin, “Artificial intelligence in education: A review,” IEEE Access 8, 75264 –75278 (2020)
work page 2020
-
[5]
Affectnet: A database f or facial expression, valence, and arousal computing in the wild,
A. Mollahosseini, B. Hasani, and M. H. Mahoor, “Affectnet: A database f or facial expression, valence, and arousal computing in the wild,” IEEE Transactions on Affective Computing 10(1), 18 –31 (2017)
work page 2017
-
[6]
Challenges in representation learning: A report on three machine learning contests,
I. J. Goodfellow, D. Erhan, P. L. Carrier, A. Courville, M. Mirza, B. Hamner, W. Cuki erski, Y. Tang, D. Thaler, D. - H. Lee, and others, “Challenges in representation learning: A report on three machine learning contests,” in Neural Information Processing: 20th International Conference, ICONIP 2013, Daegu, Korea, November 3 -7, 2013. Proceedings, Part II...
work page 2013
-
[7]
S. R. Livingstone and F. A. Russo, “The ryerson audio -visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” PLOS ONE 13, 1 –35 (05 2018)
work page 2018
-
[8]
A database of german emotional speech.,
F. Burkhardt, A. Paeschke, M . Rolfes, W. F. Sendlmeier, B. Weiss, and others, “A database of german emotional speech.,” in Interspeech, 5, 1517–1520 (2005)
work page 2005
Show all 28 references
-
[9]
Multi -modal gesture recognit ion challenge 2013: Dataset and results,
S. Escalera, J. Gonza`lez, X. Baro´, M. Reyes, O. Lopes, I. Guyon, V. Athitsos, and H . Escalante, “Multi -modal gesture recognit ion challenge 2013: Dataset and results,” in Proceedings of the 15th ACM on International conference on multimodal interaction, 445–452 (2013)
2013
-
[10]
A survey on databases for multimodal emotion recognition and an introduction to the viri (visible and infrared image) database,
M. F. H. Siddiqui, P. Dhakal, X. Yang, and A. Y. Javaid, “A survey on databases for multimodal emotion recognition and an introduction to the viri (visible and infrared image) database,” Multimodal Technologies and Interaction 6(6), 47 (2022)
2022
-
[11]
Facial expression recognition based on deep learning: a survey,
T. Zhang, “Facial expression recognition based on deep learning: a survey,” in Adv ances in Intelligent Systems and Interactive Applications: Proceedings of the 2nd International Conference on Intelligent and Interactive Systems and Applications (IISA2017), 345–352, Springer (2018)
2018
-
[12]
D eep learning techniques for speech emotion recognition, from databases to models,
B. J. Abbaschian, D. Sierra -Sosa, and A. Elmaghraby, “D eep learning techniques for speech emotion recognition, from databases to models,” Sensors 21(4), 1249 (2021)
2021
-
[13]
Speech emotion recognition using deep learning techniques: A review,
R. A. Khalil, E. Jones, M. I. Babar, T. Jan, M. H. Zafar, and T. Alhussain, “Speech emotion recognition using deep learning techniques: A review,” IEEE Access 7, 117327–117345 (2019)
2019
-
[14]
Survey on emotional body gesture recognition,
F. Noroozi, C. A. Corneanu, D. Kamin´ska, T. Sapin´ski, S. Escalera, and G. Anbarjafari, “Survey on emotional body gesture recognition,” IEEE transactions on affective computing 12(2), 505 –523 (2018)
2018
-
[15]
Deep multimodal emotion recognition on human speech: A review,
P. Koromilas and T. Giannakopoulos, “Deep multimodal emotion recognition on human speech: A review,” Applied Sciences 11(17), 7962 (2021)
2021
-
[16]
Multimodal approaches for emotion recognition: a survey,
N. Sebe, I. Cohen, T. Gevers, and T. S. Huang, “Multimodal approaches for emotion recognition: a survey,” in Internet Imaging VI, 5670, 56–67, SPIE (2005)
2005
-
[17]
K - emocon, a multimodal sensor dataset for continuous emotion recognition in naturalistic conversations,
C. Y. Park, N. Cha, S. Kang, A. Kim, A. H. Khandoker, L. Hadjileontiadis, A. Oh, Y. Jeong, and U. Lee, “K - emocon, a multimodal sensor dataset for continuous emotion recognition in naturalistic conversations,” Scientific Data 7(1), 293 (2020)
2020
-
[18]
Amigos: A dataset for affect, personality and mood research on individuals and groups,
J. A. Miranda -Correa, M. K. Abad i, N. Sebe, and I. Patras, “Amigos: A dataset for affect, personality and mood research on individuals and groups,” IEEE Transactions on Affective Computing 12(2), 479 –493 (2018)
2018
-
[19]
Cross-subject multimodal emotion recognition based on hybrid fusion,
Y. Cimtay, E. Ekmekcioglu, and S. Caglar-Ozhan, “Cross-subject multimodal emotion recognition based on hybrid fusion,” IEEE Access 8, 168865–168878 (2020)
2020
-
[20]
Meld: A multimodal multi -party dataset for emotion recognition in conversations,
S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea, “Meld: A multimodal multi -party dataset for emotion recognition in conversations,” arXiv preprint arXiv:1810.02508 (2018)
2018 arXiv
-
[21]
Iemocap: Interactive emotional dyadic motion capture database,
C. Busso, M. Bulut, C. -C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation 42, 335 –359 (2008)
2008
-
[22]
Multimodal emotion recognition using deep learning archi - tectures,
H. Ranganathan, S. Chakraborty, and S. Panchanathan, “Multimodal emotion recognition using deep learning archi - tectures,” in 2016 IEEE winter conference on applications of computer vision (WACV) , 1 –9, IEEE (2016)
2016
-
[23]
A cross -cultural, multimodal, affective corpus for gesture expressivity analysis,
G. Caridakis, J. Wagner, A. Raouzaiou, F. Lingenfelser, K. Karpouzis, and E. Andre, “A cross -cultural, multimodal, affective corpus for gesture expressivity analysis,” Journal on Multimodal User Interfaces 7, 121 –134 (2013)
2013
-
[24]
Percep tual borderline for balancing multi -class spontaneous emotional data,
L. B. Letaifa and M. I. Torres, “Percep tual borderline for balancing multi -class spontaneous emotional data,” IEEE Access 9, 55939–55954 (2021)
2021
-
[25]
Blazepose: On -device real- time body pose tracking,
V. Bazarevsky, I. Grishchenko, K. Raveendran, T. Zhu, F. Zhang, and M. Grundmann, “Blazepose: On -device real- time body pose tracking,” arXiv preprint arXiv:2006.10204 (2020)
2020 arXiv
-
[26]
Graph convolutional networks for assessment of physical rehabil- itation exercises,
S. Deb, M. F. Islam, S. Rahman, and S. Rahman, “Graph convolutional networks for assessment of physical rehabil- itation exercises,” IEEE Transactions on Neural Systems and Rehabilitation Engineering 30, 410 –419 (2022)
2022
-
[27]
Recording affect in the field: Towards methods and metrics for improving ground truth labels,
J. Healey, “Recording affect in the field: Towards methods and metrics for improving ground truth labels,” in Affec - tive Computing and Intelligent Interaction: 4th International Conference, ACII 2011, Memphis, TN, USA, October 9–12, 2011, Proceedings, Part I 4 , 107–116, Spr...
2011
-
[28]
Automatic recognition of self -reported a nd perceived emotion: Does joint modeling help?,
B. Zhang, G. Essl, and E. Mower Provost, “Automatic recognition of self -reported a nd perceived emotion: Does joint modeling help?,” in Proceedings of the 18th ACM International Conference on Multimodal Interaction, 217 – 224 (2016)
2016
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.