Pith. sign in

REVIEW 4 major objections 5 minor 28 references

CG-MER: A Card Game-based Multimodal dataset for Emotion Recognition

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper introduces CG-MER, a multimodal French emotion-recognition dataset built from ten dyadic card-game sessions, and argues it fills a gap in spontaneous, French-language affective computing benchmarks.

desk verdict CG-MER is a genuinely new French multimodal collection with an interesting card-game protocol, but the paper as written undermines its own ground-truth claim and ships no data. read the letter →

arxiv 2501.08182 v1 pith:ATBQZU4M submitted 2025-01-14 cs.AI cs.CVcs.HC

classification cs.AIcs.CVcs.HC
keywords emotionrecognitionmultimodaldatasetCG-MERspontaneousfacialexpressionspeechskeletondataaffectivecomputing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces CG-MER, a multimodal French dataset for emotion recognition assembled from ten dyadic sessions of an emotional card game. Twenty participants answered questions of low, medium, and high personal intensity while two Kinect cameras recorded RGB and depth video, separate audio tracks captured speech, and Blazepose extracted skeleton motion. After each question, the speaker and partner each labeled the speaker's emotion among seven basic emotions; three external raters later annotated the recordings over time ranges, using the self and partner labels as their starting point. The authors' claim is that this combination—spontaneous social interaction rather than acted displays, synchronized face/voice/body channels, and French-language content—fills a gap left by existing English and acted multimodal corpora, and that CG-MER can serve as a benchmark for unimodal and multimodal emotion recognition.

What carries the argument

The central object is the card-game elicitation setup itself: two participants facing each other across a table, two Kinect v1 cameras capturing RGB and depth, audio capture per participant, and Blazepose-generated skeleton tracks. The three card categories (green, yellow, red; scores +1, +3, +5) are the intensity-control mechanism meant to produce a spread of genuine emotional responses. The labeling machinery is three-layer: immediate self-annotation, immediate partner-annotation, and delayed external-observer annotation over time ranges. The paper states that the external raters used a data-driven approach relying on the self and partner annotations, so the third layer is built on the two in-game label sets rather than being fully separate from them.

What would settle it

One could test the independence of the external annotations by asking fresh raters to label the same videos without ever seeing the self or partner annotations and measuring agreement; if those independent ratings are no closer to the recorded behavior than chance, or if the published external labels simply reproduce the self/partner majority, the dataset's claim of providing a meaningful observer perspective would be refuted. A more direct check would be to run a replication of the card game, record it, and compare participants' self-labels with blind viewers' ratings of the same moments.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is a collection protocol and dataset: CG-MER captures genuinely spontaneous emotional expressions by having two people play a card game whose red, yellow, and green cards vary the intimacy of the questions. The resulting ten hours of interaction are stored per participant as RGB video, depth video, audio, and a skeleton video, with external-observer annotations in JSON and CSV plus the in-game self and partner labels. The paper presents this as a new benchmark for emotion recognition in French, one that offers synchronized multimodal evidence of how emotions appear in real social interaction rather than in posed recordings. It also positions the card game's intensity gradient as a way to study emotional dynamics, since questions with scores +1, +3, and +5 were designed to elicit progressively stronger responses.

Load-bearing premise

The load-bearing premise is that the card game made participants genuinely feel and naturally display the emotions they report, and that the self- and partner-based labels—which the external observers were told to use as their basis—therefore constitute reliable ground truth for emotion recognition.

Editorial extensions

If this is right

  • A model trained on CG-MER can be evaluated separately on facial video, depth, audio, or skeleton streams, and then on their fusion, giving a common benchmark for comparing unimodal and multimodal emotion recognition.
  • Because the corpus is French and spontaneous, it offers a testbed for whether emotion recognition systems built on English or acted data transfer to conversational French dyads.
  • The card color and score design makes it possible to relate question intensity to the strength or clarity of the emotion expressed, something most existing corpora do not encode.
  • The time-range external annotations plus per-question self/partner labels allow both discrete event-level and continuous segment-level emotion modeling.
  • If transcripts are added, the same recordings could support text-based and audio-text fusion models, an extension the authors explicitly mention.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the external raters were told to base their judgments on the participants' self and partner annotations, the three 'perspectives' are not statistically independent; the external labels are best treated as a derived or smoothed variant of the two in-game labels, not as a third source of ground truth.
  • The intensity gradient of the cards (green/yellow/red) is a natural independent variable for testing whether emotion-recognition accuracy rises with elicitation intensity, a comparison the paper does not itself run.
  • The dataset is not publicly available at the time of writing, so its benchmark value depends on release; a useful next step would be a public subset with fixed train/test splits and baseline unimodal results.
  • A future version aimed at neurodegenerative populations, as the conclusion proposes, would need to re-examine whether the same card questions are emotionally appropriate and whether the annotation scheme captures atypical expression.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper introduces CG-MER, a French multimodal emotion recognition dataset collected from 20 participants in 10 dyadic emotional card game sessions. The dataset is described as containing RGB video, depth video, audio, and skeleton data, together with annotations from three perspectives: self, partner, and external observers. The paper details the acquisition setup, the card-game protocol, the annotation procedure, and the final dataset organization, but it reports no baseline experiments, no inter-rater reliability metrics, and states that the dataset is not publicly available at this time. The central claim is that CG-MER is a comprehensive multimodal resource for emotion recognition research in French.

Significance. If the dataset were publicly accessible and the annotation protocol were genuinely independent across perspectives, CG-MER would fill a practical gap: a French spontaneous dyadic emotion corpus with RGB, depth, audio, and skeleton modalities. The game-based elicitation with varying question intensity is an original protocol, and the GDPR-compliant consent procedure is a clear strength. However, the contribution is currently unverifiable: the dataset is not released, the external-observer annotations are explicitly based on the self and partner annotations, and the final dataset listing omits self and partner annotation files. The absence of inter-rater agreement and any baseline evaluation further weakens the paper's ability to substantiate its central claim.

major comments (4)
  1. [3.4] The external-observer annotation procedure is circular as described: external raters "adopt a data-driven approach that relies on both the self and partner annotations" and these annotations "served as the basis" for the external ratings. Because the third perspective is therefore not independent of the self and partner reports, the claim in Section 3.3 of continuous annotations from three distinct perspectives is not supported. The authors must either conduct a fully blind re-annotation in which external raters see only the audiovisual data, or provide evidence such as inter-rater agreement against independent naive ratings showing that the external labels contain information beyond the self/partner labels.
  2. [3.5] The final dataset listing includes only External Observers' Annotations among the annotation files; the self and partner annotations described in Section 3.4 are not listed as part of the released dataset. If CG-MER is to support the claimed three-perspective ground truth, all three annotation sets must be included and aligned with the per-question and per-time-range segmentation.
  3. [3.3, Table 3] The reported audio duration (5h3m45s) is roughly half the video duration (10h1m24s) for the same sessions and participants. Since each participant is recorded with a separate audio track, the audio total should be comparable to the video total unless audio files are missing or the session time is double-counted. This inconsistency affects the dataset's completeness claim for the speech modality and must be explained with per-participant duration statistics.
  4. [Overall] No inter-rater reliability, label distribution statistics, or baseline emotion recognition experiments are reported. As a dataset paper, the core contribution is the annotations and the multimodal recordings, and without any agreement measure, label statistics, or a data-access mechanism, the reader cannot assess the quality or utility of the resource. The conclusion's statement that the dataset is "not publicly available at this time" further prevents verification of the dataset's contents.
minor comments (5)
  1. [Table 1] The MELD row lists 7 participants but then gives "3f, 3m +o", and the MSP-IMPROV citation [23] points to Caridakis et al. rather than to the MSP-IMPROV corpus; these entries need correction.
  2. [Abstract] The abstract mentions the potential to incorporate NLP, but no text or transcript modality is described in Section 3; please remove this claim or clarify how NLP data would be added.
  3. [3.2] The instruction that participants select two cards from each of three categories implies six questions per participant, but Table 3 reports 120 questions for 10 sessions, which is 12 per session; please clarify the number of questions per session.
  4. [Figure 2] Figure 2 is referenced as showing emotion distribution across participants in selected sessions, but the caption lacks axis labels and normalization details, making the plot difficult to interpret.
  5. [3.4] The phrase "Using "label-studio", the open-source data labeling tool" should be grammatically integrated, and the description should specify whether the label-studio output was post-processed or filtered.

Circularity Check

1 steps flagged · score 6.0 of 10

External-observer annotations are based on self/partner labels, so the claimed third perspective collapses into the other two by construction.

  1. self definitional [Section 3.4, Data annotation, External Observers Annotation; Tables 3 and 4]
    "For the external observers’ annotation, we adopted a data-driven approach that relies on both the self and partner annotations provided during the game sessions. These annotations served as the basis for external observers to assess and annotate the participants’ emotions."

    The paper presents self, partner, and external observer annotations as three independent perspectives ('three distinct perspectives... comprehensive ground truth'). However, the external raters' labels are explicitly based on the self and partner annotations: the third perspective is constructed from the first two. Any external label is therefore a transformed copy of participant/partner judgments, not an independent observation of the audiovisual signal. Models trained on these labels are indirectly fitting self-report, and the external perspective cannot validate or triangulate the other two. This is a by-construction reduction of the claimed third annotation channel to its inputs, not an empirical finding.

full rationale

The only load-bearing circular step is in Section 3.4: external observer annotations are said to rely on both the self and partner annotations and to use them as the basis for rating. Since the dataset's contribution includes continuous emotion annotations from three distinct perspectives, this makes the third perspective a derived artifact rather than an independent ground-truth source. The raw multimodal recordings and the self/partner annotations are still independent contributions, and the paper makes no equation-level derivations or fitted predictions. Citations [24, 27, 28] are background context, not load-bearing self-citation. The circularity is localized but affects the comprehensive three-perspective ground truth claim, so the score is moderate rather than zero.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No new entities or parameters are introduced. The dataset is the claimed contribution, but it is not publicly available, and the axioms above are assumed without validation.

assumptions (4)
  • domain assumption The seven basic emotions (happiness, anger, neutral, fear, sadness, surprise, disgust) are a sufficient categorical scheme for the elicited emotions.
    The annotation protocol uses these fixed categories (Section 3.2), without justification for why continuous or other categories were not used.
  • domain assumption The emotional card game elicits spontaneous and genuine emotional expressions representative of real affective states.
    The paper assumes that the game creates a 'convivial and spontaneous atmosphere' (Section 3.2) that yields authentic emotions, but no validation, such as comparing to known elicitation methods, is provided.
  • domain assumption The recording setup (two Kinect cameras and microphones) captures facial, gesture, and speech modalities with sufficient quality for emotion recognition.
    The paper describes the setup but gives no quality metrics or validation of the captured data.
  • domain assumption External observer annotations informed by self and partner annotations provide a valid ground truth.
    Section 3.4 states that external observers used self and partner annotations as a basis, so the assumption of independent third-party labels is violated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CG-MER: A Card Game-based Multimodal dataset for Emotion Recognition." pith.science (2026). https://pith.science/paper/ATBQZU4M

@misc{pith2026250108182,
  author       = {Pith},
  title        = {Pith review of: CG-MER: A Card Game-based Multimodal dataset for Emotion Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ATBQZU4M}},
  note         = {Machine review of arXiv:2501.08182}
}
read the original abstract

The field of affective computing has seen significant advancements in exploring the relationship between emotions and emerging technologies. This paper presents a novel and valuable contribution to this field with the introduction of a comprehensive French multimodal dataset designed specifically for emotion recognition. The dataset encompasses three primary modalities: facial expressions, speech, and gestures, providing a holistic perspective on emotions. Moreover, the dataset has the potential to incorporate additional modalities, such as Natural Language Processing (NLP) to expand the scope of emotion recognition research. The dataset was curated through engaging participants in card game sessions, where they were prompted to express a range of emotions while responding to diverse questions. The study included 10 sessions with 20 participants (9 females and 11 males). The dataset serves as a valuable resource for furthering research in emotion recognition and provides an avenue for exploring the intricate connections between human emotions and digital technologies.

Figures

Figures reproduced from arXiv: 2501.08182 by the authors.

Figure 1
Figure 1. The left image depicts a pair of participants engaged in the card game while seated at a table. Two Kinect cameras positioned in the middle of the table were used to capture the participants’ facial expressions, upper body movements and voice, as illustrated on the right side in the sample screenshot of the recorded footage. To configure the cameras, we installed KinectSDK-v1.8, which included the Kinect driver and … view at source ↗
Figure 2
Figure 2. Emotion distribution across participants in [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

28 extracted references · 26 canonical work pages

  1. [1]

    Deep learning -based facial emotion recognition for human –computer inter- action applications,

    M. KalpanaN. T. N. H. D . J. Chowdary, “Deep learning -based facial emotion recognition for human –computer inter- action applications,” Neural Computing and Applications (2021)

  2. [2]

    Artificial intelligenc e for mental health care: Clinical applications, barriers, facilitators, and artificial wisdom,

    E. E. Lee, J. Torous, M. De Choudhury, C. A. Depp, S. A. Graham, H. -C. Kim, M. P. Paulus, J. H. Krystal , and D. V. Jeste, “Artificial intelligenc e for mental health care: Clinical applications, barriers, facilitators, and artificial wisdom,” Biological Psychiatry: Cognitive Neuroscience and Neuroimaging 6(9), 856 –864 (2021)

  3. [3]

    Hello marketing, what can artificial intelligence help you with ?,

    N. Wirth, “Hello marketing, what can artificial intelligence help you with ?,” International Journal of Market Re - search 60(5), 435–438 (2018)

  4. [4]

    Artificial intelligence in education: A review,

    L. Chen, P. Chen, and Z. Lin, “Artificial intelligence in education: A review,” IEEE Access 8, 75264 –75278 (2020)

  5. [5]

    Affectnet: A database f or facial expression, valence, and arousal computing in the wild,

    A. Mollahosseini, B. Hasani, and M. H. Mahoor, “Affectnet: A database f or facial expression, valence, and arousal computing in the wild,” IEEE Transactions on Affective Computing 10(1), 18 –31 (2017)

  6. [6]

    Challenges in representation learning: A report on three machine learning contests,

    I. J. Goodfellow, D. Erhan, P. L. Carrier, A. Courville, M. Mirza, B. Hamner, W. Cuki erski, Y. Tang, D. Thaler, D. - H. Lee, and others, “Challenges in representation learning: A report on three machine learning contests,” in Neural Information Processing: 20th International Conference, ICONIP 2013, Daegu, Korea, November 3 -7, 2013. Proceedings, Part II...

  7. [7]

    The ryerson audio -visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,

    S. R. Livingstone and F. A. Russo, “The ryerson audio -visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” PLOS ONE 13, 1 –35 (05 2018)

  8. [8]

    A database of german emotional speech.,

    F. Burkhardt, A. Paeschke, M . Rolfes, W. F. Sendlmeier, B. Weiss, and others, “A database of german emotional speech.,” in Interspeech, 5, 1517–1520 (2005)

Show all 28 references
  1. [9]

    Multi -modal gesture recognit ion challenge 2013: Dataset and results,

    S. Escalera, J. Gonza`lez, X. Baro´, M. Reyes, O. Lopes, I. Guyon, V. Athitsos, and H . Escalante, “Multi -modal gesture recognit ion challenge 2013: Dataset and results,” in Proceedings of the 15th ACM on International conference on multimodal interaction, 445–452 (2013)

  2. [10]

    A survey on databases for multimodal emotion recognition and an introduction to the viri (visible and infrared image) database,

    M. F. H. Siddiqui, P. Dhakal, X. Yang, and A. Y. Javaid, “A survey on databases for multimodal emotion recognition and an introduction to the viri (visible and infrared image) database,” Multimodal Technologies and Interaction 6(6), 47 (2022)

  3. [11]

    Facial expression recognition based on deep learning: a survey,

    T. Zhang, “Facial expression recognition based on deep learning: a survey,” in Adv ances in Intelligent Systems and Interactive Applications: Proceedings of the 2nd International Conference on Intelligent and Interactive Systems and Applications (IISA2017), 345–352, Springer (2018)

  4. [12]

    D eep learning techniques for speech emotion recognition, from databases to models,

    B. J. Abbaschian, D. Sierra -Sosa, and A. Elmaghraby, “D eep learning techniques for speech emotion recognition, from databases to models,” Sensors 21(4), 1249 (2021)

  5. [13]

    Speech emotion recognition using deep learning techniques: A review,

    R. A. Khalil, E. Jones, M. I. Babar, T. Jan, M. H. Zafar, and T. Alhussain, “Speech emotion recognition using deep learning techniques: A review,” IEEE Access 7, 117327–117345 (2019)

  6. [14]

    Survey on emotional body gesture recognition,

    F. Noroozi, C. A. Corneanu, D. Kamin´ska, T. Sapin´ski, S. Escalera, and G. Anbarjafari, “Survey on emotional body gesture recognition,” IEEE transactions on affective computing 12(2), 505 –523 (2018)

  7. [15]

    Deep multimodal emotion recognition on human speech: A review,

    P. Koromilas and T. Giannakopoulos, “Deep multimodal emotion recognition on human speech: A review,” Applied Sciences 11(17), 7962 (2021)

  8. [16]

    Multimodal approaches for emotion recognition: a survey,

    N. Sebe, I. Cohen, T. Gevers, and T. S. Huang, “Multimodal approaches for emotion recognition: a survey,” in Internet Imaging VI, 5670, 56–67, SPIE (2005)

  9. [17]

    K - emocon, a multimodal sensor dataset for continuous emotion recognition in naturalistic conversations,

    C. Y. Park, N. Cha, S. Kang, A. Kim, A. H. Khandoker, L. Hadjileontiadis, A. Oh, Y. Jeong, and U. Lee, “K - emocon, a multimodal sensor dataset for continuous emotion recognition in naturalistic conversations,” Scientific Data 7(1), 293 (2020)

  10. [18]

    Amigos: A dataset for affect, personality and mood research on individuals and groups,

    J. A. Miranda -Correa, M. K. Abad i, N. Sebe, and I. Patras, “Amigos: A dataset for affect, personality and mood research on individuals and groups,” IEEE Transactions on Affective Computing 12(2), 479 –493 (2018)

  11. [19]

    Cross-subject multimodal emotion recognition based on hybrid fusion,

    Y. Cimtay, E. Ekmekcioglu, and S. Caglar-Ozhan, “Cross-subject multimodal emotion recognition based on hybrid fusion,” IEEE Access 8, 168865–168878 (2020)

  12. [20]

    Meld: A multimodal multi -party dataset for emotion recognition in conversations,

    S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea, “Meld: A multimodal multi -party dataset for emotion recognition in conversations,” arXiv preprint arXiv:1810.02508 (2018)

  13. [21]

    Iemocap: Interactive emotional dyadic motion capture database,

    C. Busso, M. Bulut, C. -C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation 42, 335 –359 (2008)

  14. [22]

    Multimodal emotion recognition using deep learning archi - tectures,

    H. Ranganathan, S. Chakraborty, and S. Panchanathan, “Multimodal emotion recognition using deep learning archi - tectures,” in 2016 IEEE winter conference on applications of computer vision (WACV) , 1 –9, IEEE (2016)

  15. [23]

    A cross -cultural, multimodal, affective corpus for gesture expressivity analysis,

    G. Caridakis, J. Wagner, A. Raouzaiou, F. Lingenfelser, K. Karpouzis, and E. Andre, “A cross -cultural, multimodal, affective corpus for gesture expressivity analysis,” Journal on Multimodal User Interfaces 7, 121 –134 (2013)

  16. [24]

    Percep tual borderline for balancing multi -class spontaneous emotional data,

    L. B. Letaifa and M. I. Torres, “Percep tual borderline for balancing multi -class spontaneous emotional data,” IEEE Access 9, 55939–55954 (2021)

  17. [25]

    Blazepose: On -device real- time body pose tracking,

    V. Bazarevsky, I. Grishchenko, K. Raveendran, T. Zhu, F. Zhang, and M. Grundmann, “Blazepose: On -device real- time body pose tracking,” arXiv preprint arXiv:2006.10204 (2020)

  18. [26]

    Graph convolutional networks for assessment of physical rehabil- itation exercises,

    S. Deb, M. F. Islam, S. Rahman, and S. Rahman, “Graph convolutional networks for assessment of physical rehabil- itation exercises,” IEEE Transactions on Neural Systems and Rehabilitation Engineering 30, 410 –419 (2022)

  19. [27]

    Recording affect in the field: Towards methods and metrics for improving ground truth labels,

    J. Healey, “Recording affect in the field: Towards methods and metrics for improving ground truth labels,” in Affec - tive Computing and Intelligent Interaction: 4th International Conference, ACII 2011, Memphis, TN, USA, October 9–12, 2011, Proceedings, Part I 4 , 107–116, Spr...

  20. [28]

    Automatic recognition of self -reported a nd perceived emotion: Does joint modeling help?,

    B. Zhang, G. Essl, and E. Mower Provost, “Automatic recognition of self -reported a nd perceived emotion: Does joint modeling help?,” in Proceedings of the 18th ACM International Conference on Multimodal Interaction, 217 – 224 (2016)

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.