{"id":"b067bdda-03ba-42c0-b35a-bb11880b85ec","arxiv_id":"2501.08182","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"CG-MER is a new French multimodal emotion recognition dataset built from emotional card game sessions, offering RGB, depth, audio, and skeleton data with three annotation perspectives.","lead":"This paper introduces CG-MER, a French multimodal emotion recognition dataset collected during dyadic card game sessions with 20 participants. It includes facial, speech, and gesture data with self, partner, and external annotations, but the dataset is not publicly available.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"External-observer labels are derived from self/partner annotations, so the claimed third perspective is not independent; a blind re-annotation test is needed before CG-MER can be treated as a three-perspective ground truth.","rationale":"The reader's weakest assumption is the same one I would flag. I looked for a more fundamental issue, such as dataset unavailability or the audio/video duration mismatch, but those are secondary: a dataset can still be described before release, and duration mismatches could be explained by separate audio tracks per session. The external-annotation circularity directly attacks the paper's stated rationale for three perspectives and applies to the only time-continuous annotations in Table 3. A re-annotation experiment is feasible and would settle whether the concern lands. Because this supports the reader's REJECT verdict, I recommend no change. I deliberately avoid alleging intentional deception; the phrase 'data-driven approach' is ambiguous, and the test would distinguish a benign reference-point use of self/partner labels from a genuinely non-independent external ground truth.","tokens_in":8064,"tokens_out":4125,"duration_ms":41916,"concrete_test":"Take a random subset (e.g., three sessions) and have three fresh raters annotate the same video segments with no access to the self/partner annotations or to the original external labels; then compute agreement (e.g., Cohen's kappa or weighted kappa) between the blind raters and the published external labels, and also compute how well self/partner labels predict the original external labels. If the original external labels agree almost perfectly with self/partner labels but blind raters diverge, the external perspective is redundant or contaminated; if blind raters reproduce the labels, the circularity is not practically harmful.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is CG-MER as a multimodal emotion-recognition dataset with annotations from three perspectives (self, partner, external observer), and it cites prior work [27, 28] to argue that multiple perspectives improve ground truth. The load-bearing assumption is that the external-observer labels are an independent source. Section 3.4 states that external raters 'adopted a data-driven approach that relies on both the self and partner annotations' and that these 'served as the basis' for the external annotations. If the external raters were shown or anchored by self/partner labels, the third perspective is not independent: it is a transformed copy of the participants' and partners' judgments. Any model trained to predict these external labels is indirectly fitting self-report, not learning an observer's read of the audiovisual signal. Table 3 lists external annotations as the only per-time-range labels; if they are contaminated, the 'comprehensive ground truth' claim is unsupported. No inter-rater reliability or baseline experiments are reported, so the paper provides no way to tell whether the external labels add signal beyond self/partner. This is not an accusation of fraud; the text may mean raters used self/partner labels only to define segments or as a reference. But as written, the central claim depends on a third perspective that may be circular.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces CG-MER, a French multimodal emotion recognition dataset collected from 20 participants in 10 dyadic emotional card game sessions. The dataset is described as containing RGB video, depth video, audio, and skeleton data, together with annotations from three perspectives: self, partner, and external observers. The paper details the acquisition setup, the card-game protocol, the annotation procedure, and the final dataset organization, but it reports no baseline experiments, no inter-rater reliability metrics, and states that the dataset is not publicly available at this time. The central claim is that CG-MER is a comprehensive multimodal resource for emotion recognition research in French.","tokens_in":8309,"tokens_out":5325,"duration_ms":52551,"significance":"If the dataset were publicly accessible and the annotation protocol were genuinely independent across perspectives, CG-MER would fill a practical gap: a French spontaneous dyadic emotion corpus with RGB, depth, audio, and skeleton modalities. The game-based elicitation with varying question intensity is an original protocol, and the GDPR-compliant consent procedure is a clear strength. However, the contribution is currently unverifiable: the dataset is not released, the external-observer annotations are explicitly based on the self and partner annotations, and the final dataset listing omits self and partner annotation files. The absence of inter-rater agreement and any baseline evaluation further weakens the paper's ability to substantiate its central claim.","major_comments":[{"comment":"The external-observer annotation procedure is circular as described: external raters \"adopt a data-driven approach that relies on both the self and partner annotations\" and these annotations \"served as the basis\" for the external ratings. Because the third perspective is therefore not independent of the self and partner reports, the claim in Section 3.3 of continuous annotations from three distinct perspectives is not supported. The authors must either conduct a fully blind re-annotation in which external raters see only the audiovisual data, or provide evidence such as inter-rater agreement against independent naive ratings showing that the external labels contain information beyond the self/partner labels.","section":"3.4"},{"comment":"The final dataset listing includes only External Observers' Annotations among the annotation files; the self and partner annotations described in Section 3.4 are not listed as part of the released dataset. If CG-MER is to support the claimed three-perspective ground truth, all three annotation sets must be included and aligned with the per-question and per-time-range segmentation.","section":"3.5"},{"comment":"The reported audio duration (5h3m45s) is roughly half the video duration (10h1m24s) for the same sessions and participants. Since each participant is recorded with a separate audio track, the audio total should be comparable to the video total unless audio files are missing or the session time is double-counted. This inconsistency affects the dataset's completeness claim for the speech modality and must be explained with per-participant duration statistics.","section":"3.3, Table 3"},{"comment":"No inter-rater reliability, label distribution statistics, or baseline emotion recognition experiments are reported. As a dataset paper, the core contribution is the annotations and the multimodal recordings, and without any agreement measure, label statistics, or a data-access mechanism, the reader cannot assess the quality or utility of the resource. The conclusion's statement that the dataset is \"not publicly available at this time\" further prevents verification of the dataset's contents.","section":"Overall"}],"minor_comments":[{"comment":"The MELD row lists 7 participants but then gives \"3f, 3m +o\", and the MSP-IMPROV citation [23] points to Caridakis et al. rather than to the MSP-IMPROV corpus; these entries need correction.","section":"Table 1"},{"comment":"The abstract mentions the potential to incorporate NLP, but no text or transcript modality is described in Section 3; please remove this claim or clarify how NLP data would be added.","section":"Abstract"},{"comment":"The instruction that participants select two cards from each of three categories implies six questions per participant, but Table 3 reports 120 questions for 10 sessions, which is 12 per session; please clarify the number of questions per session.","section":"3.2"},{"comment":"Figure 2 is referenced as showing emotion distribution across participants in selected sessions, but the caption lacks axis labels and normalization details, making the plot difficult to interpret.","section":"Figure 2"},{"comment":"The phrase \"Using \"label-studio\", the open-source data labeling tool\" should be grammatically integrated, and the description should specify whether the label-studio output was post-processed or filtered.","section":"3.4"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is more in the style of a data descriptor than a standard research article, yet it lacks the data availability and quality metrics expected of that genre. If the journal's scope requires citable and usable resources, the authors should be asked to clarify data access before any acceptance is considered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: the collection idea is real, the execution description is not yet credible as a dataset paper. The card-game protocol for eliciting spontaneous dyadic emotion in French, with RGB, depth, audio, and skeleton streams, fills a niche that IEMOCAP, K-EmoCon, and MELD do not. Credit where it is due: 20 participants, ten sessions, roughly ten hours of video, and a three-perspective annotation scheme is a serious data-collection effort, and the paper is honest about several limitations.\n\nThat said, the central problem is the external-observer annotation. Section 3.4 says the raters \"adopted a data-driven approach that relies on both the self and partner annotations\" and that these \"served as the basis\" for their judgments. As written, that makes the third perspective a transformed copy of the participants' self-reports, not an independent read of the audiovisual signal. The stress-test note is right: this is ambiguous and may be a harmless procedural detail, but it lands squarely on the paper's claim of a \"comprehensive ground truth\" from three perspectives. No inter-rater reliability numbers are given, and no baseline experiments are reported, so there is no way to tell whether those external labels add any signal beyond self/partner.\n\nThe other soft spots are minor by comparison. The audio duration (5h3m45s) being roughly half the video duration (10h1m24s) for the same sessions needs explanation; it may be that only one participant's audio per session was counted, but the text does not say. The dataset is not publicly available, the paper says so in the conclusion, so the central artifact cannot be inspected or used. Table 4 actually checks out; the first three sessions' totals sum correctly, so that is not a flaw.\n\nWho is this for? A reader working on French spontaneous emotion recognition might find the protocol worth studying, but the current paper gives them no data and no validated labels. As a dataset contribution, it is not usable as-is. I would not send this to review in this form; I would desk-reject and invite a resubmission with the data released, external annotations redone blind to self/partner labels, and at least a basic recognition baseline or IRR metric.\n\nIn short: worthwhile raw material, not yet a paper that supports its own claims.","headline":"CG-MER is a genuinely new French multimodal collection with an interesting card-game protocol, but the paper as written undermines its own ground-truth claim and ships no data.","tokens_in":8827,"tokens_out":2491,"would_cite":false,"duration_ms":25618,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces CG-MER, a multimodal French emotion-recognition dataset built from ten dyadic card-game sessions, and argues it fills a gap in spontaneous, French-language affective computing benchmarks.","keywords":["emotion recognition","multimodal dataset","CG-MER","spontaneous emotion","facial expression","speech emotion","skeleton data","affective computing"],"falsifier":"One could test the independence of the external annotations by asking fresh raters to label the same videos without ever seeing the self or partner annotations and measuring agreement; if those independent ratings are no closer to the recorded behavior than chance, or if the published external labels simply reproduce the self/partner majority, the dataset's claim of providing a meaningful observer perspective would be refuted. A more direct check would be to run a replication of the card game, record it, and compare participants' self-labels with blind viewers' ratings of the same moments.","tokens_in":7879,"feed_emoji":"🎴","tokens_out":7398,"duration_ms":71471,"temperature":0.7,"pith_summary":"The paper introduces CG-MER, a multimodal French dataset for emotion recognition assembled from ten dyadic sessions of an emotional card game. Twenty participants answered questions of low, medium, and high personal intensity while two Kinect cameras recorded RGB and depth video, separate audio tracks captured speech, and Blazepose extracted skeleton motion. After each question, the speaker and partner each labeled the speaker's emotion among seven basic emotions; three external raters later annotated the recordings over time ranges, using the self and partner labels as their starting point. The authors' claim is that this combination—spontaneous social interaction rather than acted displays, synchronized face/voice/body channels, and French-language content—fills a gap left by existing English and acted multimodal corpora, and that CG-MER can serve as a benchmark for unimodal and multimodal emotion recognition.","feed_headline":"A card game yields a French multimodal emotion dataset","feed_subtitle":"Ten hours of dyadic play pair facial, vocal, and skeletal signals with seven emotion labels.","key_machinery":"The central object is the card-game elicitation setup itself: two participants facing each other across a table, two Kinect v1 cameras capturing RGB and depth, audio capture per participant, and Blazepose-generated skeleton tracks. The three card categories (green, yellow, red; scores +1, +3, +5) are the intensity-control mechanism meant to produce a spread of genuine emotional responses. The labeling machinery is three-layer: immediate self-annotation, immediate partner-annotation, and delayed external-observer annotation over time ranges. The paper states that the external raters used a data-driven approach relying on the self and partner annotations, so the third layer is built on the two in-game label sets rather than being fully separate from them.","core_discovery":"On its own terms, the paper's central discovery is a collection protocol and dataset: CG-MER captures genuinely spontaneous emotional expressions by having two people play a card game whose red, yellow, and green cards vary the intimacy of the questions. The resulting ten hours of interaction are stored per participant as RGB video, depth video, audio, and a skeleton video, with external-observer annotations in JSON and CSV plus the in-game self and partner labels. The paper presents this as a new benchmark for emotion recognition in French, one that offers synchronized multimodal evidence of how emotions appear in real social interaction rather than in posed recordings. It also positions the card game's intensity gradient as a way to study emotional dynamics, since questions with scores +1, +3, and +5 were designed to elicit progressively stronger responses.","pith_inferences":["Because the external raters were told to base their judgments on the participants' self and partner annotations, the three 'perspectives' are not statistically independent; the external labels are best treated as a derived or smoothed variant of the two in-game labels, not as a third source of ground truth.","The intensity gradient of the cards (green/yellow/red) is a natural independent variable for testing whether emotion-recognition accuracy rises with elicitation intensity, a comparison the paper does not itself run.","The dataset is not publicly available at the time of writing, so its benchmark value depends on release; a useful next step would be a public subset with fixed train/test splits and baseline unimodal results.","A future version aimed at neurodegenerative populations, as the conclusion proposes, would need to re-examine whether the same card questions are emotionally appropriate and whether the annotation scheme captures atypical expression."],"forward_implications":["A model trained on CG-MER can be evaluated separately on facial video, depth, audio, or skeleton streams, and then on their fusion, giving a common benchmark for comparing unimodal and multimodal emotion recognition.","Because the corpus is French and spontaneous, it offers a testbed for whether emotion recognition systems built on English or acted data transfer to conversational French dyads.","The card color and score design makes it possible to relate question intensity to the strength or clarity of the emotion expressed, something most existing corpora do not encode.","The time-range external annotations plus per-question self/partner labels allow both discrete event-level and continuous segment-level emotion modeling.","If transcripts are added, the same recordings could support text-based and audio-text fusion models, an extension the authors explicitly mention."],"supporting_citations":[{"why":"Survey of multimodal emotion databases used to argue that unimodal datasets miss part of the emotional spectrum, motivating the dataset design.","marker":"[10]"},{"why":"The debate dataset that uses self, partner, and external annotations, providing the annotation precedent CG-MER adapts to a card game.","marker":"[17]"},{"why":"A multimodal group-affect dataset with self and external labels, used as a comparison point in the dataset overview table.","marker":"[18]"},{"why":"A widely used dyadic audiovisual benchmark that CG-MER compares against and tries to complement with spontaneous French data.","marker":"[21]"},{"why":"A spontaneous emotional corpus that includes French speakers, supporting the need for a dedicated French dataset.","marker":"[24]"},{"why":"The Blazepose algorithm used to generate the skeleton modality from the video recordings.","marker":"[25]"},{"why":"Argues that multiple annotation sources improve ground truth, directly supporting the three-perspective annotation design.","marker":"[27]"},{"why":"Shows joint modeling of self-reported and perceived emotion, supporting the decision to collect both self and partner labels.","marker":"[28]"}],"fun_headline_variants":["Card game captures spontaneous French emotions in multimodal dataset","Emotion recognition gets a card game dataset: CG-MER","French card game yields spontaneous emotion data for AI","Spontaneous emotions from French card play: new multimodal dataset","Card game yields ten hours of multimodal French emotion data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the card game made participants genuinely feel and naturally display the emotions they report, and that the self- and partner-based labels—which the external observers were told to use as their basis—therefore constitute reliable ground truth for emotion recognition.","fun_headline_variants_meta":{"raw":{"variants":["Card game captures spontaneous French emotions in multimodal dataset","Emotion recognition gets a card game dataset: CG-MER","French card game yields spontaneous emotion data for AI","Spontaneous emotions from French card play: new multimodal dataset","Card game yields ten hours of multimodal French emotion data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001214,"raw_usage":{"total_tokens":4946,"prompt_tokens":842,"completion_tokens":4104,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":458,"completion_tokens_details":{"reasoning_tokens":4027}},"tokens_in":458,"tokens_out":4104,"duration_ms":25647,"temperature":1.0,"reasoning_tokens":4027,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:28:27.414753+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One could test the independence of the external annotations by asking fresh raters to label the same videos without ever seeing the self or partner annotations and measuring agreement; if those independent ratings are no closer to the recorded behavior than chance, or if the published external labels simply reproduce the self/partner majority, the dataset's claim of providing a meaningful observer perspective would be refuted. A more direct check would be to run a replication of the card game, record it, and compare participants' self-labels with blind viewers' ratings of the same moments.","supporting_citations":[{"cited_title":"A survey on databases for multimodal emotion recognition and an introduction to the viri (visible and infrared image) database,","cited_arxiv_id":null,"evidence_quote":"Survey of multimodal emotion databases used to argue that unimodal datasets miss part of the emotional spectrum, motivating the dataset design."},{"cited_title":"K - emocon, a multimodal sensor dataset for continuous emotion recognition in naturalistic conversations,","cited_arxiv_id":null,"evidence_quote":"The debate dataset that uses self, partner, and external annotations, providing the annotation precedent CG-MER adapts to a card game."},{"cited_title":"Amigos: A dataset for affect, personality and mood research on individuals and groups,","cited_arxiv_id":null,"evidence_quote":"A multimodal group-affect dataset with self and external labels, used as a comparison point in the dataset overview table."},{"cited_title":"Iemocap: Interactive emotional dyadic motion capture database,","cited_arxiv_id":null,"evidence_quote":"A widely used dyadic audiovisual benchmark that CG-MER compares against and tries to complement with spontaneous French data."},{"cited_title":"Percep tual borderline for balancing multi -class spontaneous emotional data,","cited_arxiv_id":null,"evidence_quote":"A spontaneous emotional corpus that includes French speakers, supporting the need for a dedicated French dataset."},{"cited_title":"Recording affect in the field: Towards methods and metrics for improving ground truth labels,","cited_arxiv_id":null,"evidence_quote":"Argues that multiple annotation sources improve ground truth, directly supporting the three-perspective annotation design."},{"cited_title":"Automatic recognition of self -reported a nd perceived emotion: Does joint modeling help?,","cited_arxiv_id":null,"evidence_quote":"Shows joint modeling of self-reported and perceived emotion, supporting the decision to collect both self and partner labels."}],"review_version":1}