{"id":"39170f13-1110-465f-bc3b-0738d8dee7a6","arxiv_id":"2502.06803","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A broad survey of emotion recognition and generation across three modalities, but its reliability is hampered by misreported comparative results and citation errors.","lead":"This paper surveys artificial intelligence methods for recognizing and generating emotions from faces, speech, and text. It compiles datasets, preprocessing techniques, model architectures, and evaluation metrics to help newcomers navigate the field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central claim of reliable, comprehensive guidance is undermined by inaccuracies in its comparative evaluation tables and their textual interpretation; a targeted re-verification of Table 3 and §6.2.2 determines whether the concern lands.","rationale":"I read the paper as a survey whose only original contribution is curation and synthesis; it introduces no new method or result. The central claim is therefore exactly the reliability and comprehensiveness of that curation. The reader's weakest assumption—that cited numbers and references are transcribed accurately—is the load-bearing one, and it fails in the paper's own evaluation section. I picked Table 3/§6.2.2 as the sharpest instance because the numerical comparison is internally inconsistent (10.31 is called 'lower' than 9.38) and because the metrics in that table are heterogeneous in ways that are not disclosed. This is not a matter of disagreeing with the field's consensus or asking for a different scope; the survey contradicts itself. The concrete test is a targeted re-verification, not a re-review of the whole field, and it would settle whether the issue is isolated typos or systemic misreporting. If the test shows the table matches sources and only the prose is garbled, a correction could restore some value; as written, the review cannot serve the newcomer audience it targets. No change to the reader's REJECT is needed, hence UNCHANGED.","tokens_in":28317,"tokens_out":5386,"duration_ms":52432,"concrete_test":"Take every row of Table 3 that is sourced from [147], [105], [99], [153], and [119], retrieve the original papers, and record the metric name, scale, and value as reported in each source; then check whether the Table 3 ACC/FID/SyncNet columns are homogeneous across rows and whether the §6.2.2 sentence comparing SadTalker (10.31) and Wav2Lip (9.38) is consistent with those original values. If the ACC column mixes different metrics or the ordering is reversed, the comparative analysis is unreliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that it gives newcomers a reliable map of SOTA methods, datasets, and metrics across emotion recognition and generation. That claim stands or falls on accurate transcription and interpretation of cited results, especially in the promised comparative analyses. Section 6.2.2/Table 3 is the weakest point. The prose says SadTalker has a 'lower ACC (10.31)' than Wav2Lip's 9.38, which is numerically backwards (10.31 > 9.38), and the ACC column mixes entries such as 0.8, 9.38, 58.8, and 75.43 with no indication that they measure the same quantity, so the table cannot support the comparative statements built on it. Section 6.2.5/Table 6 shows the same failure mode: the text credits Emotion BERT with an F1 of 0.88 on EmotionLines, while the table lists only ACC=0.71 for that model, so the claimed SOTA result is unverifiable from the review itself. Section 5.4.2 describes a text-driven talking-head framework and cites [114], which is StarGANv2-VC, a voice-conversion paper, leaving the described method unattributed. These are internal inconsistencies in the review's own reporting, not disputes with the field's consensus, and they directly compromise the stated purpose of orienting researchers beginning in the area.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey of emotion recognition and generation across face, speech, and text modalities. It covers preprocessing techniques, datasets, state-of-the-art methods for recognition and generation, evaluation metrics, comparative analyses, and future research directions. The stated goal is to provide a holistic, integrated review that helps researchers beginning in the field, addressing what the authors describe as a gap in the literature covering these two domains together.","tokens_in":28543,"tokens_out":4077,"duration_ms":36645,"significance":"If the survey were accurate, it would fill a useful niche: most prior reviews treat emotion recognition and generation separately, and a single structured overview spanning modalities would be valuable to newcomers. The paper has broad coverage of datasets, methods, and metrics, and it explicitly discusses applications and open challenges. However, the survey's usefulness depends critically on the correctness of its comparative tables and their interpretation; the reported numbers and citations currently contain several internal inconsistencies that undermine the central claim of providing reliable comparative guidance. The paper does not include code or machine-checked proofs, but that is not expected for a survey; the burden instead lies on accurate reporting of cited work.","major_comments":[{"comment":"The prose states that SadTalker achieves a lower ACC (10.31) than Wav2Lip's 9.38, which is numerically incorrect since 10.31 > 9.38. More generally, the ACC column mixes values from different evaluation protocols (0.8, 9.38, 58.8, 75.43) without any explanation of what is being measured, so the table cannot support the comparative conclusions drawn from it. This directly affects the review's stated goal of offering comparative analyses.","section":"§6.2.2, Table 3"},{"comment":"The text credits Emotion BERT [176] with the highest F1 score of 0.88 on EmotionLines, but Table 6 reports only ACC=0.71 for that model and no F1 score. The same paragraph attributes an F1 score of 0.47 and accuracy of 0.5 to AutoVC on the ESD dataset, but AutoVC and ESD appear in Table 5 (speech generation), not in Table 6 (text sentiment recognition). These mismatches make the claimed SOTA results unverifiable from the review itself.","section":"§6.2.5, Table 6"},{"comment":"The text-driven talking-head framework attributed to reference [114] is described in detail (components Gmou, Gupp, Ghed, Gldmk), but reference [114] is StarGANv2-VC, a voice-conversion paper. The described method is therefore left without a correct citation, undermining the reliability of the method categorization for readers trying to locate the original work.","section":"§5.4.2"},{"comment":"The prose states that FreeVC achieves the lowest WER (5.4%) and EER (11.28%) on LibriSpeech, but Table 5 lists FreeVC with EER 35.63 and Phoneme Hallucinator with a lower WER (5.1). This is another instance where the textual interpretation contradicts the table data, making the comparative discussion internally inconsistent.","section":"§6.2.4, Table 5"}],"minor_comments":[{"comment":"The phrase 'emotion control methods accross modalities' contains a typo: 'accross' should be 'across'.","section":"§1"},{"comment":"The metric list refers to 'GPQU' but the correct abbreviation is 'GPQA'; also, the statement that 'All of these metrics are obtained from user studies' is inaccurate for MMLU, MATH, HumanEval, MGSM, and DROP, which are benchmark evaluations rather than subjective user studies.","section":"§6.1.3"},{"comment":"The prose says 'The GPT-4 model [28] achieves the highest MMLU score of 88.7%', but Table 7 attributes 88.7 to GPT-4o, not GPT-4; the model name should be corrected for consistency.","section":"§6.2.6, Table 7"},{"comment":"Reference [150] is titled 'Makelttalk: speaker-aware talking-head animation' in the bibliography; the correct title is 'MakeItTalk: Speaker-Aware Talking-Head Animation'.","section":"References"},{"comment":"The FERPlus dataset size is listed as 'Unlimited', which is undefined and not informative; the actual number of images in FER2013/FERPlus should be reported instead.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The reader's report recommends reject, and I agree that the number of transcription errors in the comparative tables is serious. However, these errors appear correctable within the scope of a revision: every table entry and every sentence interpreting a table should be re-verified against the cited source, and mismatched citations (such as [114]) should be fixed. For this reason I favor major_revision over reject. If the authors cannot provide a corrected set of tables and a point-by-point reconciliation of the prose with those tables, rejection would then be appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this survey covers a real gap — both emotion recognition and generation across face, speech, and text — and a newcomer would get a decent sense of the main methods, datasets, and evaluation metrics. But the comparative analysis, which is the core value of a survey like this, has enough internal inconsistencies that I wouldn't rely on any specific number until it's re-verified.\n\nWhat it does well: the organization is sensible (FEG/SEG/TSG taxonomy works), the preprocessing and dataset sections are broad enough to orient a new researcher, and it correctly notes the lack of standardized evaluation in FEG. The idea of integrating recognition and generation in one review is genuinely useful; most prior surveys stay on one side.\n\nThe soft spots are real and they land in the same place the reader's stress-test flagged. In §6.2.2 the text says SadTalker has a 'lower ACC (10.31)' than Wav2Lip's 9.38; numerically that's backwards, and the ACC column in Table 3 mixes quantities (0.8, 9.38, 58.8, 75.43) with no indication they're the same measure. Table 6's text credits Emotion BERT with an F1 of 0.88 on EmotionLines, but the table lists no F1 for that model. §5.4.2 describes a text-driven talking-head framework and cites [114], which is StarGANv2-VC, a voice conversion paper. There are also cross-modality confusions in §6.2.5: AutoVC (a voice conversion model) is described as a text sentiment model, and FERV39K (a video dataset) is discussed in the TSR section. The claim that MMLU/GPQA etc. are 'obtained from user studies' is wrong. These are not one-off typos; they compound into the central promise of a reliable map for beginners.\n\nI don't think the overall concept is flawed. The fix is mechanical but must be thorough: re-check every number in Tables 2–7 against the cited papers, align prose with tables, correct the [114] citation, and add a short methodology note on how papers were selected. If that's done, this becomes a genuinely useful entry point for the field.\n\nBottom line: this deserves peer review, but a major-revision one. I wouldn't cite it in its current form.","headline":"Useful survey concept, but the comparative tables are too unreliable to cite as-is; fixable with a careful re-verification pass.","tokens_in":29003,"tokens_out":3890,"would_cite":false,"duration_ms":33526,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that emotion recognition and emotion generation belong in one holistic survey across face, speech, and text, and delivers that survey with datasets, methods, evaluation metrics, and future directions.","keywords":["emotion recognition","emotion generation","facial expression recognition","speech emotion recognition","text sentiment recognition","talking-head generation","voice conversion","affective computing"],"falsifier":"Locating a published survey that already covers both emotion recognition and emotion generation across face, speech, and text with comparable scope would disprove the paper's gap claim; alternatively, spot-checking every entry in the comparative tables against its source paper and finding systematic misreporting would undermine the reliability claim.","tokens_in":28109,"feed_emoji":"🎭","tokens_out":5490,"duration_ms":53773,"temperature":0.7,"pith_summary":"Emotion recognition and emotion generation have usually been reviewed as separate topics, each limited to one technical approach or one modality. This paper contends that the field needs a single holistic survey and then provides one, covering facial, vocal, and textual modalities for both analysis and synthesis. It walks a newcomer from preprocessing and datasets through state-of-the-art methods, evaluation metrics, comparative tables, challenges, and future directions. If the paper is right, a researcher beginning in affective AI can start from one place and see how recognition and generation share structure across modalities.","feed_headline":"One review bridges emotion recognition and generation","feed_subtitle":"A single map of datasets, methods, and metrics for face, speech, and text.","key_machinery":"The organising device is a two-axis taxonomy: modality (face, speech, text) crossed with task (recognition versus generation), with a separate section for emotion-control methods within generation. Within each cell, the survey groups methods by technical family, such as attention-based, transformer-based, GAN-based, and diffusion-based, plus LLMs for text, and anchors everything to the eight basic emotions derived from Ekman's model. The comparative tables then translate heterogeneous papers into common metrics so that different approaches can be positioned against one another.","core_discovery":"The central claim is that no existing review integrates emotion recognition with emotion generation, and that a review which does so is a useful map for newcomers. On the paper's own terms, it provides that map: it categorises recent state-of-the-art research by technical approach, explains the theoretical foundations of each approach, and compares methods on common metrics such as accuracy, F1 score, FID, WER, and perplexity. It also identifies shared limitations, including scarce and biased datasets, inconsistent evaluation, difficulty of real-time and subtle emotion generation, and ethical risks, and it proposes future directions such as multimodal integration, standardised benchmarks, and responsible deployment.","pith_inferences":["In the editor's reading, the comparative tables are best treated as a guide to typical operating points rather than a strict leaderboard, since the collected results come from different evaluation protocols.","A testable consequence of the survey's structure is that multimodal systems sharing representations across recognition and generation will outperform single-modality pipelines, which the paper names as a future direction but does not itself prove.","One extension a reader could pursue is to build a unified benchmark that scores recognition and generation jointly on the same emotional episodes, directly addressing the standardisation gap the paper identifies."],"forward_implications":["A newcomer can use the survey as a single entry point covering both recognition and generation across all three modalities.","The comparative tables give baseline expectations for performance on widely used datasets such as AffectNet, RAF-DB, IEMOCAP, and LibriSpeech.","The taxonomy shows which technical families dominate each task and where gaps remain, notably text-driven facial expression generation.","The survey's stated challenges define a concrete research agenda: larger diverse in-the-wild datasets, standardised metrics, real-time emotion control, and ethical safeguards."],"supporting_citations":[{"why":"A prior survey of facial expression recognition used to show that recognition-only coverage exists separately from generation.","marker":"[17]"},{"why":"A prior survey of textual emotion recognition used to establish that text recognition has been reviewed apart from generation.","marker":"[18]"},{"why":"A prior survey of speech emotion recognition used to show the same separation for the speech modality.","marker":"[19]"},{"why":"A prior survey of AI-generated content used to represent generation-only coverage, which the paper says does not address recognition.","marker":"[20]"},{"why":"A prior survey of generative models in machine learning used to support the claim that generation has been reviewed separately from recognition.","marker":"[21]"},{"why":"A survey of deep facial expression recognition cited to explain why facial systems dominate and are usually treated alone.","marker":"[22]"},{"why":"Ekman's foundational work supplies the discrete emotion categories that structure the review across modalities.","marker":"[4]"},{"why":"Ekman's later work is cited when the paper states that most emotion recognition systems use the eight primary emotions.","marker":"[5]"}],"fun_headline_variants":["Emotion AI: one review, two tasks, three modalities","Recognition and generation: emotion AI in one survey","Face, voice, text: the all-in-one emotion AI review","Emotion tech: from recognition to generation in one survey","The survey that unites emotion recognition and generation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's value as a starting point assumes that the numbers in its comparative tables and the attribution of methods to cited papers faithfully reproduce what the cited papers actually report.","fun_headline_variants_meta":{"raw":{"variants":["Emotion AI: one review, two tasks, three modalities","Recognition and generation: emotion AI in one survey","Face, voice, text: the all-in-one emotion AI review","Emotion tech: from recognition to generation in one survey","The survey that unites emotion recognition and generation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000407,"raw_usage":{"total_tokens":2065,"prompt_tokens":844,"completion_tokens":1221,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":460,"completion_tokens_details":{"reasoning_tokens":1140}},"tokens_in":460,"tokens_out":1221,"duration_ms":11856,"temperature":1.0,"reasoning_tokens":1140,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T18:20:24.310865+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Locating a published survey that already covers both emotion recognition and emotion generation across face, speech, and text with comparable scope would disprove the paper's gap claim; alternatively, spot-checking every entry in the comparative tables against its source paper and finding systematic misreporting would undermine the reliability claim.","supporting_citations":[],"review_version":1}