{"id":"39863528-fb0a-4ece-9ca3-4cf027cf1386","arxiv_id":"2505.06685","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Emotion-Qwen reports a dual-expert multimodal architecture and a 40K-clip VER dataset that achieve strong emotion recognition while preserving general vision-language performance.","lead":"Emotion-Qwen is a new multimodal model that adds a facial-expression capture module and a mixture-of-experts visual compressor to keep emotion understanding from erasing general vision-language skills. It also contributes a 40,000-clip bilingual video emotion reasoning dataset, and reports state-of-the-art scores on emotion benchmarks while retaining strong general benchmark performance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Instruction fine-tuning and evaluation splits are undocumented; VER is built from MER2024/MAFW and the same benchmarks are then evaluated, so the SOTA numbers may reflect train/eval overlap.","rationale":"The strongest claim is empirical: Emotion-Qwen is SOTA on emotion benchmarks while retaining general VL ability. Every component of that claim is a number in Tables 2 and 4, and the least secure condition for those numbers is benchmark hygiene, specifically disjointness between training and evaluation. The paper builds VER from MAFW and MER2024, instruction-tunes on VER plus DFEW and EmoViT/EmoSet, then evaluates on MER2024, DFEW, EMER, and EmoSet. No split protocol or video-ID-level exclusion is given anywhere. If even a subset of the 36,357 MER2024-derived VER clips appears in the MER2024 test set, the 85.47 SEMI score is not a fair SOTA; the same logic applies to DFEW and EmoSet if their test folds were used in LoRA training. The paper does have real supporting evidence: detailed ablations (Tables 5-8), a parameter-counted projector comparison, and internally consistent multi-LoRA results. Those support the architecture and the forgetting-avoidance story, but they do not validate the external SOTA comparisons. The EMER ChatGPT-scored metric is a secondary concern: without human agreement or blinding, the 8.25/8.16 record is not established as better emotional reasoning. I therefore align with the reader's conditional verdict: the paper should be accepted only after the authors document or release the exact splits and overlap computation, and ideally add human validation of EMER scoring. This does not move the verdict; it sharpens the condition that must be met.","tokens_in":13840,"tokens_out":9666,"duration_ms":93760,"concrete_test":"Obtain the sample/video IDs used to build VER (8,034 MAFW + 36,357 MER2024) and the samples used for each DFEW/EmoSet LoRA adapter, then compute their overlap with the evaluation sets for MER2024 SEMI/NOISE, DFEW, EMER, and EmoSet. If any overlap exists, rerun Tables 2 and 4 after removing those samples from training; if the reported margins over prior SOTA (e.g., 78.31 vs 77.51 WAR on DFEW, 85.47 vs 78.80 on MER2024 SEMI) shrink or vanish, the SOTA claim is invalid. A clean overlap computation with a released ID list would resolve the concern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline SOTA comparisons in Tables 2 and 4 are only valid if instruction fine-tuning data and evaluation data are disjoint. The 'Emotional Instruction Fine-tuning' section says the model is fine-tuned on VER (built from MAFW and MER2024) plus DFEW and EmoViT/EmoSet; the 'Instruction Fine-tuning Evaluation' section then reports MER2024 SEMI/NOISE, DFEW WAR/UAR, EMER Clue/Label, and EmoSet accuracy. The paper never states the split protocol, never lists video IDs excluded from VER, and never says whether DFEW/EmoSet LoRA training samples are held out from the reported test folds. Because VER is built from MER2024, and the LoRA recipe includes DFEW/EmoSet, there is a direct route for training/evaluation overlap to inflate the signature numbers (85.47/79.67, 78.31/62.11, 85.49). The ablations do not close this gap, and no error bars are given. The EMER 8.25/8.16 record is additionally a ChatGPT overlap measure with no human agreement or blinding, so it may reflect stylistic fluency rather than emotional reasoning. Both gaps are fixable, but as written the SOTA claim is not independently checkable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Emotion-Qwen, a 7B video-capable large multimodal model that combines a Facial Emotion Capture (FEC) module with a Mixture-of-Experts Hybrid Compressor to route visual tokens through emotion-specialized and general-purpose experts. The authors introduce a three-stage pretraining pipeline, a new bilingual Video Emotion Reasoning (VER) dataset of over 40K clips, and a multi-LoRA instruction fine-tuning strategy. They report state-of-the-art results on emotion benchmarks (DFEW, MER2024, EMER, EmoSet) while retaining competitive performance on general vision-language benchmarks such as MMBench and TextVQA.","tokens_in":14135,"tokens_out":5540,"duration_ms":54849,"significance":"If the reported results are reproducible and the evaluation protocol is clean, the paper makes a useful contribution: it demonstrates a concrete architecture for mitigating catastrophic forgetting when adapting large multimodal models to emotion-centric tasks, contributes a substantial new emotion-reasoning dataset, and provides component-level ablations for the FEC module, the Hybrid Compressor, and the multi-LoRA training strategy. The release of code and weights is a further strength, and the routing-pattern analysis in Table 7 is a valuable sanity check for the MoE design. However, the central SOTA claims are not independently checkable as written because the instruction fine-tuning and evaluation split protocol is not documented, and the EMER metric relies on ChatGPT scoring without reported human validation.","major_comments":[{"comment":"The paper never documents how the instruction fine-tuning data and the evaluation data are disjoint. The fine-tuning section states that VER (built from MER2024 and MAFW) is used together with DFEW and EmoViT, and the evaluation section then reports MER2024 SEMI/NOISE, DFEW WAR/UAR, EMER, and EmoSet accuracy. Since VER is constructed from MER2024, and since DFEW and EmoSet are used in fine-tuning, there is a direct route for training/evaluation overlap to inflate the headline numbers (85.47/79.67, 78.31, 85.49). The manuscript must specify the exact train/test splits, state which video IDs were excluded from VER, and describe which LoRA adapter is used for each evaluation benchmark. Without this, the SOTA claims in Table 4 are not verifiable.","section":"Emotional Instruction Fine-tuning / Instruction Fine-tuning Evaluation, Tables 2 and 4"},{"comment":"The EMER Clue and Label Overlap scores are assessed by ChatGPT, but the paper reports no human agreement, no blinding protocol, and no correlation between ChatGPT scores and human judgments. The 8.25/8.16 EMER results are central to the claim of state-of-the-art emotional reasoning; if the metric rewards stylistic fluency rather than substantive emotional inference, the comparison is not meaningful. The authors should add a human-evaluation study or report the official EMER scoring procedure and its reliability, and ideally compare ChatGPT scores against human annotations on a subset.","section":"Table 2 note and Instruction Fine-tuning Evaluation"},{"comment":"No error bars, confidence intervals, or multiple-seed results are reported anywhere. Several headline margins are small (e.g., 78.31 vs. 77.51 WAR on DFEW in Table 4), and the multi-LoRA vs. single-LoRA comparison in Table 8 could depend on initialization and adapter selection. The authors should report at least three runs for the fine-tuning experiments and, where applicable, for the ablations, so that the claimed improvements can be distinguished from run-to-run variance.","section":"General experimental reporting, Tables 2, 4, 5, and 8"}],"minor_comments":[{"comment":"The text states that FEC yields a 1.92% gain on DFEW, but Table 5 shows only +0.10 for WAR and +1.82 for UAR; the reported number appears to be the sum of the two metric gains. Please clarify whether the text refers to a combined score or to a single metric.","section":"Ablation Study, Table 5"},{"comment":"The phrase 'such as DFEW and EmoViT' calls EmoViT a dataset, but elsewhere EmoViT is presented as a model (Xie et al., 2024). The authors should clarify which concrete emotion datasets are used for instruction fine-tuning (e.g., EmoSet, AffectNet, or another resource) and cite the corresponding dataset papers.","section":"Emotional Instruction Fine-tuning"},{"comment":"The table header 'H-Params' should be 'Hyperparameters', and 'DeepSpeed Zero2' should be written as 'ZeRO-2' for consistency with the cited DeepSpeed paper.","section":"Table 3"},{"comment":"The table has a typo ('Emoiton Expert') and the column headers are not fully legible in the submitted text; please format the table so that each benchmark has clear 'Emo' and 'Gen' subcolumns.","section":"Table 7"},{"comment":"Since code and weights are promised, please include the exact dataset versions, filtering thresholds, and the prompts used for VER construction and for ChatGPT-based EMER scoring in the supplementary material.","section":"Reproducibility"}],"recommendation":"major_revision","confidential_remarks":"The main risk is data overlap between instruction fine-tuning and evaluation, especially because VER is built from MER2024 and the paper evaluates on MER2024. This is fixable with a clear split protocol, but without it the SOTA numbers cannot be trusted. I would also like the editor to consider whether the novelty of the VER dataset is sufficiently distinguished from the AffectGPT data-collection pipeline, which uses a similar model-guided, human-assisted filtering approach."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent empirical system paper with a genuinely new resource and a plausible architecture, but the load-bearing SOTA claims currently rest on undocumented splits and small differences without significance checks. I would send it to peer review, but only with a required revision that documents the train/eval disjointness and releases the artifacts.\n\nWhat is actually new: the VER dataset is a real contribution. 40K+ clips, 80K bilingual reasoning annotations, built from MAFW and MER2024 with TalkNet filtering and human ratings. That is valuable regardless of the rest. The Hybrid Compressor (two MLP experts with attention gating) plus the FEC face-focused input is a reasonable variation on MoE connectors, and the ablations are honest and useful: FEC helps emotion tasks, HC beats MLP and Fusion, multi-LoRA beats FFT and single LoRA. The routing analysis in Table 7 supports the expert-balance story, and the LLaMA 3 backbone swap shows the method is not Qwen-only. The citation pattern is standard for the area; no obvious gaps.\n\nThe soft spots, in order of severity. First, the split overlap issue is real. VER is built from MAFW and MER2024, and the instruction fine-tuning recipe includes DFEW and EmoSet; the evaluation then reports MER2024, DFEW, and EmoSet numbers. The paper never states which videos are excluded, so if the full MER2024/DFEW/EmoSet videos feed VER and the same benchmark folds are evaluated, the SOTA deltas are inflated. This is the central concern and it lands. Second, EMER Clue/Label scores are ChatGPT overlap metrics; reporting new records on them without human agreement or blinding is weaker evidence than the prose implies. Third, no error bars. DFEW 78.31 vs 77.51 and EmoSet 85.49 vs 83.36 are differences that need significance testing. Minor point: the conclusion claims public weights/code, but only an anonymous repo link is given and no data link appears; that should be fixed.\n\nWho this is for: people building emotion-adapted LMMs and affective computing researchers who might reuse VER. The architecture is incremental, but the dataset is useful and the forgetting-control comparison is worth seeing. It deserves a serious referee, and I would conditionally recommend acceptance with the splits documented, error bars added, and artifacts released.","headline":"A solid system paper with a genuinely new dataset (VER) and a plausible connector design, but the headline SOTA numbers are not independently checkable until the fine-tuning/evaluation split is documented and error bars are added.","tokens_in":14663,"tokens_out":3472,"would_cite":false,"duration_ms":36168,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Emotion-Qwen claims that emotional understanding and general vision-language reasoning can coexist in a single 7B multimodal model, and reports state-of-the-art scores on both families of benchmarks.","keywords":["multimodal emotion recognition","vision-language models","mixture of experts","catastrophic forgetting","emotion reasoning","Video Emotion Reasoning dataset","instruction tuning","facial expression recognition"],"falsifier":"Run the instruction-tuning evaluation on a cleaned version of the benchmarks from which every clip that resembles a VER training clip (same source video, scene, or speaker from MAFW or MER2024) has been removed; if the reported margins over Emotion-LLaMA and MMA-DFER shrink or disappear, the balanced-capability claim is not established. Independently, collect human ratings on a random sample of EMER responses and compare them with the ChatGPT-assigned Clue/Label scores; low agreement would put the emotion-reasoning metric in doubt.","tokens_in":13651,"feed_emoji":"🎭","tokens_out":9866,"duration_ms":90102,"temperature":0.7,"pith_summary":"This paper argues that emotion understanding and general vision-language ability do not have to be traded off in a multimodal model. It introduces Emotion-Qwen, a 7B vision-and-text model whose visual tokens are routed through two experts—one tuned for emotional content, one for general scenes—and which is pre-trained in three stages so the emotion expert can be added without erasing general knowledge. To support fine-grained reasoning, the authors build the Video Emotion Reasoning (VER) dataset, over 40,000 bilingual clips with context-aware annotations. The reported results are state-of-the-art on emotion benchmarks (DFEW, EMER, EmoSet, MER2024) while MMBench, TextVQA, and ScienceQA remain competitive with the best open general LMMs. If correct, this gives a practical recipe for emotion-specialized assistants that do not lose general competence.","feed_headline":"One 7B model tops emotion tests, keeps general vision skill","feed_subtitle":"Emotion-Qwen reports state-of-the-art results on DFEW, EMER, EmoSet, and MER2024 while holding MMBench at 87.3.","key_machinery":"The load-bearing mechanism is the Hybrid Compressor, a mixture-of-experts projector with two MLP experts—an Emotion Expert and a General Expert—whose outputs are combined by an attention-based gating network: $V_{out} = G \\odot V_{emo} + (1-G) \\odot V_{gen}$. The gate decides per input how much of the visual representation should be processed emotionally versus generally, and an ablation records task-dependent routing (for example, 0.63 of the gate weight goes to the Emotion Expert on MER2024 versus 0.37 on MMBench). Around this sits the Facial Emotion Capture module, which keeps key facial-expression frames and masks backgrounds, and a three-stage pretraining schedule that first aligns the general expert, then warms the emotion expert, then fine-tunes the LLM on instruction data. During emotion instruction tuning, separate low-rank adapters (LoRA) per dataset let the Qwen2.5 backbone specialize without full retraining.","core_discovery":"On the paper's own terms, the central discovery is that a single 7B model can specialize in emotion without generalizing worse: Emotion-Qwen reaches 87.3 on MMBench, 87.9 on TextVQA, and 77.2 on ScienceQA in zero-shot evaluation, and after instruction fine-tuning with per-dataset LoRA adapters it reports 78.31 WAR on DFEW, 8.25 Clue and 8.16 Label overlap on EMER, 85.49 accuracy on EmoSet, and 85.47/79.67 on MER2024 SEMI/NOISE. This is achieved with vision and text only—no audio—beating audio-capable emotion models. The mechanism is a dynamic routing of compressed visual tokens between an emotion expert and a general expert, guided by an attention-based gate, so the model can keep scene-level reasoning while adding facial-emotion analysis. The authors credit the new VER dataset and the staged pretraining pipeline for making the emotion expert learn without damaging the general one.","pith_inferences":["The routing weights suggest a softer, input-dependent specialization than a hard switch; probing the experts' internal representations for facial-action versus scene-layout features would test whether the two experts truly encode complementary information.","Because the model uses vision and text only, adding an audio expert is a natural next experiment; if the bottleneck is modality coverage, EMER and MER2024 NOISE scores should rise with speech input.","Since Clue and Label overlap scores come from an LLM judge, a human-agreement study on a sample of EMER outputs would tell whether the metric reflects human judgments of emotional reasoning quality."],"forward_implications":["A 7B open-source model can match or beat far larger systems on both emotion and general vision-language benchmarks, so emotion-specialized assistants no longer have to sacrifice general competence.","Fine-grained emotion reasoning can be trained and evaluated as a vision-and-text task, and the VER dataset provides over 40,000 bilingual clips with human-verified explanations for doing so.","Per-dataset LoRA adapters on a frozen backbone make it practical to extend the model to new emotion domains without full retraining.","Task-dependent gating means one deployed model can serve mixed workloads, weighting the emotion expert for affective queries and the general expert for scene-level ones.","Reported gains on DFEW, EMER, EmoSet, and MER2024, if the evaluation splits are clean, indicate that catastrophic forgetting during emotion fine-tuning is avoidable rather than inevitable."],"supporting_citations":[{"why":"Supplies the Emotion-LLaMA baseline that Emotion-Qwen must beat on emotion reasoning and the instruction-tuning approach it extends.","marker":"Cheng et al. 2024"},{"why":"Provides the MER2024 dataset used both as a source of VER video clips and as the SEMI/NOISE benchmark for fine-tuned evaluation.","marker":"Lian et al. 2024b"},{"why":"Gives the AffectGPT prior state of the art and the two-stage filtering recipe (TalkNet plus human-assisted refinement) reused to build VER from MER2024.","marker":"Lian et al. 2025"},{"why":"Defines the EMER benchmark with Clue and Label overlap scores that are the basis for the emotion-reasoning state-of-the-art claim.","marker":"Lian et al. 2024a"},{"why":"Supplies the DFEW dynamic facial expression dataset used for instruction fine-tuning and for the WAR/UAR benchmark results.","marker":"Jiang et al. 2020"},{"why":"Provides the MAFW dataset, the second source of VER video clips with emotion annotations and facial action descriptions.","marker":"Liu et al. 2022"},{"why":"Supplies the EmoViT baseline and the EmoViT instruction-tuning data used in the emotional fine-tuning stage.","marker":"Xie et al. 2024"},{"why":"Serves as the Qwen2-VL general vision-language baseline and a design reference for vision-language alignment.","marker":"Wang et al. 2024"},{"why":"Supplies the Qwen2.5 LLM backbone whose frozen general capabilities are central to the catastrophic-forgetting comparison.","marker":"Yang et al. 2024"},{"why":"Provides GPT-4, used to generate VER annotations and as a proprietary zero-shot comparison model.","marker":"Achiam et al. 2023"}],"fun_headline_variants":["Emotion-Qwen: no audio, still tops emotion benchmarks","One 7B model: emotion SOTA plus general vision intact","MoE gate lets Emotion-Qwen master emotion, keep vision","Emotion-Qwen: vision-text only beats audio emotion models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the evaluation uses clean train/test separation, so the VER training clips drawn from MAFW and MER2024 do not overlap the DFEW, EmoSet, MER2024, or EMER test sets, and that the ChatGPT-assigned Clue/Label overlap scores faithfully measure emotional reasoning quality.","fun_headline_variants_meta":{"raw":{"variants":["Emotion-Qwen: no audio, still tops emotion benchmarks","One 7B model: emotion SOTA plus general vision intact","MoE gate lets Emotion-Qwen master emotion, keep vision","Emotion-Qwen: vision-text only beats audio emotion models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000772,"raw_usage":{"total_tokens":3436,"prompt_tokens":983,"completion_tokens":2453,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":2381}},"tokens_in":599,"tokens_out":2453,"duration_ms":18476,"temperature":1.0,"reasoning_tokens":2381,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:35:54.061306+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the instruction-tuning evaluation on a cleaned version of the benchmarks from which every clip that resembles a VER training clip (same source video, scene, or speaker from MAFW or MER2024) has been removed; if the reported margins over Emotion-LLaMA and MMA-DFER shrink or disappear, the balanced-capability claim is not established. Independently, collect human ratings on a random sample of EMER responses and compare them with the ChatGPT-assigned Clue/Label scores; low agreement would put the emotion-reasoning metric in doubt.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Emotion-LLaMA baseline that Emotion-Qwen must beat on emotion reasoning and the instruction-tuning approach it extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the DFEW dynamic facial expression dataset used for instruction fine-tuning and for the WAR/UAR benchmark results."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the MAFW dataset, the second source of VER video clips with emotion annotations and facial action descriptions."}],"review_version":1}