{"id":"de9376ca-6704-4296-aa2e-a9ee37c5d2f8","arxiv_id":"2505.14143","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"MMoLRE, a low-rank mixture-of-experts model for joint multimodal sentiment and emotion prediction, reports state-of-the-art sentiment scores on CMU-MOSI and CMU-MOSEI while emotion recognition remains competitive but not best.","lead":"This paper proposes a multi-task model that jointly performs sentiment analysis and emotion recognition using shared and task-specific low-rank expert networks. On the CMU-MOSI and CMU-MOSEI benchmarks it reports improved sentiment accuracy over prior published results, with competitive emotion recognition.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The CMU-MOSI SOTA result is confounded because the emotion pseudo-labels are generated using the sentiment labels being predicted; the reported MSA gain may come from target-derived auxiliary supervision rather than from the MMoLRE architecture.","rationale":"The reader's conditional verdict identifies pseudo-label quality as the weakest assumption; I sharpen that concern to a structural confound. The polarity filter in the label-unification method makes the auxiliary emotion labels a deterministic transform of the sentiment target, so the CMU-MOSI comparison is not architecture-only. This does not require assuming author misconduct; it is a design property of the protocol. The proposed permutation test would settle whether the matched pseudo-labels are actually responsible for the reported gains. If they are, the 'state-of-the-art on MSA' claim must be narrowed to CMU-MOSEI or the method must be re-benchmarked against a same-feature MMML baseline that also uses the pseudo-labels. The reader's CONDITIONAL verdict remains appropriate, so I recommend no change to the verdict: with code, variance reporting, and the pseudo-label control, the paper could be accepted; without these, the headline empirical claim is not yet established.","tokens_in":9500,"tokens_out":15715,"duration_ms":158329,"concrete_test":"Re-train the exact CMU-MOSI configuration of MMoLRE with the pseudo emotion labels randomly permuted across training samples (preserving marginal label frequencies and the same train/validation split), so the auxiliary labels are no longer matched by sentiment polarity. If Acc5/Acc7 stay near 56.51/47.52 and above MMML's 55.83/47.03, the matching protocol is not the driver; if they drop toward or below MMML, the reported SOTA depends on the target-derived pseudo-label procedure and the headline claim should be restricted to CMU-MOSEI until a controlled re-implementation is available.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline MSA claim includes CMU-MOSI (Table I), but CMU-MOSI has no native emotion labels. Section IV-B says samples are first classified by sentiment polarity and then SimCSE assigns emotion labels from the most similar same-polarity sample in MOSEI/MELD/IEMOCAP. The auxiliary MER targets are therefore not independent labels: the polarity filter makes them a function of the primary sentiment regression target. Table I compares against single-task MSA baselines such as MMML that receive no such target-derived auxiliary supervision, and on CMU-MOSI MMoLRE does not even beat MMML on MAE (0.666 vs 0.663) or Corr (0.837 vs 0.837). Because Section IV-C explicitly credits the pseudo-labels for the MOSI result, the Acc5/Acc7/Acc2 gains cannot be attributed to the shared/task-specific low-rank expert design unless the pseudo-label protocol is controlled. The Fig. 3 ablation compares MTL sharing variants, but the text does not report numeric same-feature baselines that separate the label-generation effect from the architecture effect.","agreement_with_reader":"partial"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the architecture is plausible and the MOSEI results are interesting, but the paper's headline claim about CMU-MOSI is confounded. The pseudo-labels for emotion on CMU-MOSI are generated by first filtering samples by sentiment polarity, then matching with SimCSE to other datasets. That makes the auxiliary emotion labels a function of the sentiment regression target, so the MSA gains on MOSI (Acc5/Acc7/Acc2) could come from the leaked supervision rather than from the shared/task-specific low-rank experts. The paper even says 'leveraging pseudo-labels for emotion classification leads to SOTA results,' which acknowledges the effect. The comparison to single-task MMML is therefore unfair on MOSI; the ablation in Fig. 3 is on MOSEI with real labels, so it doesn't control for this. The MOSEI results are cleaner, but even there MMoLRE loses to MMML on Acc2Non0 and F1Non0, and the claim of SOTA relies on selected metrics. Also, hyperparameters n, r_n, k are chosen from test-set sensitivity curves, so the reported numbers are partially tuned on the test set. No standard deviations, no code release.\n\nWhat's genuinely useful: the low-rank expert design cuts parameters and FLOPs by roughly 6x versus standard MoE, the method is clearly described, and the ablation showing early separation before fusion helps is a reasonable data point. The writing is honest about the PLE similarity and the lack of context modeling for MER.\n\nOne more thing: the model uses only text and audio, while most baselines (including MMML) are tri-modal. That's an extra confound in the comparison, though it cuts against the paper in the sense that two modalities beating three is a surprise.\n\nBottom line: the paper deserves peer review, but it needs a major revision. Ask for code, standard deviations, a MOSI ablation where a single-task baseline gets the same auxiliary pseudo-label loss, and hyperparameter selection on validation. If the MOSEI gains survive a stricter protocol, there's a decent empirical contribution here.","headline":"A plausible low-rank MoE for multimodal affect MTL, but the CMU-MOSI SOTA is partly an artifact of pseudo-label leakage and the paper needs controlled ablations, variance, and code before the claims hold.","tokens_in":10284,"tokens_out":4217,"would_cite":false,"duration_ms":40564,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-07T15:39:02.764226+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}