{"id":"d5bd38b5-bd9a-46ab-9044-9ed61832ec95","arxiv_id":"2608.10796","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A new video benchmark jointly evaluates expressed and evoked emotion understanding in multimodal LLMs using perception, open-vocabulary recognition, and VAD rating tasks, with Bayesian pairwise alignment for scalable annotation.","lead":"The paper introduces E3mo-Bench, a video benchmark with 12,314 questions that tests whether multimodal AI models understand both emotions expressed by people on screen and emotions a video evokes in a viewer. It also proposes a Bayesian method to collect continuous emotion ratings from simple pairwise comparisons, and shows current AI models handle expressed emotions better than evoked ones, with particular weakness in fine-grained recognition and dominance ratings.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline expressed/evoked performance skew may be an artifact of face-based content routing: the two pools differ in human presence and source dataset, so the observed gap is confounded with content difficulty.","rationale":"The reader's weakest_assumption identifies exactly the structural issue I consider most load-bearing: perspective assignment by face detection yields subsets that differ in content, source, and difficulty, so the main expressed/evoked skew cannot be cleanly attributed to affective perspective. This is not a minor sampling detail. The paper's own limitation statement (Supp. H.1) confirms the lack of paired within-video annotations, which is the strongest independent evidence that the confound is acknowledged and unresolved. I also note that the E3mo-Score validation should state explicitly that the Golden Anchor subset used for committee confidence calibration is disjoint from the samples whose scores are inferred and compared in Table 3; the main text is ambiguous on this point, and an overlap would inflate the reported advantage of E3mo-Score. However, the paper does have real independent support: the human pairwise annotation recovery on expressed valence (SRCC 0.91, PLCC 0.90), the fixed Golden Anchor construction with 15 annotators and ICC reporting, and the BPA ablations showing that iterative comparison improves correlation. These support the benchmark's annotation machinery even though they do not resolve the content confound. Because the confound threatens the headline interpretation but is testable with a matched dual-perspective subset, the appropriate verdict remains CONDITIONAL rather than outright rejection: the central benchmark artifact is plausibly valuable, but the paper's signature comparative claim needs a controlled demonstration before it can be taken at face value.","tokens_in":41503,"tokens_out":3408,"duration_ms":36918,"concrete_test":"Construct a dual-perspective held-out subset: take roughly 200 videos from the expressed pool that contain clear faces and obtain human Golden-Anchor-style OV and VAD annotations under both prompts—'what does the subject express' and 'what does the video evoke in you'—with the evoked instruction explicitly ignoring the subject's face. Then run the same model evaluation protocol on both annotated versions of the identical videos. If the expressed-over-evoked gap shrinks or reverses on this matched set, the headline skew is substantially a content/difficulty artifact; if it persists on identical content, perspective is the causal factor. As a secondary check, report all headline metrics stratified by the easy/medium/hard difficulty tiers used in Section 'Broad-Coverage Set Construction' to rule out sampling imbalance between the two pools.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central finding that MLLMs are better at expressed than evoked emotion rests on comparing two video pools that are constructed by a content proxy, not by affective perspective alone. In 'Video Collection and Candidate Curation', YOLO11n face screening routes samples with a detected face into the expressed pool and face-free samples into the evoked pool; the expressed pool is drawn almost entirely from DFEW, MELD, CAER, and MAFW (actors and dialogue scenes), while the evoked pool comes mainly from VGGSound, LIRIS-ACCEDE, VideoEmotion, and MediaEval (natural scenes, movies, user-generated content). Face presence therefore correlates with genre, shot composition, emotional intensity, and per-dataset labeling difficulty, all of which can affect MLLM accuracy independently of whether the question asks what a subject expresses or what a viewer feels. The paper's own Supplement H.1 concedes that the benchmark does not provide paired annotations for both perspectives on the same video, so the expressed-versus-evoked gap in Tables 1-3 cannot be separated from the distribution shift created by the routing rule. The claim 'models perceive expressed emotions slightly better than evoked emotions' is thus not yet established as a statement about affective perspective; it is at least equally consistent with a statement about human-centered versus non-human-centered video content.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces E$^3$mo-Bench, a benchmark of 12,314 question-answer pairs over 2,524 videos intended to evaluate both expressed and evoked emotion understanding in multimodal large language models (MLLMs). The benchmark organizes evaluation into three tasks: emotion perception, open-vocabulary recognition, and valence-arousal-dominance (VAD) assessment. To obtain scalable continuous VAD labels, the authors propose Bayesian Pairwise Alignment (BPA), which aggregates sparse pairwise human judgments into anchor-referenced scores, and an automated variant, E$^3$mo-Score, which uses a five-model committee as virtual annotators. The paper reports an evaluation of 16 MLLMs and claims that the framework is effective, that BPA-based annotation is reliable and cost-efficient, and that current models exhibit a performance skew between expressed and evoked emotion paradigms.","tokens_in":41724,"tokens_out":6564,"duration_ms":64150,"significance":"If the benchmark and annotation method hold up, this is a potentially useful contribution: it is among the first audio-visual benchmarks to combine expressed and evoked perspectives with open-vocabulary and dimensional annotations, and the pairwise-comparison approach to VAD annotation is a plausible route to scalable affective labels. The Golden Anchor human annotation effort, with 15 annotators and reported ICC values, is a real strength, as is the BPA validation on the expressed-valence subtask (SRCC 0.91, PLCC 0.90). The E$^3$mo-Score committee agent also shows improvements over individual models in several subtasks. However, the central expressed-versus-evoked comparison is currently confounded by the face-based routing rule, and several validation claims outrun the evidence reported in the paper. The benchmark is not released, which limits independent verification.","major_comments":[{"comment":"The central claim that MLLMs are better at expressed than evoked emotion is not identifiable from the reported comparisons. The expressed pool is selected by YOLO11n face detection and drawn almost entirely from DFEW, MELD, CAER, and MAFW, whereas the evoked pool is drawn from VGGSound, LIRIS-ACCEDE, VideoEmotion, and MediaEval. Face presence therefore correlates with source dataset, genre, shot composition, and labeling difficulty, so the differences in Tables 1 and 2 (e.g., GPT-5.4: 58.32% evoked vs. 61.69% expressed overall perception) can be explained by content shift rather than by affective perspective. Supp. H.1 explicitly states that the design 'does not provide paired annotations for both perspectives on every video,' so the expressed/evoked comparison cannot separate perspective from content. The authors should either add a dual-perspective subset with identical videos annotated under both perspectives, or restrict the benchmark's claims to perspective-conditioned evaluation and remove the causal-sounding skew conclusions.","section":"Video Collection and Candidate Curation; Supp. H.1"},{"comment":"BPA validation is reported only for the expressed-valence subtask (SRCC 0.91, PLCC 0.90), while the benchmark and E$^3$mo-Score cover six subtasks (evoked/expressed times valence/arousal/dominance). Since dominance shows markedly lower annotator consistency (ICC(3,k)=0.56) and E$^3$mo-Score's gains are smallest on dominance in Table 3, the claim that BPA 'reliably scales continuous dimensional annotations' across all dimensions is not established. Please provide anchor-recovery or other validation for each of the six subtasks, or explicitly scope the reliability claim.","section":"Human Pairwise Annotation; Supp. C.6"},{"comment":"The evaluation of E$^3$mo-Score is potentially circular. Committee reliability weights are estimated on the Golden Anchor Set (Supp. E.1), and the reported correlations in Table 3 are then computed by masking scores from the same Golden Anchor Set (14 fixed anchors, 369 candidates). If the anchor subset used for weight calibration overlaps the candidate subset used for evaluation, the committee weights are fit to the evaluation data. In addition, candidate scores in BPA are initialized from model-generated bootstrap VAD estimates and pulled toward those values by the adaptive prior in Eq. (8), so E$^3$mo-Score's margin over baselines that predict VAD directly may partly reflect prior information from the same model family. Please clarify the exact disjointness of calibration and evaluation anchors, and add an ablation that removes the model-initialized prior.","section":"The E3mo-Score Agent; Supp. E.1; Table 3"},{"comment":"The main text states that 'Evaluation on the full E$^3$mo-Bench yields consistent findings, with detailed results... provided in Supp.' However, Supp. Table 5 reports only the five individual models on the full benchmark and does not include E$^3$mo-Score, so the claimed generalization of E$^3$mo-Score's superiority to the full benchmark is not supported by any reported result.","section":"Evaluation on Assessment Task; Supp. F.2"}],"minor_comments":[{"comment":"The statement that 'models perceive expressed emotions slightly better than evoked emotions' should be scoped to the Perception task, because the Recognition results in Tables 1 and 2 often show the opposite direction; as written, the two paragraphs appear inconsistent.","section":"Overall Analysis in Experiments"},{"comment":"The cost comparison reports 'approximately 64%' of conventional MOS cost and then says 'roughly 60% when reported at a coarse level'; the abstract and main text use only the 60% figure, so please reconcile these numbers or present a single estimate with its uncertainty.","section":"Supp. C.7"},{"comment":"The row label 'Random guess w/o recognition' is confusing because recognition scores are listed as NA; please clarify that the random baseline applies only to the Perception task.","section":"Tables 1 and 2"},{"comment":"The benchmark itself is not released and no URL or data-access statement is provided, which prevents independent verification of the QA pairs, VAD scores, and the reported model evaluations; please add an availability statement.","section":"General"},{"comment":"The F_VAD metric depends on several data-adaptive parameters (tau_dup, sigma_s, gamma, beta); please state whether the reported conclusions are robust to reasonable perturbations of these parameters.","section":"Eq. (4) and Supp. E.1"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about face-based routing lands: it is the load-bearing weakness of the paper's main empirical claim. The paper has useful components, especially the human Golden Anchor annotation and the BPA validation on one subtask, but the authors need either a dual-perspective control or a substantially more careful framing before the benchmark claim is publishable. The lack of benchmark release and the circularity in the E3mo-Score evaluation also need to be addressed. I would not reject the paper outright, but the revision required is substantial."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real contribution to affective benchmarking, but the headline finding is not yet established. The benchmark is genuinely new — first audio-visual testbed that evaluates both expressed and evoked emotion with open-vocabulary recognition and continuous VAD under one protocol. I'd send it to reviewers if the data ships.\n\nThe strongest evidence is the Golden Anchor set: 15 annotators, z-scored normalization, reliability thresholds, 383 high-consensus samples, and the BPA recovery of expressed-valence from pairwise data (SRCC 0.91/PLCC 0.90). That is a legitimate result. The BPA machinery — TrueSkill-style scheduling, Rao-Kupper ties, annotator reliability, adaptive priors — is not conceptually new, but it is carefully assembled and the ablations show each piece earns its keep. The paper also gives a clear, honest limitations section.\n\nThe soft spots are real. The biggest one is the line separating expressed from evoked. Videos are routed by YOLO face detection: face present goes to expressed, no face to evoked. The expressed pool is mostly DFEW, MELD, CAER, MAFW — actors and scripted dialogue — while the evoked pool is mostly VGGSound, LIRIS-ACCEDE, VideoEmotion, MediaEval. So the two pools differ in face presence, source genre, and likely difficulty all at once. The paper itself concedes (Supplement H.1) there are no paired annotations of both perspectives on the same video. That means the central claim — \"models perceive expressed emotions slightly better than evoked\" — is confounded. It might be true, but this design cannot demonstrate it. The authors should either add a paired subset or openly present the gap as a property of these two content distributions, not of affective perspective.\n\nSecond, the benchmark is not released. For a resource paper that is the main product; without the data, the evaluation numbers are unverifiable. Third, the E3mo-Score evaluation uses a split of the Golden Anchor set for calibration and inference. That is not fatal — the committee weights are estimated on screening pairs, and the BPA validation on expressed valence is on held-out anchors — but the paper should state explicitly that calibration and evaluation sets are disjoint. The adaptive prior initialized from model-generated bootstrap scores is a mild circularity; the prior weakens with evidence, so it is manageable.\n\nBottom line: if the data is released and the skew claim is reframed, this becomes a useful benchmark for affective MLLM evaluation. As is, I would accept it for peer review with a request for revision on those two points.","headline":"Genuinely new audio-visual benchmark for expressed and evoked emotion, but the headline expressed-vs-evoked gap is confounded by face-based content routing; referee-worthy if the data ships.","tokens_in":42342,"tokens_out":2314,"would_cite":false,"duration_ms":23917,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A unified benchmark, E3mo-Bench, measures both evoked and expressed emotion understanding in MLLMs and finds models handle the two perspectives unevenly.","keywords":["E3mo-Bench","evoked emotion","expressed emotion","Bayesian pairwise alignment","valence-arousal-dominance","multimodal large language models","open-vocabulary emotion recognition","audio-visual emotion benchmark"],"falsifier":"Re-annotate a matched subset of the same videos under both affective perspectives and compare model performance on perspective-matched pairs; if the expressed/evoked gap disappears or reverses when content and difficulty are equalized, the reported skew is an artifact of video routing rather than a genuine property of the models. Concretely, the paper's own proposed 'dual-perspective subset' (stated as future work in the Limitations section) would settle this directly.","tokens_in":41260,"feed_emoji":"🎭","tokens_out":7835,"duration_ms":71608,"temperature":0.7,"pith_summary":"E3mo-Bench is a new audio-visual benchmark that asks whether multimodal large language models (MLLMs) can handle both sides of emotion: the emotion a person in a video expresses (shown on a face or body) and the emotion a video evokes in a viewer. It contains 12,314 question-answer pairs over 2,524 videos, each video assigned one predefined affective perspective, and covers three tasks: perception, open-vocabulary recognition, and valence-arousal-dominance (VAD) assessment. To scale the continuous VAD labels cheaply, the paper introduces Bayesian Pairwise Alignment (BPA), which turns sparse, low-burden pairwise judgments into anchor-calibrated 1–9 scores, plus a training-free committee agent, E3mo-Score, that automates the same protocol. The headline empirical finding is that current MLLMs perform unevenly across the two perspectives: expressed perception generally beats evoked perception, fine-grained evoked recognition beats expressed recognition, and dominance lags across the board.","feed_headline":"Benchmark reveals AI's split between felt and shown emotion","feed_subtitle":"12,314 QA pairs over 2,524 videos expose lopsided multimodal emotional intelligence.","key_machinery":"The engine of the paper is Bayesian Pairwise Alignment (BPA), a framework for turning sparse \"which video is higher / lower / similar?\" judgments into continuous valence-arousal-dominance (VAD) scores on a 1–9 scale. Each sample carries a latent score $\\mu$ and an uncertainty $\\sigma$; Golden Anchors (383 human-consensus videos) are fixed references, and candidates are initialized from percentile-mapped model bootstraps. Comparisons are scheduled by a tripartite policy — TrueSkill-inspired exploration (50%), anchor calibration (30%), and bridge sampling (20%) — and periodically a global maximum a posteriori optimization fits a Rao–Kupper tie-aware pairwise model with per-annotator reliability weights, using an adaptive quadratic prior that weakens as evidence accumulates. Uncertainty is estimated by a diagonal empirical-Fisher approximation, temporally smoothed, and used as the stopping criterion. E3mo-Score is BPA's model-based instantiation: five MLLMs vote on each pairwise comparison, weighted by per-task reliability calibrated on held-out anchors, and the same BPA optimizer yields the final VAD predictions.","core_discovery":"On its own terms, the paper establishes that a single benchmark can test expressed and evoked emotion understanding under a shared taxonomy, and that doing so exposes a reliable imbalance in today's models. Using eight source datasets and a two-tier annotation design — a 383-sample Golden Anchor Set with strict human consensus and a 2,141-sample Broad-Coverage Set validated by a human-in-the-loop pipeline — the authors construct 12,314 QA pairs. They report that the best open and proprietary MLLMs score around 60–70% on perception tasks, cluster near random on pairwise video comparisons, and show a pronounced skew: expressed perception is easier than evoked perception (best models: ~62% vs ~60%), while open-vocabulary recognition is stronger for evoked than expressed emotion (best: 61.52% vs 53.31% on the paper's F_VAD metric). The paper also claims that BPA recovers anchor-aligned VAD scores from sparse comparisons at roughly 60% of the cost of conventional mean-opinion-score annotation, and that the training-free E3mo-Score committee outperforms every individual model it aggregates.","pith_inferences":["One testable extension is to re-annotate a matched subset of videos under both perspectives; if the expressed/evoked gap persists with content and difficulty held constant, it is a genuine model limitation, while if it shrinks the gap is partly an artifact of face-detection routing.","Because the preliminary VAD bootstrap relies on English lexical norms (13,915 lemmas), the benchmark likely inherits English-language affective biases; re-norming with multilingual or cross-cultural emotion lexicons could shift both the scores and the model rankings.","The success of a training-free 'committee that compares rather than scores' suggests a general recipe for subjective estimation: replace direct regression with pairwise preference aggregation, which may transfer to other fine-grained LLM evaluation dimensions.","The paper's own limitations list admits that no video has annotations for both perspectives, so the benchmark cannot yet say how expressed and evoked emotions interact within a single clip; a dual-perspective subset would be the natural next release."],"forward_implications":["A single 'emotional intelligence' score is misleading: E3mo-Bench implies MLLMs should be evaluated separately for expressed and evoked understanding, since the two are not learned in tandem.","Pairwise-comparison annotation with Bayesian calibration could replace expensive absolute-rating protocols for other subjective continuous labels, such as aesthetics or humor, at a reported ~60% of conventional cost.","The persistent dominance deficit points to a concrete training target: models need explicit cues of agency, control, and submissiveness, which current emotion datasets mostly ignore.","The redundancy-aware VAD set similarity metric (F_VAD) makes open-vocabulary emotion errors comparable by severity, not just by semantic category, which could become a standard grading scheme for emotion generation.","The dual-perspective design (evoked vs expressed) sets a template for future benchmarks that want to separate a model's theory of mind about others from its own affective response."],"supporting_citations":[{"why":"Defines the 9-point Self-Assessment Manikin (SAM) VAD rating scale used for all dimensional labels.","marker":"Bradley and Lang 1994"},{"why":"Supplies the 13,915-lemma VAD norms used to bootstrap preliminary scores from open-vocabulary emotion words.","marker":"Warriner, Kuperman, and Brysbaert 2013"},{"why":"TrueSkill, the Bayesian rating system that drives provisional online updates and uncertainty-based pair scheduling.","marker":"Herbrich, Minka, and Graepel 2006"},{"why":"The tie-aware paired-comparison model underlying the pairwise likelihood in BPA.","marker":"Rao and Kupper 1967"},{"why":"Crowd-BT-style annotator-reliability mixture that downweights unreliable judgments.","marker":"Chen et al. 2013"},{"why":"AffectGPT: supplies the five emotion-wheel alignment protocol (FEW) and motivates reflective multimodal annotation.","marker":"Lian et al. 2025"},{"why":"EEmo-Bench: the prior image-evoked benchmark for VAD/OV annotation that E3mo-Bench extends to audio-visual dual-perspective.","marker":"Gao et al. 2025"},{"why":"Q-Bench: the probability-based logit-pooling protocol used to score VAD for models not trained to regress.","marker":"Wu et al. 2023"}],"fun_headline_variants":["AI emotion benchmark: expressed vs evoked gap","12,314 questions reveal AI's emotion understanding split","E3mo-Bench: AI better at shown than felt emotion perception","Benchmark exposes AI's uneven emotion recognition","AI emotion test: felt vs shown is not symmetric"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire expressed-versus-evoked comparison rests on treating face presence as an operational proxy: videos with a detected face are labeled expressed and videos without one are labeled evoked, with the assumption that the two pools differ only in affective perspective rather than in content, source dataset, or difficulty.","fun_headline_variants_meta":{"raw":{"variants":["AI emotion benchmark: expressed vs evoked gap","12,314 questions reveal AI's emotion understanding split","E3mo-Bench: AI better at shown than felt emotion perception","Benchmark exposes AI's uneven emotion recognition","AI emotion test: felt vs shown is not symmetric"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000766,"raw_usage":{"total_tokens":3425,"prompt_tokens":1001,"completion_tokens":2424,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":2348}},"tokens_in":617,"tokens_out":2424,"duration_ms":19473,"temperature":1.0,"reasoning_tokens":2348,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:12:27.050443+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a matched subset of the same videos under both affective perspectives and compare model performance on perspective-matched pairs; if the expressed/evoked gap disappears or reverses when content and difficulty are equalized, the reported skew is an artifact of video routing rather than a genuine property of the models. Concretely, the paper's own proposed 'dual-perspective subset' (stated as future work in the Limitations section) would settle this directly.","supporting_citations":[],"review_version":1}