REVIEW 2 major objections 5 minor 39 references
EmoPrefer's preference scores can be matched by a logistic regression that uses only description length and generator identity — never reading the text, video, or audio — showing the benchmark rewards style over grounded content.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 15:11 UTC pith:7FX6FDDZ
load-bearing objection Careful, reproducible audit showing EmoPrefer scores can be gamed without video grounding; the position-bias worry does not overturn the main finding, but the missing random baseline on the counter-stereotypical slice is a real blemish. the 2 major comments →
Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that EmoPrefer's preference labels are predicted largely by description length and generator identity, not by video-grounded content. The authors demonstrate this with content-blind probes: a logistic regression on 18 features (lengths and one-hot generator IDs) achieves 65.8 WAF, statistically tied with finetuned 7B judges that process text, video, and audio (66.5–66.8). They identify the mechanism: generator identity is 99.5% recoverable from surface text, every candidate pair contrasts different generators, and the human labels agree with a per-generator win-rate prior on 66% of pairs. On the 552 pairs that counter the prior, trained judges follow the prior on 63–80%
What carries the argument
The load-bearing device is the content-blind logistic-regression probe — 18 features: character lengths of the two descriptions plus one-hot generator identities — which never sees the text, video, or audio. It matches finetuned judges because of a three-part mechanism: (1) writing style reveals the generator (a TF–IDF classifier recovers the source with 99.5% accuracy); (2) every pair contrasts two distinct generators, so the cue always discriminates; (3) the labels track a per-generator win-rate prior, agreeing with it on 66% of pairs. The ODIN-style diagnostic formalizes the decomposition r(d) = r_C(d) + r_L(d) + r_G(d) — content, length, generator-prior heads — with a decorrelation objec
Load-bearing premise
That the ODIN-style decomposition and decorrelation objective separate style from content without discarding genuine quality signals; if decorrelation also removes real content that co-varies with style, the 'near-chance content head' overstates how stylistic the labels are.
What would settle it
A judge that beats the content-blind probe by a statistically significant margin on length-matched counter-stereotypical pairs (e.g., >10 WAF above the probe while the probe stays low) would falsify the claim that scores are reachable without video grounding.
If this is right
- Scores on EmoPrefer-V2, and by extension the MER-Prefer track, largely reflect generator style and length, not grounded emotion understanding; the human preference signal as currently collected is confounded.
- Any model or judge evaluated on the benchmark can achieve high scores by picking up on source identity or verbosity, so leaderboard rankings should not be read as evidence of video understanding.
- Adding more data (V1) or more media (frames, audio) yields no statistically significant gain on length-matched pairs, indicating the tested configurations do not extract additional content signal.
- Chain-of-thought reasoning does not help; it just interpolates the prior and loses accuracy.
- The recommended fixes — source-balanced pairing, strict length control, counter-stereotypical slices, multi-annotator consensus — would produce a benchmark that rewards grounded content rather than style.
Where Pith is reading between the lines
- The audit protocol itself (content-blind probes + counter-stereotypical slices + deconfounding) transfers to other pairwise preference benchmarks, especially LLM-as-a-judge evaluations; any benchmark where candidates come from an identifiable set of generators will exhibit the same shortcut unless pairing is source-balanced.
- Beyond the judge side, the finding implies that preference labels used to train reward models may encode generator style rather than human emotion judgment; if so, alignment via preference optimization on such labels could reward stylistic mimicry rather than emotional accuracy.
- A natural test the paper leaves open: re-annotate EmoPrefer pairs with strictly length-matched, source-blind descriptions produced by the same generator, and check whether human agreement rates collapse; if they do, the human labels themselves are style-driven.
- The near-chance content head suggests either that the descriptions contain little video-specific emotional detail, or that the deconfounding removes genuine content that co-varies with style; the paper notes the latter caveat, so the 'style over substance' verdict is best read as a lower bound on groundedness.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper audits EmoPrefer, a pairwise emotion-description preference benchmark used in the MER2026 MER-Prefer track. The authors show that a content-blind logistic regression using only description length and generator identity reaches 65.8 WAF on EmoPrefer-V2, statistically tied with LoRA-finetuned Qwen2.5-7B text (66.5) and Omni-7B audio-visual (66.8) judges. Supporting analyses show that generator identity is 99.5% recoverable from text via a TF-IDF classifier, that every V2 pair contrasts two distinct generators, and that human labels agree with a fold-exclusive per-generator win-rate prior on 66% of pairs. On counter-stereotypical pairs (where the label contradicts the prior), trained judges side with the prior on 63–80% of pairs; on length-matched slices, media configurations yield no significant improvement; and an ODIN-style deconfounding diagnostic leaves the content head near chance. The paper recommends source-balanced pairing, strict length control, counter-stereotypical sliced reporting, and multi-annotator consensus for future benchmark iterations.
Significance. If the findings hold, they are practically important: preference evaluation in the MER-Prefer track may be gameable by style priors without any video grounding, which affects both leaderboard interpretation and reward-model training. The paper makes a methodological contribution by articulating a reusable shortcut-audit protocol: content-blind probes, counter-stereotypical slices, and permutation/counterfactual controls. Strengths include fold-exclusive priors, bootstrapped CIs, dual-order inference, and a released codebase. The main risk is an unaddressed candidate-presentation-order confound that could change the mechanistic interpretation from 'style over substance' to 'order over substance,' though the narrow claim that current scores can be reached without video would survive even if order were the true cue. The ODIN-based 'little grounded signal' claim also depends on the decorrelation objective's assumptions, which the paper acknowledges only parenthetically.
major comments (2)
- [Section 3 and Appendix A] The binary V2 set has a 61.5% first-position label preference (1,000/1,625 pairs), yet the paper never analyzes candidate presentation order. The content-blind probe uses one-hot generator identity; if generator identity correlates with which description is shown first, the 'generator prior' and the 'style prior' defining counter-stereotypical pairs may be proxies for an order prior. The paper should report the generator-by-position balance, add a position-only probe, and/or condition the key analyses on presentation position. Without this, the mechanistic claim that the benchmark ranks generator style rather than video-grounded content is unsupported, although the narrower claim that scores can be reached without video would survive.
- [Section 4.2 and Fig. 2b] The text states that judges on counter-stereotypical pairs score 20–42 WAF, 'far below the ≈50 of random guessing.' WAF is a support-weighted F1 and its random reference depends on the slice's label distribution; given the 1,000/625 position imbalance, 50 is not an established baseline. The figure caption mentions a 'stratified-random reference' (dotted) but the text never reports it. Please report the stratified-random WAF for each slice and replace or qualify the 'far below random' claim.
minor comments (5)
- [Section 4.2] 'The wordiest system wins only 63%' is vague; specify which generator this refers to or provide the per-generator win-rate/length table.
- [Fig. 2b caption] The dotted 'stratified-random reference' in the figure is not referenced in the main text; please point readers to it where the counter-stereotypical WAF values are discussed.
- [Appendix A] Clarify how ties in the generator prior (π(g) equal for both candidates) are handled in the 66% agreement calculation and in the counter-stereotypical slice; the text says 'ties abstain' but does not state the resulting counts.
- [Table 1] The 'Generator win-rate rule' row's 0.0 on the counter-stereotypical slice is by construction; add a footnote to prevent reader confusion.
- [Abstract and Section 7] The caveat in Section 4.3 that decorrelation might remove legitimate quality signals should be reflected in the abstract and conclusion, which currently state that 'little grounded signal survives' without that qualification.
Circularity Check
No significant circularity: headline probe results are out-of-fold, and the ODIN-style diagnostic is an explicitly caveated external method.
full rationale
All headline numbers are computed on held-out folds: the content-blind logistic regression is fit on four folds and evaluated on the fifth (Sec. 3; Appendix A: 'Each is l2-regularized logistic regression ... fit on four folds'), and the generator win-rate prior is estimated out-of-fold (Eq. 2; Appendix A), so the 66% agreement with labels is not an in-sample fit. The TF-IDF source classifier is also fitted only on training folds (Appendix B). The counter-stereotypical slice is defined by the same fold-exclusive prior, and the paper explicitly states 'a pure style prior scores 0 here by definition' (Table 1 note); this is a diagnostic slice, not a predicted quantity. The ODIN diagnostic (Eq. 1) is adapted from an external paper [3] and is presented as a diagnostic with explicit caveats: 'decorrelation might remove legitimate quality signals co-varying with style' (Sec. 4.2) and r_C 'cannot establish video grounding' (Appendix B). The central claim that current scores can be reached without video grounding rests on the content-blind probe parity (65.8 vs 66.8 WAF), which is an out-of-fold comparison, not a fitted value renamed as a prediction. No load-bearing self-citation chain appears; references to ODIN, EmoPrefer, and MER2026 are external prior work. The unaddressed 61.5% first-position label imbalance is a potential confound for the mechanistic 'style over substance' interpretation, but it is a validity threat, not a circular derivation.
Axiom & Free-Parameter Ledger
free parameters (3)
- Per-generator win-rate prior π(g) =
7 values spanning 32–79% (EmoPrefer-V2)
- Logistic regression probe weights (length + generator) =
18-coefficient vector, re-fit per fold
- ODIN/LoRA judge weights and heads =
not enumerated; trained checkpoints
axioms (5)
- domain assumption Human preference labels in EmoPrefer are the reference outcome for measuring shortcut behavior.
- standard math The Bradley–Terry model links reward differences to preference probabilities.
- ad hoc to paper De-correlating r_C from length ℓ and generator prior π is sufficient to remove the style shortcut, leaving a content head.
- domain assumption The tested 7B LoRA judges are representative enough to detect any video-grounded signal if it existed.
- domain assumption Weighted F1 (WAF) is an appropriate task metric.
read the original abstract
Preference over model-generated emotion descriptions is emerging as a standard evaluation metric for multimodal emotion understanding, exemplified by the MER2026 MER-Prefer track on EmoPrefer. Such benchmarks assume that predicting the preferred description requires grounded cross-modal understanding of the video. We conduct a systematic shortcut audit of EmoPrefer using content-blind probes. A simple logistic regression using only description length and generator identity, without processing the text, video, or audio, performs comparably to LoRA-finetuned 7B text and audio-visual judges (65.8 versus 66.8 WAF on EmoPrefer-V2). Generator identity is recoverable from description text with 99.5 percent accuracy, every candidate pair contrasts two distinct generators, and the human preference labels agree with a fold-exclusive per-generator win-rate prior on 66 percent of the evaluated pairs. When the human label conflicts with this prior, trained judges still follow the style prior on 63 to 80 percent of the pairs. On a length-matched subset that neutralizes verbosity bias, the tested media configurations yield no statistically significant improvement, while an ODIN-inspired diagnostic that decouples the style shortcut leaves its content head near chance. These results do not imply that human preferences are inherently stylistic or that the descriptions contain no emotional information. Instead, they show that the current scores can be reached without verifying either description against the video. We recommend source-balanced pairing, strict length control, counter-stereotypical sliced reporting, and multi-annotator consensus for future cross-generator evaluations. Code is available at https://github.com/jiabingyang01/EmoPrefer-Audit.
Figures
Reference graph
Works this paper leans on
-
[1]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...
Pith/arXiv arXiv 2025
-
[2]
Ralph Allan Bradley and Milton E Terry. 1952. Rank analysis of incomplete block designs: I. the method of paired comparisons.Biometrika39, 3/4 (1952), 324–345
1952
-
[3]
Lichang Chen, Chen Zhu, Jiuhai Chen, Davit Soselia, Tianyi Zhou, Tom Gold- stein, Heng Huang, Mohammad Shoeybi, and Bryan Catanzaro. 2024. ODIN: Disentangled Reward Mitigates Hacking in RLHF. InInternational Conference on Machine Learning. PMLR, 7935–7952
2024
-
[4]
Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto. 2024. Length-controlled alpacaeval: A simple way to debias automatic evaluators.arXiv preprint arXiv:2404.04475(2024)
Pith/arXiv arXiv 2024
-
[5]
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al
-
[6]
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020. Shortcut learning in deep neural networks.Nature Machine Intelligence2, 11 (2020), 665–673
2020
-
[7]
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R Bowman, and Noah A Smith. 2018. Annotation artifacts in natural language inference data. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers). 107–112
2018
-
[8]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, et al. 2022. Lora: Low-rank adaptation of large language models.Iclr1, 2 (2022), 3
2022
-
[9]
Seungone Kim, Jay Shin, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Ryan Shin, Sungdong Kim, James Thorne, Minjoon Seo, et al. 2024. Prometheus: Inducing fine-grained evaluation capability in language models. InInternational Conference on Learning Representations, Vol. 2024. 29927–29962
2024
-
[10]
Zheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun, Licai Sun, Yong Ren, Zebang Cheng, Bin Liu, Rui Liu, Xiaojiang Peng, et al. 2025. AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models. InInternational Conference on Machine Learning. PMLR, 36993– 37014
2025
-
[11]
Zheng Lian, Rui Liu, Kele Xu, Bin Liu, Xuefei Liu, Yazhou Zhang, Xin Liu, Yong Li, Zebang Cheng, Haolin Zuo, et al. 2025. Mer 2025: When affective comput- ing meets large language models. InProceedings of the 33rd ACM International Conference on Multimedia. 13837–13842
2025
-
[12]
Zheng Lian, Xiaojiang Peng, Kele Xu, Ziyu Jia, Xinyi Che, Zebang Cheng, Fei Ma, Laizhong Cui, Yazhou Zhang, Xin Liu, et al. 2026. MER 2026: From Discriminative Emotion Recognition to Generative Emotion Understanding.arXiv preprint arXiv:2604.19417(2026)
Pith/arXiv arXiv 2026
-
[13]
Zheng Lian, Haiyang Sun, Licai Sun, Kang Chen, Mingyu Xu, Kexin Wang, Ke Xu, Yu He, Ying Li, Jinming Zhao, et al. 2023. Mer 2023: Multi-label learning, modality robustness, and semi-supervised learning. InProceedings of the 31st ACM international conference on multimedia. 9610–9614
2023
-
[14]
Zheng Lian, Haiyang Sun, Licai Sun, Zhuofan Wen, Siyuan Zhang, Shun Chen, Hao Gu, Jinming Zhao, Ziyang Ma, Xie Chen, et al . 2024. Mer 2024: Semi- supervised learning, noise robustness, and open-vocabulary multimodal emotion recognition. InProceedings of the 2nd International Workshop on Multimodal and Responsible Affective Computing. 41–48
2024
-
[15]
Zheng Lian, Licai Sun, Lan Chen, Haoyu Chen, Zebang Cheng, Fan Zhang, Ziyu Jia, Ziyang Ma, Fei Ma, Xiaojiang Peng, et al. 2025. EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?arXiv preprint arXiv:2507.04278(2025)
arXiv 2025
-
[16]
Tianqi Liu, Wei Xiong, Jie Ren, Lichang Chen, Junru Wu, Rishabh Joshi, Yang Gao, Jiaming Shen, Zhen Qin, Tianhe Yu, et al . 2025. Rrm: Robust reward model training mitigates reward hacking. InInternational Conference on Learning Representations, Vol. 2025. 62682–62700
2025
-
[17]
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: NLG evaluation using gpt-4 with better human alignment. InProceedings of the 2023 conference on empirical methods in natural language processing. 2511–2522
2023
-
[18]
Yantao Liu, Zijun Yao, Rui Min, Yixin Cao, Lei Hou, and Juanzi Li. 2025. Rm- bench: Benchmarking reward models of language models with subtlety and style. InInternational Conference on Learning Representations, Vol. 2025. 44323–44355
2025
-
[19]
R Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. InProceedings of the 57th annual meeting of the association for computational linguistics. 3428–3448
2019
-
[20]
Quinn McNemar. 1947. Note on the sampling error of the difference between correlated proportions or percentages.Psychometrika12, 2 (1947), 153–157
1947
-
[21]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback.Advances in neural information processing systems35 (2022), 27730–27744
2022
-
[22]
Junsoo Park, Seungyeon Jwa, Ren Meiying, Daeyoung Kim, and Sanghyuk Choi
-
[23]
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Ben- jamin Van Durme. 2018. Hypothesis only baselines in natural language inference. InProceedings of the seventh joint conference on lexical and computational seman- tics. 180–191
2018
-
[24]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li,...
Pith/arXiv arXiv 2025
-
[25]
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems36 (2023), 53728–53741
2023
-
[26]
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with CheckList. InProceed- ings of the 58th annual meeting of the association for computational linguistics. 4902–4912
2020
-
[27]
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. 2020. Distributionally Robust Neural Networks. InInternational Conference on Learning Representations
2020
-
[28]
Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett. 2023. A long way to go: Investigating length correlations in rlhf.arXiv preprint arXiv:2310.03716 (2023)
Pith/arXiv arXiv 2023
-
[29]
Pragya Srivastava, Harman Singh, Rahul Madhavan, Gandharv Patil, Sravanti Addepalli, Arun Suggala, Rengarajan Aravamudhan, Soumya Sharma, Anirban Laha, Aravindan Raghuveer, et al. 2025. Robust Reward Modeling via Causal Rubrics. In2nd Workshop on Models of Human Feedback for AI Alignment
2025
-
[30]
Chaoqi Wang, Zhuokai Zhao, Yibo Jiang, Zhaorun Chen, Chen Zhu, Yuxin Chen, Jiayi Liu, Lizhu Zhang, Xiangjun Fan, Hao Ma, et al . 2025. Beyond reward hacking: Causal rewards for large language model alignment.arXiv preprint arXiv:2501.09620(2025)
Pith/arXiv arXiv 2025
-
[31]
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, et al. 2024. Large language models are not fair evaluators. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 9440–9450
2024
-
[32]
Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, Bin Zhang, Xiong Wang, Yunfei Chu, and Junyang Lin. 2025. Qwen2.5-Omni Technical Report. arXiv:2503.20215 [cs.CL] https://arxiv.org/abs/2503.20215
Pith/arXiv arXiv 2025
-
[33]
Wenqian Ye, Guangtao Zheng, and Aidong Zhang. 2026. Rectifying shortcut behaviors in preference-based reward learning.Advances in Neural Information Processing Systems38 (2026), 64712–64740
2026
-
[34]
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022. Star: Boot- strapping reasoning with reasoning.Advances in Neural Information Processing Systems35 (2022), 15476–15488
2022
-
[35]
Lunjun Zhang, Arian Hosseini, Hritik Bansal, Seyed Mehran Kazemi, Aviral Kumar, and Rishabh Agarwal. 2025. Generative verifiers: Reward modeling as next-token prediction. InInternational Conference on Learning Representations, Vol. 2025. 12476–12505
2025
-
[36]
Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. 2024. LLaVA-NeXT: A Strong Zero-shot Video Understanding Model. https://llava-vl.github.io/blog/2024-04-30-llava-next- video/
2024
-
[37]
same” labels, so all judges and probes share held-out members. V1 contains 574 unanimous three-annotator pairs; V2 contains 2,096 individually labeled pairs. Removing 471 V2 “same
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems36 (2023), 46595–46623. 6 Appendix This appendix provides the evaluation and implementation details omitted fr...
2023
-
[2020]
InFindings of the Association for Computational Linguistics: EMNLP 2020
Evaluating models’ local decision boundaries via contrast sets. InFindings of the Association for Computational Linguistics: EMNLP 2020. 1307–1323
2020
-
[2024]
InFindings of the Association for Computational Linguistics: EMNLP 2024
Offsetbias: Leveraging debiased data for tuning evaluators. InFindings of the Association for Computational Linguistics: EMNLP 2024. 1043–1067
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.