Pith. sign in

REVIEW 2 major objections 5 minor 39 references

EmoPrefer's preference scores can be matched by a logistic regression that uses only description length and generator identity — never reading the text, video, or audio — showing the benchmark rewards style over grounded content.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 15:11 UTC pith:7FX6FDDZ

load-bearing objection Careful, reproducible audit showing EmoPrefer scores can be gamed without video grounding; the position-bias worry does not overturn the main finding, but the missing random baseline on the counter-stereotypical slice is a real blemish. the 2 major comments →

arxiv 2607.18508 v1 pith:7FX6FDDZ submitted 2026-07-20 cs.CV

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation

classification cs.CV
keywords emotion-description preferencebenchmark auditshortcut learningcontent-blind probesgenerator identityODIN deconfoundingLLM-as-a-judgereward models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper audits EmoPrefer, a benchmark that scores how well models predict which of two machine-generated emotion descriptions of a video a human prefers. The authors show that a content-blind logistic regression using only description length and generator identity — no text, video, or audio — scores on par with finetuned 7B text and audio-visual judges (65.8 vs 66.8 WAF). They trace this to generator identity leaking through writing style and to labels following a per-generator win-rate prior on 66% of pairs; where the label contradicts the prior, judges still side with the prior 63–80% of the time. After deconfounding content from style, the content head lands near chance. The conclusion: the current benchmark scores can be reached without verifying either description against the video, so the benchmark conflates generator style with grounded accuracy.

Core claim

The central claim is that EmoPrefer's preference labels are predicted largely by description length and generator identity, not by video-grounded content. The authors demonstrate this with content-blind probes: a logistic regression on 18 features (lengths and one-hot generator IDs) achieves 65.8 WAF, statistically tied with finetuned 7B judges that process text, video, and audio (66.5–66.8). They identify the mechanism: generator identity is 99.5% recoverable from surface text, every candidate pair contrasts different generators, and the human labels agree with a per-generator win-rate prior on 66% of pairs. On the 552 pairs that counter the prior, trained judges follow the prior on 63–80%

What carries the argument

The load-bearing device is the content-blind logistic-regression probe — 18 features: character lengths of the two descriptions plus one-hot generator identities — which never sees the text, video, or audio. It matches finetuned judges because of a three-part mechanism: (1) writing style reveals the generator (a TF–IDF classifier recovers the source with 99.5% accuracy); (2) every pair contrasts two distinct generators, so the cue always discriminates; (3) the labels track a per-generator win-rate prior, agreeing with it on 66% of pairs. The ODIN-style diagnostic formalizes the decomposition r(d) = r_C(d) + r_L(d) + r_G(d) — content, length, generator-prior heads — with a decorrelation objec

Load-bearing premise

That the ODIN-style decomposition and decorrelation objective separate style from content without discarding genuine quality signals; if decorrelation also removes real content that co-varies with style, the 'near-chance content head' overstates how stylistic the labels are.

What would settle it

A judge that beats the content-blind probe by a statistically significant margin on length-matched counter-stereotypical pairs (e.g., >10 WAF above the probe while the probe stays low) would falsify the claim that scores are reachable without video grounding.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Scores on EmoPrefer-V2, and by extension the MER-Prefer track, largely reflect generator style and length, not grounded emotion understanding; the human preference signal as currently collected is confounded.
  • Any model or judge evaluated on the benchmark can achieve high scores by picking up on source identity or verbosity, so leaderboard rankings should not be read as evidence of video understanding.
  • Adding more data (V1) or more media (frames, audio) yields no statistically significant gain on length-matched pairs, indicating the tested configurations do not extract additional content signal.
  • Chain-of-thought reasoning does not help; it just interpolates the prior and loses accuracy.
  • The recommended fixes — source-balanced pairing, strict length control, counter-stereotypical slices, multi-annotator consensus — would produce a benchmark that rewards grounded content rather than style.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The audit protocol itself (content-blind probes + counter-stereotypical slices + deconfounding) transfers to other pairwise preference benchmarks, especially LLM-as-a-judge evaluations; any benchmark where candidates come from an identifiable set of generators will exhibit the same shortcut unless pairing is source-balanced.
  • Beyond the judge side, the finding implies that preference labels used to train reward models may encode generator style rather than human emotion judgment; if so, alignment via preference optimization on such labels could reward stylistic mimicry rather than emotional accuracy.
  • A natural test the paper leaves open: re-annotate EmoPrefer pairs with strictly length-matched, source-blind descriptions produced by the same generator, and check whether human agreement rates collapse; if they do, the human labels themselves are style-driven.
  • The near-chance content head suggests either that the descriptions contain little video-specific emotional detail, or that the deconfounding removes genuine content that co-varies with style; the paper notes the latter caveat, so the 'style over substance' verdict is best read as a lower bound on groundedness.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper audits EmoPrefer, a pairwise emotion-description preference benchmark used in the MER2026 MER-Prefer track. The authors show that a content-blind logistic regression using only description length and generator identity reaches 65.8 WAF on EmoPrefer-V2, statistically tied with LoRA-finetuned Qwen2.5-7B text (66.5) and Omni-7B audio-visual (66.8) judges. Supporting analyses show that generator identity is 99.5% recoverable from text via a TF-IDF classifier, that every V2 pair contrasts two distinct generators, and that human labels agree with a fold-exclusive per-generator win-rate prior on 66% of pairs. On counter-stereotypical pairs (where the label contradicts the prior), trained judges side with the prior on 63–80% of pairs; on length-matched slices, media configurations yield no significant improvement; and an ODIN-style deconfounding diagnostic leaves the content head near chance. The paper recommends source-balanced pairing, strict length control, counter-stereotypical sliced reporting, and multi-annotator consensus for future benchmark iterations.

Significance. If the findings hold, they are practically important: preference evaluation in the MER-Prefer track may be gameable by style priors without any video grounding, which affects both leaderboard interpretation and reward-model training. The paper makes a methodological contribution by articulating a reusable shortcut-audit protocol: content-blind probes, counter-stereotypical slices, and permutation/counterfactual controls. Strengths include fold-exclusive priors, bootstrapped CIs, dual-order inference, and a released codebase. The main risk is an unaddressed candidate-presentation-order confound that could change the mechanistic interpretation from 'style over substance' to 'order over substance,' though the narrow claim that current scores can be reached without video would survive even if order were the true cue. The ODIN-based 'little grounded signal' claim also depends on the decorrelation objective's assumptions, which the paper acknowledges only parenthetically.

major comments (2)
  1. [Section 3 and Appendix A] The binary V2 set has a 61.5% first-position label preference (1,000/1,625 pairs), yet the paper never analyzes candidate presentation order. The content-blind probe uses one-hot generator identity; if generator identity correlates with which description is shown first, the 'generator prior' and the 'style prior' defining counter-stereotypical pairs may be proxies for an order prior. The paper should report the generator-by-position balance, add a position-only probe, and/or condition the key analyses on presentation position. Without this, the mechanistic claim that the benchmark ranks generator style rather than video-grounded content is unsupported, although the narrower claim that scores can be reached without video would survive.
  2. [Section 4.2 and Fig. 2b] The text states that judges on counter-stereotypical pairs score 20–42 WAF, 'far below the ≈50 of random guessing.' WAF is a support-weighted F1 and its random reference depends on the slice's label distribution; given the 1,000/625 position imbalance, 50 is not an established baseline. The figure caption mentions a 'stratified-random reference' (dotted) but the text never reports it. Please report the stratified-random WAF for each slice and replace or qualify the 'far below random' claim.
minor comments (5)
  1. [Section 4.2] 'The wordiest system wins only 63%' is vague; specify which generator this refers to or provide the per-generator win-rate/length table.
  2. [Fig. 2b caption] The dotted 'stratified-random reference' in the figure is not referenced in the main text; please point readers to it where the counter-stereotypical WAF values are discussed.
  3. [Appendix A] Clarify how ties in the generator prior (π(g) equal for both candidates) are handled in the 66% agreement calculation and in the counter-stereotypical slice; the text says 'ties abstain' but does not state the resulting counts.
  4. [Table 1] The 'Generator win-rate rule' row's 0.0 on the counter-stereotypical slice is by construction; add a footnote to prevent reader confusion.
  5. [Abstract and Section 7] The caveat in Section 4.3 that decorrelation might remove legitimate quality signals should be reflected in the abstract and conclusion, which currently state that 'little grounded signal survives' without that qualification.

Circularity Check

0 steps flagged

No significant circularity: headline probe results are out-of-fold, and the ODIN-style diagnostic is an explicitly caveated external method.

full rationale

All headline numbers are computed on held-out folds: the content-blind logistic regression is fit on four folds and evaluated on the fifth (Sec. 3; Appendix A: 'Each is l2-regularized logistic regression ... fit on four folds'), and the generator win-rate prior is estimated out-of-fold (Eq. 2; Appendix A), so the 66% agreement with labels is not an in-sample fit. The TF-IDF source classifier is also fitted only on training folds (Appendix B). The counter-stereotypical slice is defined by the same fold-exclusive prior, and the paper explicitly states 'a pure style prior scores 0 here by definition' (Table 1 note); this is a diagnostic slice, not a predicted quantity. The ODIN diagnostic (Eq. 1) is adapted from an external paper [3] and is presented as a diagnostic with explicit caveats: 'decorrelation might remove legitimate quality signals co-varying with style' (Sec. 4.2) and r_C 'cannot establish video grounding' (Appendix B). The central claim that current scores can be reached without video grounding rests on the content-blind probe parity (65.8 vs 66.8 WAF), which is an out-of-fold comparison, not a fitted value renamed as a prediction. No load-bearing self-citation chain appears; references to ODIN, EmoPrefer, and MER2026 are external prior work. The unaddressed 61.5% first-position label imbalance is a potential confound for the mechanistic 'style over substance' interpretation, but it is a validity threat, not a circular derivation.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

The central claim rests on standard empirical baselines and a public benchmark; the only substantive modeling assumptions come from the ODIN diagnostic and the choice of WAF. No new physical or ontological entities are posited; the content head r_C is a model component, not a new entity.

free parameters (3)
  • Per-generator win-rate prior π(g) = 7 values spanning 32–79% (EmoPrefer-V2)
    Fit on training folds from binary preference labels; used to measure prior-label agreement (66%) and to define counter-stereotypical slices. This is a deliberate baseline, but it is a number fitted to data.
  • Logistic regression probe weights (length + generator) = 18-coefficient vector, re-fit per fold
    Content-blind probes are fitted to training-fold labels to demonstrate WAF parity; individual coefficients are not reported.
  • ODIN/LoRA judge weights and heads = not enumerated; trained checkpoints
    All finetuned judge and diagnostic head weights are fitted; they are standard model parameters, not hand-set constants.
axioms (5)
  • domain assumption Human preference labels in EmoPrefer are the reference outcome for measuring shortcut behavior.
    The audit measures how well predictors recover the human label without video; it does not independently validate the labels as emotion ground truth (Section 3, WAF metric).
  • standard math The Bradley–Terry model links reward differences to preference probabilities.
    Used in the ODIN diagnostic (Eq. 1); standard choice from prior literature.
  • ad hoc to paper De-correlating r_C from length ℓ and generator prior π is sufficient to remove the style shortcut, leaving a content head.
    The decomposition r = r_C + r_L + r_G and objectives L_dec/L_orth (Eq. 1) are adapted from ODIN; the identification of r_C as 'content' is an assumption the authors partially flag.
  • domain assumption The tested 7B LoRA judges are representative enough to detect any video-grounded signal if it existed.
    Used in Section 5 to conclude media inputs yield no detectable gain and r_C is near chance; limited capacity, frame counts, and n=310 low power could obscure a true content signal.
  • domain assumption Weighted F1 (WAF) is an appropriate task metric.
    Adopted from the challenge; WAF is not accuracy, so the paper's own chance-level caveat applies.

pith-pipeline@v1.3.0-alltime-deepseek · 12920 in / 18494 out tokens · 207158 ms · 2026-08-01T15:11:37.894307+00:00 · methodology

0 comments
read the original abstract

Preference over model-generated emotion descriptions is emerging as a standard evaluation metric for multimodal emotion understanding, exemplified by the MER2026 MER-Prefer track on EmoPrefer. Such benchmarks assume that predicting the preferred description requires grounded cross-modal understanding of the video. We conduct a systematic shortcut audit of EmoPrefer using content-blind probes. A simple logistic regression using only description length and generator identity, without processing the text, video, or audio, performs comparably to LoRA-finetuned 7B text and audio-visual judges (65.8 versus 66.8 WAF on EmoPrefer-V2). Generator identity is recoverable from description text with 99.5 percent accuracy, every candidate pair contrasts two distinct generators, and the human preference labels agree with a fold-exclusive per-generator win-rate prior on 66 percent of the evaluated pairs. When the human label conflicts with this prior, trained judges still follow the style prior on 63 to 80 percent of the pairs. On a length-matched subset that neutralizes verbosity bias, the tested media configurations yield no statistically significant improvement, while an ODIN-inspired diagnostic that decouples the style shortcut leaves its content head near chance. These results do not imply that human preferences are inherently stylistic or that the descriptions contain no emotional information. Instead, they show that the current scores can be reached without verifying either description against the video. We recommend source-balanced pairing, strict length control, counter-stereotypical sliced reporting, and multi-annotator consensus for future cross-generator evaluations. Code is available at https://github.com/jiabingyang01/EmoPrefer-Audit.

Figures

Figures reproduced from arXiv: 2607.18508 by Jiabing Yang, Liang Wang, Peiyan Li, Qisen Ma, Tao Yu, Yan Huang, Yingda Li, Yixiang Chen, Yuan Xu.

Figure 1
Figure 1. Figure 1: Overview of our shortcut audit of EmoPrefer. Stage 1 (task setup): a judge receives the video, its audio, and two candidate [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: (a) Mean absolute head–confound correlations across folds; error bars show one standard deviation. Decorrelation [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 7 linked inside Pith

  1. [1]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...

  2. [2]

    Ralph Allan Bradley and Milton E Terry. 1952. Rank analysis of incomplete block designs: I. the method of paired comparisons.Biometrika39, 3/4 (1952), 324–345

  3. [3]

    Lichang Chen, Chen Zhu, Jiuhai Chen, Davit Soselia, Tianyi Zhou, Tom Gold- stein, Heng Huang, Mohammad Shoeybi, and Bryan Catanzaro. 2024. ODIN: Disentangled Reward Mitigates Hacking in RLHF. InInternational Conference on Machine Learning. PMLR, 7935–7952

  4. [4]

    Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto. 2024. Length-controlled alpacaeval: A simple way to debias automatic evaluators.arXiv preprint arXiv:2404.04475(2024)

  5. [5]

    Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al

  6. [6]

    Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020. Shortcut learning in deep neural networks.Nature Machine Intelligence2, 11 (2020), 665–673

  7. [7]

    Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R Bowman, and Noah A Smith. 2018. Annotation artifacts in natural language inference data. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers). 107–112

  8. [8]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, et al. 2022. Lora: Low-rank adaptation of large language models.Iclr1, 2 (2022), 3

  9. [9]

    Seungone Kim, Jay Shin, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Ryan Shin, Sungdong Kim, James Thorne, Minjoon Seo, et al. 2024. Prometheus: Inducing fine-grained evaluation capability in language models. InInternational Conference on Learning Representations, Vol. 2024. 29927–29962

  10. [10]

    Zheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun, Licai Sun, Yong Ren, Zebang Cheng, Bin Liu, Rui Liu, Xiaojiang Peng, et al. 2025. AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models. InInternational Conference on Machine Learning. PMLR, 36993– 37014

  11. [11]

    Zheng Lian, Rui Liu, Kele Xu, Bin Liu, Xuefei Liu, Yazhou Zhang, Xin Liu, Yong Li, Zebang Cheng, Haolin Zuo, et al. 2025. Mer 2025: When affective comput- ing meets large language models. InProceedings of the 33rd ACM International Conference on Multimedia. 13837–13842

  12. [12]

    Zheng Lian, Xiaojiang Peng, Kele Xu, Ziyu Jia, Xinyi Che, Zebang Cheng, Fei Ma, Laizhong Cui, Yazhou Zhang, Xin Liu, et al. 2026. MER 2026: From Discriminative Emotion Recognition to Generative Emotion Understanding.arXiv preprint arXiv:2604.19417(2026)

  13. [13]

    Zheng Lian, Haiyang Sun, Licai Sun, Kang Chen, Mingyu Xu, Kexin Wang, Ke Xu, Yu He, Ying Li, Jinming Zhao, et al. 2023. Mer 2023: Multi-label learning, modality robustness, and semi-supervised learning. InProceedings of the 31st ACM international conference on multimedia. 9610–9614

  14. [14]

    Zheng Lian, Haiyang Sun, Licai Sun, Zhuofan Wen, Siyuan Zhang, Shun Chen, Hao Gu, Jinming Zhao, Ziyang Ma, Xie Chen, et al . 2024. Mer 2024: Semi- supervised learning, noise robustness, and open-vocabulary multimodal emotion recognition. InProceedings of the 2nd International Workshop on Multimodal and Responsible Affective Computing. 41–48

  15. [15]

    Zheng Lian, Licai Sun, Lan Chen, Haoyu Chen, Zebang Cheng, Fan Zhang, Ziyu Jia, Ziyang Ma, Fei Ma, Xiaojiang Peng, et al. 2025. EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?arXiv preprint arXiv:2507.04278(2025)

  16. [16]

    Tianqi Liu, Wei Xiong, Jie Ren, Lichang Chen, Junru Wu, Rishabh Joshi, Yang Gao, Jiaming Shen, Zhen Qin, Tianhe Yu, et al . 2025. Rrm: Robust reward model training mitigates reward hacking. InInternational Conference on Learning Representations, Vol. 2025. 62682–62700

  17. [17]

    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: NLG evaluation using gpt-4 with better human alignment. InProceedings of the 2023 conference on empirical methods in natural language processing. 2511–2522

  18. [18]

    Yantao Liu, Zijun Yao, Rui Min, Yixin Cao, Lei Hou, and Juanzi Li. 2025. Rm- bench: Benchmarking reward models of language models with subtlety and style. InInternational Conference on Learning Representations, Vol. 2025. 44323–44355

  19. [19]

    R Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. InProceedings of the 57th annual meeting of the association for computational linguistics. 3428–3448

  20. [20]

    Quinn McNemar. 1947. Note on the sampling error of the difference between correlated proportions or percentages.Psychometrika12, 2 (1947), 153–157

  21. [21]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback.Advances in neural information processing systems35 (2022), 27730–27744

  22. [22]

    Junsoo Park, Seungyeon Jwa, Ren Meiying, Daeyoung Kim, and Sanghyuk Choi

  23. [23]

    Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Ben- jamin Van Durme. 2018. Hypothesis only baselines in natural language inference. InProceedings of the seventh joint conference on lexical and computational seman- tics. 180–191

  24. [24]

    Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li,...

  25. [25]

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems36 (2023), 53728–53741

  26. [26]

    Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with CheckList. InProceed- ings of the 58th annual meeting of the association for computational linguistics. 4902–4912

  27. [27]

    Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. 2020. Distributionally Robust Neural Networks. InInternational Conference on Learning Representations

  28. [28]

    Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett. 2023. A long way to go: Investigating length correlations in rlhf.arXiv preprint arXiv:2310.03716 (2023)

  29. [29]

    Pragya Srivastava, Harman Singh, Rahul Madhavan, Gandharv Patil, Sravanti Addepalli, Arun Suggala, Rengarajan Aravamudhan, Soumya Sharma, Anirban Laha, Aravindan Raghuveer, et al. 2025. Robust Reward Modeling via Causal Rubrics. In2nd Workshop on Models of Human Feedback for AI Alignment

  30. [30]

    Chaoqi Wang, Zhuokai Zhao, Yibo Jiang, Zhaorun Chen, Chen Zhu, Yuxin Chen, Jiayi Liu, Lizhu Zhang, Xiangjun Fan, Hao Ma, et al . 2025. Beyond reward hacking: Causal rewards for large language model alignment.arXiv preprint arXiv:2501.09620(2025)

  31. [31]

    Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, et al. 2024. Large language models are not fair evaluators. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 9440–9450

  32. [32]

    Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, Bin Zhang, Xiong Wang, Yunfei Chu, and Junyang Lin. 2025. Qwen2.5-Omni Technical Report. arXiv:2503.20215 [cs.CL] https://arxiv.org/abs/2503.20215

  33. [33]

    Wenqian Ye, Guangtao Zheng, and Aidong Zhang. 2026. Rectifying shortcut behaviors in preference-based reward learning.Advances in Neural Information Processing Systems38 (2026), 64712–64740

  34. [34]

    Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022. Star: Boot- strapping reasoning with reasoning.Advances in Neural Information Processing Systems35 (2022), 15476–15488

  35. [35]

    Lunjun Zhang, Arian Hosseini, Hritik Bansal, Seyed Mehran Kazemi, Aviral Kumar, and Rishabh Agarwal. 2025. Generative verifiers: Reward modeling as next-token prediction. InInternational Conference on Learning Representations, Vol. 2025. 12476–12505

  36. [36]

    Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. 2024. LLaVA-NeXT: A Strong Zero-shot Video Understanding Model. https://llava-vl.github.io/blog/2024-04-30-llava-next- video/

  37. [37]

    same” labels, so all judges and probes share held-out members. V1 contains 574 unanimous three-annotator pairs; V2 contains 2,096 individually labeled pairs. Removing 471 V2 “same

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems36 (2023), 46595–46623. 6 Appendix This appendix provides the evaluation and implementation details omitted fr...

  38. [2020]

    InFindings of the Association for Computational Linguistics: EMNLP 2020

    Evaluating models’ local decision boundaries via contrast sets. InFindings of the Association for Computational Linguistics: EMNLP 2020. 1307–1323

  39. [2024]

    InFindings of the Association for Computational Linguistics: EMNLP 2024

    Offsetbias: Leveraging debiased data for tuning evaluators. InFindings of the Association for Computational Linguistics: EMNLP 2024. 1043–1067