Pith. sign in

REVIEW 4 cited by

Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.15268 v3 pith:OH4WFENF submitted 2024-09-23 cs.LG cs.AI

classification cs.LGcs.AI
keywords alignmentconcretepreferencesstylefollowingknowledgellm-judgellm-judges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The release of ChatGPT in November 2022 sparked an explosion of interest in post-training and an avalanche of new preference optimization (PO) methods. These methods claim superior alignment by virtue of better correspondence with human pairwise preferences, often measured by LLM-judges. In this work, we attempt to answer the following question -- do LLM-judge preferences translate to progress on other, more concrete metrics for alignment, and if not, why not? We define a concrete metric for alignment, and introduce SOS-Bench (Substance Outweighs Style Benchmark), which is to the best of our knowledge the largest standardized, reproducible LLM meta-benchmark to date. We find that (1) LLM-judge preferences do not correlate with concrete measures of safety, world knowledge, and instruction following; (2) LLM-judges have powerful implicit biases, prioritizing style over factuality and safety; and (3) the supervised fine-tuning (SFT) stage of post-training, and not the PO stage, has the greatest impact on alignment, with data scaling and prompt diversity as the driving factors. Our codebase and complete results can be found at https://github.com/penfever/sos-bench.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    CN-PR learns reward functions from LLM-derived preferences over clinical trajectories to improve RL policies for sequential treatment decisions, showing correlation with quality scores and better recovery outcomes.

  2. LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Small reward models trained on LitBench reach 78% agreement with upvote-derived human preferences in creative writing, beating all zero-shot LLM judges tested.

  3. Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures

    cs.LG 2025-05 accept novelty 6.0 of 10

    A new benchmark shows that leading LLMs perform poorly on data structure reasoning tasks, with the top model scoring 0.46 on challenging instances.

  4. RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition

    cs.IR 2025-06 conditional novelty 4.0 of 10

    A RAG pipeline using InstructRAG, Pinecone, and BGE placed third in the 2025 LiveRAG Challenge, though internal evaluation only weakly predicted official scores.

Pith tools