Pith. sign in

REVIEW 16 cited by

Prometheus: Inducing Fine-grained Evaluation Capability in Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08491 v2 pith:ZJDBFMWH submitted 2023-10-12 cs.CL cs.LG

classification cs.CLcs.LG
keywords prometheusgpt-4scorebenchevaluatorfeedbackhumancustomized
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, using a powerful proprietary Large Language Model (LLM) (e.g., GPT-4) as an evaluator for long-form responses has become the de facto standard. However, for practitioners with large-scale evaluation tasks and custom criteria in consideration (e.g., child-readability), using proprietary LLMs as an evaluator is unreliable due to the closed-source nature, uncontrolled versioning, and prohibitive costs. In this work, we propose Prometheus, a fully open-source LLM that is on par with GPT-4's evaluation capabilities when the appropriate reference materials (reference answer, score rubric) are accompanied. We first construct the Feedback Collection, a new dataset that consists of 1K fine-grained score rubrics, 20K instructions, and 100K responses and language feedback generated by GPT-4. Using the Feedback Collection, we train Prometheus, a 13B evaluator LLM that can assess any given long-form text based on customized score rubric provided by the user. Experimental results show that Prometheus scores a Pearson correlation of 0.897 with human evaluators when evaluating with 45 customized score rubrics, which is on par with GPT-4 (0.882), and greatly outperforms ChatGPT (0.392). Furthermore, measuring correlation with GPT-4 with 1222 customized score rubrics across four benchmarks (MT Bench, Vicuna Bench, Feedback Bench, Flask Eval) shows similar trends, bolstering Prometheus's capability as an evaluator LLM. Lastly, Prometheus achieves the highest accuracy on two human preference benchmarks (HHH Alignment & MT Bench Human Judgment) compared to open-sourced reward models explicitly trained on human preference datasets, highlighting its potential as an universal reward model. We open-source our code, dataset, and model at https://kaistai.github.io/prometheus/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Pluralis v0.1 is a culture-first, multimodal, multilingual VLM safety benchmark spanning 6 APAC locales with 6,448 prompts and an agreement-gated LLM judge that disentangles safety from cultural appropriateness.

  2. Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

    cs.LG 2026-07 conditional novelty 6.5 of 10

    LLM-as-judge scoring biases concentrate in low-dimensional, type-specific activation subspaces that support bidirectional causal steering and cross-domain failure prediction.

  3. Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG

    cs.CL 2026-07 conditional novelty 6.5 of 10

    Answer-paired analysis of a 3×3 GPT/Grok/Gemini judge matrix finds near-zero same-model recall bias for induced RAG grounding errors; remaining flag gaps reflect label-task mismatch.

  4. SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    SVGEval benchmarks and trains an explainable multimodal scorer for perceptual quality of text-to-SVG generation, showing a consistent gap on spatial and structural judgments.

  5. Role Steering of Language Models for Social Simulations

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A role-steering screening workflow on 275 roles shows role-specific activation directions beat a non-scale-matched assistant-direction control (63.2 vs 41.1 judged alignment) and flags 38 roles as 'anti-controllable'.

  6. Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Cheaper LLM judges match frontier models on citation-quality F1 but differ substantially in false positive and false negative rates, meaning reward signal calibration matters more than model cost.

  7. Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

    cs.CL 2026-07 unverdicted novelty 6.0 of 10

    LLM evaluators reach clinician-level agreement on a new German medical benchmark but fail to abstain on difficult items and show lineage-dependent scoring biases.

  8. Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

    cs.CL 2026-05 conditional novelty 6.0 of 10

    On binary verdicts, Pearson, Spearman, Kendall's tau-b, phi, and the Matthews correlation are a single statistic, so most multi-metric agreement reports repeat one number under different names.

  9. FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new benchmark and 3D decomposition paradigm for factuality evaluation of interpretive claims about contact center conversations, with best LLM-judge F1 of 0.86.

  10. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  11. When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Judge upgrades are not interchangeable: only Qwen3 1.7B→4B yields robust adjacent gains, MiniMax adjacent releases do not, and stronger judges reduce but do not remove bias or correlated jury errors.

  12. Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Hybrid RAG over UK public health guidance sharply raises MCQA accuracy and free-form faithfulness, letting smaller open models match larger closed models without retrieval.

  13. CRAFT: Learn the Schema, Execute the Plan

    cs.AI 2026-06 conditional novelty 5.0 of 10

    CRAFT, a two-stage post-training recipe that strips schema documentation from prompts and uses execution-grounded reinforcement learning, reports improved enterprise coding-agent quality at roughly 9x lower input-token cost.

  14. InfluMatch: Frontier-Quality KOL Search at 4B-Model Cost

    cs.CL 2026-07 conditional novelty 4.0 of 10

    A 4B-model cascade for Thai KOL matching reaches 94.1% P@5 on 11 queries, matching a frontier model, with pairwise SimPO training transferring end-to-end while pointwise SFT+GRPO does not.

  15. Beyond Text: Aligning Vision and Language for Multimodal E-Commerce Retrieval

    cs.IR 2026-03 conditional novelty 3.0 of 10

    A commercial 150-control system-prompt governance layer (MDBC) is reported to cut aggregate LLM risk exposure by 36.8% relative to base, far more than a generic safety prompt, under LLM-judge red-teaming.

  16. PARAM-1 BharatGen 2.9B Model

    cs.CL 2025-07 reject novelty 3.0 of 10

    A technical report on a 2.9B English-Hindi model whose headline evaluation numbers are internally inconsistent and whose promoted tokenizer was not used to train the final model.

Pith tools