Pith. sign in

REVIEW 6 cited by

Calibrating LLM-Based Evaluator

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.13308 v1 pith:4TR6JJAR submitted 2023-09-23 cs.CL

classification cs.CL
keywords humanevaluatorlanguagecalibratecriteriaevaluationllm-basedthem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in large language models (LLMs) on language modeling and emergent capabilities make them a promising reference-free evaluator of natural language generation quality, and a competent alternative to human evaluation. However, hindered by the closed-source or high computational demand to host and tune, there is a lack of practice to further calibrate an off-the-shelf LLM-based evaluator towards better human alignment. In this work, we propose AutoCalibrate, a multi-stage, gradient-free approach to automatically calibrate and align an LLM-based evaluator toward human preference. Instead of explicitly modeling human preferences, we first implicitly encompass them within a set of human labels. Then, an initial set of scoring criteria is drafted by the language model itself, leveraging in-context learning on different few-shot examples. To further calibrate this set of criteria, we select the best performers and re-draft them with self-refinement. Our experiments on multiple text quality evaluation datasets illustrate a significant improvement in correlation with expert evaluation through calibration. Our comprehensive qualitative analysis conveys insightful intuitions and observations on the essence of effective scoring criteria.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Statutory Construction and Interpretation for Artificial Intelligence

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Prompt-based legal canons and iterative rule refinement reduce disagreement among LLM judges about whether a response complies with natural-language rules.

  2. Verifiable Format Control for Large Language Model Generations

    cs.CL 2025-02 conditional novelty 5.0 of 10

    A fully verifiable format-following dataset and a progressive SFT-plus-DPO self-improvement pipeline improve 7B LLMs' format control, with mixed out-of-domain transfer.

  3. Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator

    cs.CL 2024-11 conditional novelty 5.0 of 10

    A multi-agent, self-training LLM framework called MESA evaluates meeting summaries by detecting eight error types and reports higher correlation with human scores than existing automatic metrics.

  4. Engineering AI Judge Systems

    cs.SE 2024-11 conditional novelty 5.0 of 10

    A constitution-based, search-driven development framework for AI judge systems improves judged accuracy by up to 6.2% on commit message generation, with about 58% of general principles reused across five languages.

  5. Towards unearthing neglected climate innovations from scientific literature using Large Language Models

    cs.IR 2024-11 conditional novelty 5.0 of 10

    A GPT-4o-based screening workflow with contextual prompting and control-fitted weighting can rank known climate spin-out abstracts highly, though validation is limited to a small sample.

  6. A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A review that organizes LLM uncertainty quantification into token-level, self-verbalized, semantic-similarity, and mechanistic interpretability categories.

Pith tools