Pith. sign in

REVIEW 1 cited by

Can Many-Shot In-Context Learning Help LLMs as Evaluators? A Preliminary Empirical Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11629 v6 pith:YRHVAUTX submitted 2024-06-17 cs.CL

classification cs.CL
keywords llmsmany-shotevaluationin-contextevaluatorspromptresultsbetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Utilizing Large Language Models (LLMs) as evaluators to assess the performance of LLMs has garnered attention. However, this kind of evaluation approach is affected by potential biases within LLMs, raising concerns about the accuracy and reliability of the evaluation results of LLMs. To address this problem, we propose and study two many-shot In-Context Learning (ICL) prompt templates to help LLM evaluators mitigate potential biases: Many-Shot with Reference (MSwR) and Many-Shot without Reference (MSoR). Specifically, the former utilizes in-context examples with model-generated evaluation rationales as references, while the latter does not include these references. Using these prompt designs, we investigate the impact of increasing the number of in-context examples on the consistency and quality of the evaluation results. Experimental results show that advanced LLMs, such as GPT-4o, perform better in the many-shot regime than in the zero-shot and few-shot regimes. Furthermore, when using GPT-4o as an evaluator in the many-shot regime, adopting MSwR as the prompt template performs better than MSoR.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MiMoTable: A Multi-scale Spreadsheet Benchmark with Meta Operations for Table Reasoning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    MiMoTable is a real-world spreadsheet benchmark with 1,719 bilingual question-answer pairs and a meta-operation difficulty criterion on which the best LLM scores 77.4%.

Pith tools