Pith. sign in

REVIEW 1 cited by

Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.14044 v1 pith:W65SF6X2 submitted 2024-10-17 cs.IR cs.AI

classification cs.IRcs.AI
keywords relevanceevaluationexplorelabelsllmspassagespromptalternative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Traditional evaluation of information retrieval (IR) systems relies on human-annotated relevance labels, which can be both biased and costly at scale. In this context, large language models (LLMs) offer an alternative by allowing us to directly prompt them to assign relevance labels for passages associated with each query. In this study, we explore alternative methods to directly prompt LLMs for assigned relevance labels, by exploring two hypotheses: Hypothesis 1 assumes that it is helpful to break down "relevance" into specific criteria - exactness, coverage, topicality, and contextual fit. We explore different approaches that prompt large language models (LLMs) to obtain criteria-level grades for all passages, and we consider various ways to aggregate criteria-level grades into a relevance label. Hypothesis 2 assumes that differences in linguistic style between queries and passages may negatively impact the automatic relevance label prediction. We explore whether improvements can be achieved by first synthesizing a summary of the passage in the linguistic style of a query, and then using this summary in place of the passage to assess its relevance. We include an empirical evaluation of our approaches based on data from the LLMJudge challenge run in Summer 2024, where our "Four Prompts" approach obtained the highest scores in Kendall's tau.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment

    cs.IR 2024-12 conditional novelty 5.0 of 10

    Ensembling small open-source LLMs as relevance judges achieves human-correlation scores competitive with GPT-4-based judges on the LLMJudge benchmark.

Pith tools