Pith. sign in

Leveraging LLMs to Evaluate Usefulness of Document

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The conventional Cranfield paradigm struggles to effectively capture user satisfaction due to its weak correlation between relevance and satisfaction, alongside the high costs of relevance annotation in building test collections. To tackle these issues, our research explores the potential of leveraging large language models (LLMs) to generate multilevel usefulness labels for evaluation. We introduce a new user-centric evaluation framework that integrates users' search context and behavioral data into LLMs. This framework uses a cascading judgment structure designed for multilevel usefulness assessments, drawing inspiration from ordinal regression techniques. Our study demonstrates that when well-guided with context and behavioral information, LLMs can accurately evaluate usefulness, allowing our approach to surpass third-party labeling methods. Furthermore, we conduct ablation studies to investigate the influence of key components within the framework. We also apply the labels produced by our method to predict user satisfaction, with real-world experiments indicating that these labels substantially improve the performance of satisfaction prediction models.

citation-role summary

other 1

citation-polarity summary

fields

cs.IR 1

years

2025 1

verdicts

CONDITIONAL 1

roles

other 1

polarities

unclear 1

representative citing papers

LLMs for estimating positional bias in logged interaction data

cs.IR · 2025-09-03 · conditional · novelty 6.0

An LLM-as-a-judge relevance score lets the authors estimate examination propensities from logged clicks, exposing row-column effects in a grid layout and giving an IPS-trained reranker a roughly 2% wNDCG@10 gain.

citing papers explorer

Showing 1 of 1 citing paper.

  • LLMs for estimating positional bias in logged interaction data cs.IR · 2025-09-03 · conditional · none · ref 17 · internal anchor

    An LLM-as-a-judge relevance score lets the authors estimate examination propensities from logged clicks, exposing row-column effects in a grid layout and giving an IPS-trained reranker a roughly 2% wNDCG@10 gain.