Pith. sign in

REVIEW 2 cited by

Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.00873 v2 pith:F3RAQP7C submitted 2024-10-01 cs.HC

classification cs.HC
keywords assessmentevaluationsevaluatorsjudgmentsmodelcriteriacriticaldata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Evaluation of large language model (LLM) outputs requires users to make critical judgments about the best outputs across various configurations. This process is costly and takes time given the large amounts of data. LLMs are increasingly used as evaluators to filter training data, evaluate model performance or assist human evaluators with detailed assessments. To support this process, effective front-end tools are critical for evaluation. Two common approaches for using LLMs as evaluators are direct assessment and pairwise comparison. In our study with machine learning practitioners (n=15), each completing 6 tasks yielding 131 evaluations, we explore how task-related factors and assessment strategies influence criteria refinement and user perceptions. Findings show that users performed more evaluations with direct assessment by making criteria task-specific, modifying judgments, and changing the evaluator model. We conclude with recommendations for how systems can better support interactions in LLM-assisted evaluations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring the Potential of LLMs for Serendipity Evaluation in Recommender Systems

    cs.IR 2025-07 conditional novelty 6.0 of 10

    Basic and multi-model LLM prompts can evaluate recommendation serendipity as well as or better than standard proxy formulas, reaching 21.5% Pearson correlation with user-study ratings.

  2. Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks

    cs.HC 2025-07 conditional novelty 6.0 of 10

    Data scientists construct prediction targets through bricolage, applying five reformulation strategies (piggybacking, composing, swapping, bridging, refining) to balance five criteria: validity, simplicity, predictabi...

Pith tools