Pith. sign in

REVIEW 2 cited by

ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.18638 v2 pith:GY5JIL7W submitted 2024-05-28 cs.CL cs.AI

classification cs.CLcs.AI
keywords evaluationhumancognitivegenerativelanguagelargemodelsbiases
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this position paper, we argue that human evaluation of generative large language models (LLMs) should be a multidisciplinary undertaking that draws upon insights from disciplines such as user experience research and human behavioral psychology to ensure that the experimental design and results are reliable. The conclusions from these evaluations, thus, must consider factors such as usability, aesthetics, and cognitive biases. We highlight how cognitive biases can conflate fluent information and truthfulness, and how cognitive uncertainty affects the reliability of rating scores such as Likert. Furthermore, the evaluation should differentiate the capabilities and weaknesses of increasingly powerful large language models -- which requires effective test sets. The scalability of human evaluation is also crucial to wider adoption. Hence, to design an effective human evaluation system in the age of generative NLP, we propose the ConSiDERS-The-Human evaluation framework consisting of 6 pillars -- Consistency, Scoring Criteria, Differentiating, User Experience, Responsible, and Scalability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Extension Decisions in Open Source Software Ecosystem

    cs.SE 2025-07 reject novelty 6.0 of 10

    A GitHub Actions graph study reports that most new CI tools duplicate existing functionality and that a handful of early tools become the templates for later copies, although the supporting calculation is missing from...

  2. The Impact of Foundational Models on Patient-Centric e-Health Systems

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Most patient-centric e-health apps on Google Play remain at early stages of AI maturity; only 13.79% show advanced AI integration.

Pith tools