Pith. sign in

REVIEW 1 cited by

Can ChatGPT and Bard Generate Aligned Assessment Items? A Reliability Analysis against Human Performance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.05372 v1 pith:H7AMOGA3 submitted 2023-04-09 cs.CL cs.AI

Can ChatGPT and Bard Generate Aligned Assessment Items? A Reliability Analysis against Human Performance

classification cs.CL cs.AI
keywords assessmentbardchatgpthumanreliabilityapplicationsautomatedbeen
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

ChatGPT and Bard are AI chatbots based on Large Language Models (LLM) that are slated to promise different applications in diverse areas. In education, these AI technologies have been tested for applications in assessment and teaching. In assessment, AI has long been used in automated essay scoring and automated item generation. One psychometric property that these tools must have to assist or replace humans in assessment is high reliability in terms of agreement between AI scores and human raters. In this paper, we measure the reliability of OpenAI ChatGP and Google Bard LLMs tools against experienced and trained humans in perceiving and rating the complexity of writing prompts. Intraclass correlation (ICC) as a performance metric showed that the inter-reliability of both the OpenAI ChatGPT and the Google Bard were low against the gold standard of human ratings.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Data-driven Progressive Discovery of Physical Laws

    cs.LG 2026-03 unverdicted novelty 5.0

    CoSR discovers physical laws via progressive chains of symbolic knowledge units, recovering Kepler-to-Newton and improving scaling laws in convection, pipe flow, laser-metal interaction, and aircraft aerodynamics.