Pith. sign in

REVIEW 1 cited by

The ICL Consistency Test

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.04945 v1 pith:TZ7WSOGT submitted 2023-12-08 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords consistencysetupsdifferentlackmetricmodelmodelstest
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Just like the previous generation of task-tuned models, large language models (LLMs) that are adapted to tasks via prompt-based methods like in-context-learning (ICL) perform well in some setups but not in others. This lack of consistency in prompt-based learning hints at a lack of robust generalisation. We here introduce the ICL consistency test -- a contribution to the GenBench collaborative benchmark task (CBT) -- which evaluates how consistent a model makes predictions across many different setups while using the same data. The test is based on different established natural language inference tasks. We provide preprocessed data constituting 96 different 'setups' and a metric that estimates model consistency across these setups. The metric is provided on a fine-grained level to understand what properties of a setup render predictions unstable and on an aggregated level to compare overall model consistency. We conduct an empirical analysis of eight state-of-the-art models, and our consistency metric reveals how all tested LLMs lack robust generalisation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GPAI Evaluations Standards Taskforce: Towards Effective AI Governance

    cs.CY 2024-11 conditional novelty 5.0 of 10

    The paper proposes an EU GPAI Evaluation Standards Taskforce to develop adaptive standards for AI evaluations, based on four desiderata: internal validity, external validity, reproducibility, and portability.

Pith tools