Pith. sign in

REVIEW 1 cited by

ChatGPT Prompting Cannot Estimate Predictive Uncertainty in High-Resource Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.06427 v1 pith:C74UUWKF submitted 2023-11-10 cs.CL cs.LG

classification cs.CLcs.LG
keywords languageschatgpthigh-resourceconfidenceperformtasksabilitiesanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

ChatGPT took the world by storm for its impressive abilities. Due to its release without documentation, scientists immediately attempted to identify its limits, mainly through its performance in natural language processing (NLP) tasks. This paper aims to join the growing literature regarding ChatGPT's abilities by focusing on its performance in high-resource languages and on its capacity to predict its answers' accuracy by giving a confidence level. The analysis of high-resource languages is of interest as studies have shown that low-resource languages perform worse than English in NLP tasks, but no study so far has analysed whether high-resource languages perform as well as English. The analysis of ChatGPT's confidence calibration has not been carried out before either and is critical to learn about ChatGPT's trustworthiness. In order to study these two aspects, five high-resource languages and two NLP tasks were chosen. ChatGPT was asked to perform both tasks in the five languages and to give a numerical confidence value for each answer. The results show that all the selected high-resource languages perform similarly and that ChatGPT does not have a good confidence calibration, often being overconfident and never giving low confidence values.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One For All: LLM-based Heterogeneous Mission Planning in Precision Agriculture

    cs.RO 2025-06 conditional novelty 4.0 of 10

    An LLM-based planner with XML-schema validation generates behavior-tree missions in plain language for both a Husky rover and a Kinova arm in agricultural scenarios.

Pith tools