Pith. sign in

REVIEW 3 cited by

In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.08979 v2 pith:VEOQCOVP submitted 2023-04-18 cs.CR cs.LG

classification cs.CRcs.LG
keywords chatgptreliabilityusersacrossdomainsquestionsacquireadvent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The way users acquire information is undergoing a paradigm shift with the advent of ChatGPT. Unlike conventional search engines, ChatGPT retrieves knowledge from the model itself and generates answers for users. ChatGPT's impressive question-answering (QA) capability has attracted more than 100 million users within a short period of time but has also raised concerns regarding its reliability. In this paper, we perform the first large-scale measurement of ChatGPT's reliability in the generic QA scenario with a carefully curated set of 5,695 questions across ten datasets and eight domains. We find that ChatGPT's reliability varies across different domains, especially underperforming in law and science questions. We also demonstrate that system roles, originally designed by OpenAI to allow users to steer ChatGPT's behavior, can impact ChatGPT's reliability in an imperceptible way. We further show that ChatGPT is vulnerable to adversarial examples, and even a single character change can negatively affect its reliability in certain cases. We believe that our study provides valuable insights into ChatGPT's reliability and underscores the need for strengthening the reliability and security of large language models (LLMs).

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A new Persian-Islamic trustworthiness benchmark ranks Claude highest and Qwen lowest across eight LLMs and finds safety is the weakest dimension.

  2. Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Gender-diverse users perceive ChatGPT's gender bias differently, with non-binary/transgender participants reporting condescending and stereotypical responses, and men reporting higher trust.

  3. What Shapes User Trust in ChatGPT? A Mixed-Methods Study of User Attributes, Trust Dimensions, Task Context, and Societal Perceptions among University Students

    cs.HC 2025-07 conditional novelty 4.0 of 10

    Frequent use, perceived expertise, and ethical risk perceptions predict university students' trust in ChatGPT, while trust varies strongly by task type.

Pith tools