Pith. sign in

REVIEW 10 cited by

A Categorical Archive of ChatGPT Failures

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.03494 v8 pith:IDQGBKIH submitted 2023-02-06 cs.CL cs.AIcs.LG

A Categorical Archive of ChatGPT Failures

classification cs.CL cs.AIcs.LG
keywords chatgptfailuresbeenchatbotscomprehensivehumanlanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large language models have been demonstrated to be valuable in different fields. ChatGPT, developed by OpenAI, has been trained using massive amounts of data and simulates human conversation by comprehending context and generating appropriate responses. It has garnered significant attention due to its ability to effectively answer a broad range of human inquiries, with fluent and comprehensive answers surpassing prior public chatbots in both security and usefulness. However, a comprehensive analysis of ChatGPT's failures is lacking, which is the focus of this study. Eleven categories of failures, including reasoning, factual errors, math, coding, and bias, are presented and discussed. The risks, limitations, and societal implications of ChatGPT are also highlighted. The goal of this study is to assist researchers and developers in enhancing future language models and chatbots.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Consistency Training while Mitigating Obfuscation via Rate Matching

    cs.CL 2026-06 unverdicted novelty 6.0

    RMCT matches the rate of target behaviors like bias-following across input perturbations to reduce sycophancy in LLMs while preserving verbalization of bias cues.

  2. U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning

    cs.AI 2026-05 unverdicted novelty 5.0

    U-Define improves user control in LLM planning by letting people define hard rules and soft preferences in natural language with matching verification methods, raising usefulness and satisfaction scores.

  3. Framing Effects in Independent-Agent Large Language Models: A Cross-Family Behavioral Analysis

    cs.CL 2026-03 unverdicted novelty 5.0

    Prompt framing significantly shifts LLM choices toward risk-averse options in a threshold voting task even when the prompts are logically equivalent.

  4. Assessing, Exploiting, and Mitigating Syntactic Robustness Failures in LLM-Based Code Generation

    cs.SE 2024-04 unverdicted novelty 5.0

    LLM code generation lacks syntactic robustness on math-formula prompts, but formula-reduction pre-processing raises it from 54.05% to 74.42%.

  5. TrustLLM: Trustworthiness in Large Language Models

    cs.CL 2024-01 unverdicted novelty 5.0

    TrustLLM defines eight trustworthiness principles, creates a six-dimension benchmark, and evaluates 16 LLMs showing proprietary models generally lead but some open-source ones are close while over-calibration can hurt...

  6. Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

    cs.AI 2023-08 accept novelty 5.0

    Survey organizes LLM trustworthiness into seven categories and 29 sub-categories, measures eight sub-categories on popular models, and finds that more aligned models generally score higher but with varying effectiveness.

  7. Bridging Symbolic Control and Neural Reasoning in LLM Agents -- The Structured Cognitive Loop

    cs.AI 2025-11 reject novelty 4.0

    A five-module LLM agent loop (retrieval, cognition, control, action, memory) is claimed to eliminate policy violations and redundant calls, though validation does not compare against real baselines.

  8. How Secure is Code Generated by ChatGPT?

    cs.CR 2023-04 unverdicted novelty 4.0

    ChatGPT often generates code vulnerable to attacks even when prompted to produce secure code.

  9. Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering

    cs.CY 2026-05 unverdicted novelty 3.0

    LLM graders achieve substantial human agreement on math and science MCAS items but vary on ELA, performing best as sources of formative narrative feedback rather than summative numerical scores.

  10. Enhancing Instructional Quality: Leveraging Computer-Assisted Textual Analysis to Generate In-Depth Insights from Educational Artifacts

    cs.AI 2024-03 unverdicted novelty 3.0

    AI and NLP applied to educational artifacts within the Instructional Core Framework can identify advantages for teacher coaching, student support, and personalized learning.