Pith. sign in

REVIEW 2 cited by

ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.14106 v2 pith:XEX2BKIF submitted 2023-04-27 cs.CL cs.AI

ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time

classification cs.CL cs.AI
keywords chatgptbenchmarkschatlogfeaturesdetectionevaluatingevaluationfind
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

ChatGPT has achieved great success and can be considered to have acquired an infrastructural status. There are abundant works for evaluating ChatGPT on benchmarks. However, existing benchmarks encounter two challenges: (1) Disregard for periodical evaluation and (2) Lack of fine-grained features. In this paper, we construct ChatLog, an ever-updating dataset with large-scale records of diverse long-form ChatGPT responses for 21 NLP benchmarks from March, 2023 to now. We conduct a comprehensive performance evaluation to find that most capabilities of ChatGPT improve over time except for some abilities, and there exists a step-wise evolving pattern of ChatGPT. We further analyze the inherent characteristics of ChatGPT by extracting the knowledge and linguistic features. We find some stable features that stay unchanged and apply them on the detection of ChatGPT-generated texts to improve the robustness of cross-version detection. We will continuously maintain our project at \url{https://github.com/THU-KEG/ChatLog/}.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Relativistic Quantum Thermal Machine: Harnessing Relativistic Effects to Surpass Carnot Efficiency

    quant-ph 2025-08 unverdicted novelty 6.0

    Relativistic motion of the reservoirs in a three-level maser is claimed to yield a generalized Carnot bound that allows efficiency above the ordinary Carnot limit.

  2. Large Databases Need Small, Open-Weight Language Models

    cs.AI 2026-06 unverdicted novelty 4.0

    Quantized open-weight LMs on consumer hardware match closed-source API accuracy for LM-enhanced relational operators while delivering 390x lower cost and 3.8x lower latency in the BlendSQL framework.