Pith. sign in

REVIEW 53 cited by

How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.07597 v1 pith:H7OYAH2Q submitted 2023-01-18 cs.CL

classification cs.CL
keywords chatgpthumanexpertscomparisondatasetablecollectedcomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The introduction of ChatGPT has garnered widespread attention in both academic and industrial communities. ChatGPT is able to respond effectively to a wide range of human questions, providing fluent and comprehensive answers that significantly surpass previous public chatbots in terms of security and usefulness. On one hand, people are curious about how ChatGPT is able to achieve such strength and how far it is from human experts. On the other hand, people are starting to worry about the potential negative impacts that large language models (LLMs) like ChatGPT could have on society, such as fake news, plagiarism, and social security issues. In this work, we collected tens of thousands of comparison responses from both human experts and ChatGPT, with questions ranging from open-domain, financial, medical, legal, and psychological areas. We call the collected dataset the Human ChatGPT Comparison Corpus (HC3). Based on the HC3 dataset, we study the characteristics of ChatGPT's responses, the differences and gaps from human experts, and future directions for LLMs. We conducted comprehensive human evaluations and linguistic analyses of ChatGPT-generated content compared with that of humans, where many interesting results are revealed. After that, we conduct extensive experiments on how to effectively detect whether a certain text is generated by ChatGPT or humans. We build three different detection systems, explore several key factors that influence their effectiveness, and evaluate them in different scenarios. The dataset, code, and models are all publicly available at https://github.com/Hello-SimpleAI/chatgpt-comparison-detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 53 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 293 citations worldwide. Full citation record

  1. Segmenting Human-LLM Co-authored Text via Change Point Detection

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    A weighted change point detection algorithm locates human- and LLM-written segments in mixed text and attains minimax-optimal localization under heterogeneous sentence scores.

  2. DWT-Fusion: A Signal-Based Framework for Training-Free LLM-Generated Text Detection

    cs.CL 2026-07 reject novelty 6.0 of 10

    A training-free detector using discrete wavelet analysis of token log-probability sequences reports AUROC 0.99/0.85/0.75 on HC3/M4/MAGE, but configuration selection on the test split weakens the numbers.

  3. Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it

    cs.CL 2026-07 conditional novelty 6.0 of 10

    LLMs overuse the 'not X, but Y' self-correction pattern in persuasive registers and underuse it in informal Q&A; a prompt or a detachable LoRA dial adjusts it to human levels.

  4. MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark

    cs.CL 2026-01 conditional novelty 6.0 of 10

    Adding human-alignment augmentation (roleplaying, BPO, self-refine, RLDF) to machine-generated text both fools existing detectors and improves the generalization of detectors fine-tuned on it.

  5. Black-Box Detection of LLM-Generated Text Using Generalized Jensen-Shannon Divergence

    cs.LG 2025-10 unverdicted novelty 6.0 of 10

    SurpMark detects machine-generated text by estimating state-transition matrices from discretized surprisals and scoring them with generalized Jensen-Shannon divergence to human versus machine references.

  6. Stack Overflow Is Not Dead Yet: Crowd Answers Still Matter

    cs.CY 2025-09 conditional novelty 6.0 of 10

    After ChatGPT's launch, Stack Overflow posts became longer and shifted toward medium-difficulty questions, suggesting users turn to the crowd for more complex problems.

  7. CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review Detection

    cs.CL 2025-08 conditional novelty 6.0 of 10

    CoCoNUTS is a six-mode peer-review benchmark and CoCoDet a multi-task detector that classifies reviews by content origin, reaching 98% macro F1 in-domain.

  8. EMMM, Explain Me My Model! Explainable Machine Generated Text Detection in Dialogues

    cs.CL 2025-08 conditional novelty 6.0 of 10

    An explanation-then-detection framework for machine-generated text in dialogues that uses dialogue acts and token attributions to produce fast, non-expert-friendly explanations.

  9. Relativistic Quantum Thermal Machine: Harnessing Relativistic Effects to Surpass Carnot Efficiency

    quant-ph 2025-08 unverdicted novelty 6.0 of 10

    Relativistic motion of the reservoirs in a three-level maser is claimed to yield a generalized Carnot bound that allows efficiency above the ordinary Carnot limit.

  10. RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A detector that projects a text's hidden neural activations onto a direction learned from human versus AI writing differences reports higher accuracy than baselines, including on unseen generators and attacked texts.

  11. Beyond Easy Wins: A Text Hardness-Aware Benchmark for LLM-generated Text Detection

    cs.CL 2025-07 conditional novelty 6.0 of 10

    SHIELD shows that standard AUROC overstates AI-text detector quality, and that six zero-shot detectors collapse under a controllable word-replacement humanification.

  12. Detecting LLM-generated Code with Subtle Modification by Adversarial Training

    cs.SE 2025-07 conditional novelty 6.0 of 10

    CodeGPTSensor+, trained with adversarial samples that combine identifier renaming and structure transformation, is substantially more robust to subtle modifications of LLM-generated code than the original CodeGPTSensor.

  13. PhantomHunter: Detecting Unseen Privately-Tuned LLM-Generated Text via Family-Aware Learning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    PhantomHunter detects text from privately fine-tuned LLMs by learning shared token-probability traits within LLaMA, Gemma and Mistral families, reporting F1 above 96% on held-out derivatives.

  14. A General Method for Detecting Information Generated by Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new LLM-text detector built from twin memory networks and domain-generalization losses outperforms prior detectors on held-out LLMs and domains in the paper's benchmark.

  15. When Detection Fails: The Power of Fine-Tuned Models to Generate Human-Like Social Media Text

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Fine-tuned LLMs generate social media text that evades state-of-the-art detectors and human readers, dropping detection accuracy from up to 99.9% to near chance.

  16. DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domains

    cs.CL 2025-06 conditional novelty 6.0 of 10

    DivScore detects AI-written medical and legal text by dividing a domain-tuned model's entropy by its disagreement with a general model, beating baselines on a new benchmark.

  17. Measuring Human Involvement in AI-Generated Text: A Case Study on Academic Writing

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Human involvement in AI-generated academic text can be estimated continuously by training a RoBERTa regressor on BERTScore-derived labels, outperforming binary detectors on a new synthetic dataset.

  18. HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Current machine-generated text detectors, especially metric-based ones, perform poorly on word-level detection in coauthored texts, while finetuned DeBERTa achieves strong but imperfect performance.

  19. The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Arabic text written by LLMs carries detectable stylometric signatures, and fine-tuned XLM-RoBERTa detectors reach near-perfect F1 on academic abstracts but degrade on social media.

  20. Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation

    cs.CR 2025-04 conditional novelty 6.0 of 10

    Q-faker generates adversarial examples for NLU classifiers using surrogate-model gradients and controlled generation, requiring zero queries to the target.

  21. Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A three-method auditing framework detects with roughly 87 to 97 percent accuracy whether classifiers, generators, and t-SNE plots were trained on or derived from LLM-generated synthetic data.

  22. AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A new 171-speech Greek campaign corpus with 31,674 human-validated paragraph annotations for six NLP tasks, plus a ChatGPT-versus-expert accuracy analysis.

  23. Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A learned token-weighting head over next-token probabilities improves AI-text detection accuracy and generalization with about one million trainable parameters.

  24. Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A detector trained on a new multi-LLM social media benchmark estimates that AI-generated text on Medium and Quora grew from about 2% to roughly 37 to 39 percent between 2022 and 2024, while Reddit stayed near 2%.

  25. On the Generalization and Adaptation Ability of Machine-Generated Text Detectors in Academic Writing

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A new 749K-sample academic-writing benchmark shows that supervised detectors outperform metric-based ones, attribution is difficult for metric methods, and adapting to new LLMs leaves a sizable gap.

  26. GPT as ghostwriter at the White House

    cs.CL 2024-11 conditional novelty 6.0 of 10

    ChatGPT-3.5-generated State of the Union speeches are stylistically distinguishable from real presidential addresses across word, sentence, part-of-speech, and rhetorical measures.

  27. ASK-NN: An Asymmetric Nearest-Neighbor Test that detects Distribution Drifts in Natural Language

    cs.LG 2026-07 conditional novelty 5.0 of 10

    ASK-NN, a one-sided nearest-neighbor coincidence test, detects distribution drift between reference and query samples with theoretical guarantees and competitive empirical performance on LLM hallucination and artifici...

  28. Latent Trajectory Discrimination for AI-Generated Text Detection

    cs.CL 2026-07 conditional novelty 5.0 of 10

    A sliding-window, trajectory-difference contrastive learner beats six AI-text detectors on RAID, NYT-AI, and OpenReview reviews.

  29. An Audit and Analysis of LLM-Assisted Health Misinformation Jailbreaks Against LLMs

    cs.CL 2025-08 conditional novelty 5.0 of 10

    LLM-generated jailbreak prompts elicited health misinformation from GPT-3.5, Llama 3.1-8B, and Gemini 2.0 Flash at high rates, and both LLM judges and simple classifiers detected the resulting texts with high accuracy.

  30. Stylometry recognizes human and LLM-generated texts in short samples

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Stylometric features and tree-based classifiers separate human-written Wikipedia summaries from LLM-generated texts with high cross-validated accuracy on a new seven-class benchmark, though performance drops on other ...

  31. Controlling Language Confusion in Multilingual LLMs

    cs.CL 2025-05 conditional novelty 5.0 of 10

    ORPO fine-tuning, which explicitly penalizes disfavored language-mixed responses, nearly eliminates language confusion in Korean-generation LLMs without hurting QA accuracy.

  32. Reliably Bounding False Positives: A Zero-Shot Machine-Generated Text Detection Framework via Multiscaled Conformal Prediction

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A conformal-prediction wrapper with length-binned thresholds bounds the false-positive rate of machine-generated-text detectors at a user-chosen level alpha while improving true-positive rate in experiments.

  33. AI with Emotions: Exploring Emotional Expressions in Large Language Models

    cs.AI 2025-04 conditional novelty 5.0 of 10

    LLMs given numerical arousal and valence coordinates produce text that sentiment analysis places in the same region of Russell's circumplex, demonstrating limited but real control over emotional tone.

  34. Enhancing Health Information Retrieval with RAG by Prioritizing Topical Relevance and Factual Accuracy

    cs.IR 2025-02 conditional novelty 5.0 of 10

    A three-stage RAG pipeline generates a cited summary (GenText) from PubMed Central passages and ranks health documents by topical relevance plus alignment with that summary, outperforming baselines on CLEF eHealth and...

  35. GenAI Content Detection Task 1: English and Multilingual Machine-Generated Text Detection: AI vs. Human

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A COLING 2025 shared task benchmark showing that current machine-generated text detectors reach only moderate accuracy and degrade badly on out-of-domain and humanized AI text.

  36. A Reproducibility and Generalizability Study of Large Language Models for Query Generation

    cs.IR 2024-11 conditional novelty 5.0 of 10

    LLM-generated Boolean queries for systematic reviews are unstable across seeds, and the original ChatGPT results could not be reproduced with the documented setup.

  37. Large Language Models as Robust Data Generators in Software Analytics: Are We There Yet?

    cs.SE 2024-11 conditional novelty 5.0 of 10

    Pre-trained models fine-tuned on LLM-generated data are less robust to adversarial attacks than models fine-tuned on human-written data in three software analytics tasks.

  38. Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration

    cs.CL 2026-08 conditional novelty 4.0 of 10

    EchoPrompt detects LLM-generated text by measuring the likelihood gain from restoring a generic assistant-style prompt, calibrated against a base model.

  39. A Unified Detection Framework for AI-Related Content and Artifacts

    stat.ML 2026-07 conditional novelty 4.0 of 10

    The paper develops joint casewise and cellwise MCD estimators for multi-class robust covariance estimation and applies them within a Mahalanobis-distance-based detection framework across four AI-content detection task...

  40. A Comprehensive Dataset for Human vs. AI Generated Text Detection

    cs.CL 2025-10 reject novelty 4.0 of 10

    A dataset of ~58k NYT articles plus AI rewrites from six LLMs, evaluated with a rewrite-distance baseline reaching 58.35% detection and 8.92% attribution accuracy.

  41. MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds

    cs.CL 2025-09 conditional novelty 4.0 of 10

    MoSEs improves AI-text detection by replacing a single score threshold with a threshold learned from style-matched reference texts and linguistic statistics.

  42. Artificial Intelligence and Civil Discourse: How LLMs Moderate Climate Change Conversations

    cs.CY 2025-06 reject novelty 4.0 of 10

    LLM replies to climate change posts are more emotionally neutral and lower in intensity than the human posts they respond to.

  43. A Practical Guide for Supporting Formative Assessment and Feedback Using Generative AI

    cs.CY 2025-05 conditional novelty 4.0 of 10

    A narrative review that aligns generative AI tools with formative assessment principles, provides classroom prompt examples, and identifies missing evaluation metrics for AI feedback.

  44. Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems

    cs.IR 2025-05 conditional novelty 4.0 of 10

    A reranker fine-tuned on hard negatives selected by two cosine-distance criteria outperforms older negative sampling methods on enterprise and domain-specific retrieval benchmarks.

  45. Towards Structurally Explainable Machine-Generated Text Detection: A Graph-Perspective Framework

    cs.CL 2025-05 reject novelty 4.0 of 10

    LM2OTIFS uses word co-occurrence graphs and GNNExplainer to detect and explain machine-generated text, with strong in-domain accuracy but unsupported faithfulness claims and a flawed theoretical proof.

  46. Kill two birds with one stone: generalized and robust AI-generated text detection via dynamic perturbations

    cs.CL 2025-04 conditional novelty 4.0 of 10

    DP-Net uses a DDPG reinforcement-learning agent to dynamically tune Gaussian perturbation of embeddings, improving average cross-domain accuracy and adversarial robustness of AI-generated text detection.

  47. GenAI Content Detection Task 2: AI vs. Human -- Academic Essay Authenticity Challenge

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A bilingual shared-task benchmark shows that fine-tuned classifiers distinguish AI-generated from human-written academic essays with macro-F1 above 0.98 within the challenge's test sets.

  48. How good is GPT at writing political speeches for the White House?

    cs.CL 2024-12 conditional novelty 4.0 of 10

    GPT-generated State of the Union speeches differ stylistically from real presidential speeches, with more "we", more positive and abstract terms, and lower authenticity.

  49. PickLLM: Context-Aware RL-Assisted Large Language Model Routing

    cs.LG 2024-12 conditional novelty 4.0 of 10

    A reinforcement-learning router that converges to one LLM per query session, reducing cost and latency while keeping answer quality competitive.

  50. Human Variability vs. Machine Consistency: A Linguistic Analysis of Texts Generated by Humans and Large Language Models

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Human-written texts show higher linguistic variability, richer vocabulary and emotional content, and lower syntactic depth than texts from five LLMs.

  51. StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features

    cs.CL 2025-07 conditional novelty 3.0 of 10

    A non-neural stylometric detector using LightGBM on spaCy-derived features reached a final mean score of 0.897 on the PAN 2025 task, below the 0.922 TF-IDF SVM baseline.

  52. Survey on AI-Generated Media Detection: From Non-MLLM to MLLM

    cs.CV 2025-02 unverdicted novelty 3.0 of 10

    A survey organizing AI-generated media detection into Non-MLLM and MLLM based methods, with task and benchmark taxonomies.

  53. A Survey on Large Language Models with some Insights on their Capabilities and Limitations

    cs.CL 2025-01 unverdicted novelty 3.0 of 10

    A broad survey of LLM methods and applications, plus an empirical section on how code-rich pretraining may influence chain-of-thought reasoning, the details of which are not visible in the supplied text.

Pith tools