Pith. sign in

REVIEW 7 cited by

TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.13296 v1 pith:EX72AIKT submitted 2021-09-27 cs.CL

TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation

classification cs.CL
keywords gpt-2turingbenchbenchmarkmodelsenvironmentfairgroverlarge
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recent progress in generative language models has enabled machines to generate astonishingly realistic texts. While there are many legitimate applications of such models, there is also a rising need to distinguish machine-generated texts from human-written ones (e.g., fake news detection). However, to our best knowledge, there is currently no benchmark environment with datasets and tasks to systematically study the so-called "Turing Test" problem for neural text generation methods. In this work, we present the TuringBench benchmark environment, which is comprised of (1) a dataset with 200K human- or machine-generated samples across 20 labels {Human, GPT-1, GPT-2_small, GPT-2_medium, GPT-2_large, GPT-2_xl, GPT-2_PyTorch, GPT-3, GROVER_base, GROVER_large, GROVER_mega, CTRL, XLM, XLNET_base, XLNET_large, FAIR_wmt19, FAIR_wmt20, TRANSFORMER_XL, PPLM_distil, PPLM_gpt2}, (2) two benchmark tasks -- i.e., Turing Test (TT) and Authorship Attribution (AA), and (3) a website with leaderboards. Our preliminary experimental results using TuringBench show that FAIR_wmt20 and GPT-3 are the current winners, among all language models tested, in generating the most human-like indistinguishable texts with the lowest F1 score by five state-of-the-art TT detection models. The TuringBench is available at: https://turingbench.ist.psu.edu/

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AEyeDE: An Attention-Based Attribution Framework for AI-Generated Text Detection

    cs.CL 2026-04 conditional novelty 6.0

    Attention attribution maps from a white-box proxy Transformer, classified by a lightweight CNN, provide a competitive and interpretable signal for AI-generated text detection.

  2. MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark

    cs.CL 2026-01 conditional novelty 6.0

    Adding human-alignment augmentation (roleplaying, BPO, self-refine, RLDF) to machine-generated text both fools existing detectors and improves the generalization of detectors fine-tuned on it.

  3. GigaCheck: Detecting LLM-generated Content via Object-Centric Span Localization

    cs.CL 2024-10 unverdicted novelty 6.0

    GigaCheck detects LLM-generated text at both document and span levels by combining fine-tuned language-model embeddings with a DETR-like architecture that treats generated intervals as detectable objects.

  4. C-ReD: A Comprehensive Chinese Benchmark for AI-Generated Text Detection Derived from Real-World Prompts

    cs.CL 2026-04 unverdicted novelty 5.0

    C-ReD is a new Chinese benchmark for AI-generated text detection built from diverse real-world prompts to improve in-domain performance and generalization to unseen models and datasets.

  5. C-ReD: A Comprehensive Chinese Benchmark for AI-Generated Text Detection Derived from Real-World Prompts

    cs.CL 2026-04 unverdicted novelty 4.0

    C-ReD is a Chinese AI-text detection benchmark built from diverse real-world prompts and multiple LLMs that shows strong in-domain performance and generalization to unseen models and external datasets.

  6. A Comprehensive Dataset for Human vs. AI Generated Text Detection

    cs.CL 2025-10 reject novelty 4.0

    A dataset of ~58k NYT articles plus AI rewrites from six LLMs, evaluated with a rewrite-distance baseline reaching 58.35% detection and 8.92% attribution accuracy.

  7. Introduction to the artificial neural network-based variational Monte Carlo method

    physics.comp-ph 2026-03 unverdicted novelty 3.0

    The paper introduces neural-network trial wave functions for variational Monte Carlo, frames the variational method as unsupervised learning, and illustrates the approach on the Yukawa potential and hydrogen molecule.