Pith. sign in

REVIEW 8 cited by

Large Language Models Pass the Turing Test

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.23674 v1 pith:Z45HVTJX submitted 2025-03-31 cs.CL cs.HC

classification cs.CLcs.HC
keywords humanmodelssignificantlysystemsturingelizagpt-4gpt-4o
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We evaluated 4 systems (ELIZA, GPT-4o, LLaMa-3.1-405B, and GPT-4.5) in two randomised, controlled, and pre-registered Turing tests on independent populations. Participants had 5 minute conversations simultaneously with another human participant and one of these systems before judging which conversational partner they thought was human. When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time -- not significantly more or less often than the humans they were being compared to -- while baseline models (ELIZA and GPT-4o) achieved win rates significantly below chance (23% and 21% respectively). The results constitute the first empirical evidence that any artificial system passes a standard three-party Turing test. The results have implications for debates about what kind of intelligence is exhibited by Large Language Models (LLMs), and the social and economic impacts these systems are likely to have.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Too Human to Model:The Uncanny Valley of LLMs in Social Simulation -- When Generative Language Agents Misalign with Modelling Principles

    cs.CY 2025-07 conditional novelty 7.0 of 10

    A position paper contends that LLM agents, despite their human-like talk, are often too rich in detail to serve as scientific models, and proposes conditions where they still excel.

  2. More Human or More AI? Visualizing Human-AI Collaboration Disclosures in Journalistic News Production

    cs.HC 2026-01 conditional novelty 6.0 of 10

    Disclosure visualization format systematically shifts readers' perceptions of human vs AI contribution: role-based timelines amplify perceived AI role in mostly human articles, while task-based timelines make mostly A...

  3. Playstyle and Artificial Intelligence: An Initial Blueprint Through the Lens of Video Games

    cs.AI 2025-08 conditional novelty 6.0 of 10

    This dissertation formalizes playstyle as the decision-making style of an agent, introduces a discrete-state playstyle distance that distinguishes behaviors in racing games and Atari, and proposes a blueprint for usin...

  4. The Other Mind: How Language Models Exhibit Human Temporal Cognition

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.

  5. A Collaborative Framework Integrating Large Language Model and Chemical Fragment Space: Mutual Inspiration for Lead Design

    q-bio.BM 2025-07 conditional novelty 6.0 of 10

    A closed-loop design framework that uses chemical fragments to guide a large language model in generating novel, high-affinity lead compounds outperformed existing de novo drug design methods.

  6. Large Language Models are Near-Optimal Decision-Makers with a Non-Human Learning Behavior

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Across uncertainty, risk, and set-shifting tasks, LLMs generally outperformed humans and neared optimality while exhibiting distinctly non-human decision-making processes.

  7. CoVeR: Conformal Calibration for Versatile and Reliable Autoregressive Next-Token Prediction

    cs.LG 2025-09 reject novelty 5.0 of 10

    CoVeR is a cluster-aware conformal decoding method that claims full-sequence coverage for LLM outputs without the (1-alpha)^L decay of prior conformal beam search.

  8. Artificially intelligent agents in the social and behavioral sciences: A history and outlook

    cs.AI 2025-10 conditional novelty 2.0 of 10

    AI and social science have co-evolved for 75 years through rapid technological adoption and slower scientific consolidation, with direct human-focused AI studies still scarce.

Pith tools