REVIEW 8 cited by
Large Language Models Pass the Turing Test
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We evaluated 4 systems (ELIZA, GPT-4o, LLaMa-3.1-405B, and GPT-4.5) in two randomised, controlled, and pre-registered Turing tests on independent populations. Participants had 5 minute conversations simultaneously with another human participant and one of these systems before judging which conversational partner they thought was human. When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time -- not significantly more or less often than the humans they were being compared to -- while baseline models (ELIZA and GPT-4o) achieved win rates significantly below chance (23% and 21% respectively). The results constitute the first empirical evidence that any artificial system passes a standard three-party Turing test. The results have implications for debates about what kind of intelligence is exhibited by Large Language Models (LLMs), and the social and economic impacts these systems are likely to have.
Forward citations
Cited by 8 Pith papers
-
Too Human to Model:The Uncanny Valley of LLMs in Social Simulation -- When Generative Language Agents Misalign with Modelling Principles
A position paper contends that LLM agents, despite their human-like talk, are often too rich in detail to serve as scientific models, and proposes conditions where they still excel.
-
More Human or More AI? Visualizing Human-AI Collaboration Disclosures in Journalistic News Production
Disclosure visualization format systematically shifts readers' perceptions of human vs AI contribution: role-based timelines amplify perceived AI role in mostly human articles, while task-based timelines make mostly A...
-
Playstyle and Artificial Intelligence: An Initial Blueprint Through the Lens of Video Games
This dissertation formalizes playstyle as the decision-making style of an agent, introduces a discrete-state playstyle distance that distinguishes behaviors in racing games and Atari, and proposes a blueprint for usin...
-
The Other Mind: How Language Models Exhibit Human Temporal Cognition
Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.
-
A Collaborative Framework Integrating Large Language Model and Chemical Fragment Space: Mutual Inspiration for Lead Design
A closed-loop design framework that uses chemical fragments to guide a large language model in generating novel, high-affinity lead compounds outperformed existing de novo drug design methods.
-
Large Language Models are Near-Optimal Decision-Makers with a Non-Human Learning Behavior
Across uncertainty, risk, and set-shifting tasks, LLMs generally outperformed humans and neared optimality while exhibiting distinctly non-human decision-making processes.
-
CoVeR: Conformal Calibration for Versatile and Reliable Autoregressive Next-Token Prediction
CoVeR is a cluster-aware conformal decoding method that claims full-sequence coverage for LLM outputs without the (1-alpha)^L decay of prior conformal beam search.
-
Artificially intelligent agents in the social and behavioral sciences: A history and outlook
AI and social science have co-evolved for 75 years through rapid technological adoption and slower scientific consolidation, with direct human-focused AI studies still scarce.
Discussion (0). Sign in to comment.