REVIEW 4 major objections 5 minor 13 references
Using AI to replicate human experimental results: a motion study
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims that a large language model given the same tasks as human participants reaches the same interpretive conclusions across four psycholinguistic experiments on time and motion verbs.
desk verdict A useful but under-specified proof-of-concept: the authors show LLM-human convergence on four psycholinguistic tasks, but the missing prompts and model selection details make the central substitution claim untestable as reported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the comparison between two response distributions: human participants recruited through online platforms and the LLM ChatGPT o1 prompted with what the authors state was the exact instruction text given to humans. The statistical machinery is rank-order correlation (Spearman's $\rho$) between human and LLM ratings across items, supplemented by item-by-item difference heatmaps and percentage agreement on categorical choices. The underlying assumption is that if the LLM's relative ordering and qualitative pattern match human ordering on the same prompts, then the LLM captures the interpretive regularities at issue—here, the mapping of motion-verb speed onto affective evaluations of time.
What would settle it
Rerun the four experiments with the authors' exact prompt text on a new set of manner-of-motion verbs and with an open-weight language model; if the Spearman correlations fall below 0.5 or the qualitative valence-polarization pattern reverses, the claim that LLM responses replicate human interpretive outcomes for these tasks would be falsified.
Extended reading notes
Core claim
The central claim is that, for the restricted domain of affective meanings triggered by manner-of-motion verbs in temporal expressions, the tested LLM (ChatGPT o1) reproduces human experimental outcomes closely enough that the same conclusions would be drawn if only the LLM data were used. In Study 1, ratings of nine semantic dimensions across fifteen sentences correlated with human ratings at $\rho = .73$, with most item-by-item deviations below one point on a seven-point scale. In Study 2, valence ratings in literal versus metaphorical contexts correlated at $\rho = .96$ and showed the same polarizing pattern: fast verbs became more positive and slow verbs more negative in temporal settings. In Study 3, the LLM's forced-choice verb selections matched human choices exactly, and its probability estimates fell within 10–26 percentage points of human choice percentages. In Study 4, emoji choices matched human choices 100% of the time, with a Spearman correlation of .69 that rose to .73 when two emojis the model misread were excluded. The paper concludes that the LLM's answers are fully compatible with human data and that the divergences, where present, do not change the interpretation.
Load-bearing premise
The conclusion depends on the assumption that the prompt given to the LLM was exactly equivalent to the instructions given to human participants, so that any differences in responses come from the model rather than from a different statement of the task.
Editorial extensions
If this is right
- If the replication holds, psycholinguistic stimulus sets can be scaled from dozens of sentences to hundreds, allowing broader empirical coverage without new human data collection.
- LLM-generated responses can serve as a hypothesis-generation tool, suggesting new affective meanings that researchers then validate with targeted human studies.
- The convergence supports a 'convergent evidence' interpretation, strengthening the empirical foundation of the original human-based findings on manner-of-motion verbs.
- The success of an LLM on these tasks indicates that the affective information is largely encoded in language statistics, not only in embodied experience.
- Researchers can use LLM probability outputs to approximate the strength of human preferences in forced-choice tasks, as the paper demonstrates with the fill-in-the-blank and emoji studies.
Reading between the lines
- If the convergence generalizes beyond the tested verb set, the same method could be used as a cheap pretest before running costly human experiments, flagging stimuli that need human validation.
- The dependence on prompt wording is untested; a natural extension would be to vary the instruction phrasing and measure how much the LLM-human agreement shifts, since replication results may be sensitive to framing.
- The authors' speculation about the verb 'race' being skewed by noun frequency could be tested directly with probing methods or frequency-controlled stimuli, providing a concrete next step.
- A stronger test would be to run the same replication in a different language or with a non-GPT model; convergence there would argue for a language-general phenomenon rather than a model-specific artifact.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports four psycholinguistic studies on affective meanings of manner-of-motion verbs in temporal expressions, each conducted with human participants and then replicated using ChatGPT o1. The studies cover emergent meaning ratings, valence shifts in literal vs. metaphorical contexts, forced-choice verb selection, and sentence-emoji associations. The authors report strong Spearman correlations between human and AI responses (rho = .73-.96) and conclude that the same conclusions would have been reached if the LLM had replaced the human pool, arguing that LLMs can augment or substitute human experimentation in linguistic research. The central claim is that LLM-based replication can preserve interpretative validity across these four tasks.
Significance. If the central claim were established, the paper would make a useful contribution to a live methodological debate about using LLMs as replacements for human participants in psycholinguistic research. The comparative design across four distinct tasks is a strength, as is the inclusion of full stimulus lists in the appendices. The paper also attempts a promising extension by asking the AI to provide justifications in Study 4. However, the current evidence is largely descriptive, and the lack of prompt transparency and model-selection disclosure makes the substitution claim difficult to evaluate. The paper is likely to interest researchers working on LLM-based experimentation and on the empirical grounding of linguistic semantics.
major comments (4)
- [Sections 2.1.2, 2.2.2, 2.3.2, 2.4.2] The statement that "the prompt presented to the system was exactly the one presented to human subjects" is not verifiable because no prompt text is included or deposited. Human participants used sliders, a "Not apply" option, and emoji displays, which cannot be reproduced verbatim as plain-text prompts. The transformation of the human instrument into an LLM prompt is a consequential design choice, and small wording changes can alter LLM ratings. Please provide the exact prompts for all four studies, along with a description of how the human interface was converted into text and how the "Not apply" option and emoji choices were handled.
- [Section 2.1.2] The sentence "After some trials we settled on ChatGPT o1 for this task" indicates undisclosed model selection. If the model and prompt were chosen after inspecting human data, the reported correlations reflect an exploratory fit rather than a pre-specified replication. Please disclose the full selection history, including models tried, the criteria used, and whether any human data were used to make the selection. Alternatively, pre-register the protocol for future replications.
- [Sections 2.1.2, 2.2.2, 2.4.2] The statistical evidence consists mainly of Spearman's rho values (0.73, 0.96, 0.69) without confidence intervals, p-values, or a measure of human between-subject variance. The claim in Section 2.1.2 that differences are "non-significant" is unsupported because no test is named. A stronger analysis would compare the human data distribution with the AI point estimates, for example using mixed-effects models with participant random effects, or at least report bootstrap intervals around the correlations and an equivalence test for the substitution claim.
- [Section 2.4.2] The post hoc exclusion of emojis "the system did not fully understand" is ad hoc; reporting that the correlation rises from 0.69 to 0.73 after exclusion does not establish the main claim because the exclusion criterion is not pre-specified and could inflate the apparent agreement. Please report results with and without exclusion, justify the criterion, or conduct a sensitivity analysis across different exclusion rules.
minor comments (5)
- [Appendices 3 and 4] The appendix headers are transposed: Appendix 3 is titled "Stimuli for study 4: Fill in the blanks with path and manner verbs" and Appendix 4 is titled "Stimuli for study 3: Sentence-emoji association." The cross-references in the text should be corrected accordingly.
- [Figures 8 and 9] Figure 8 appears twice: the sentence-emoji choice figure and the Spearman correlation figure for Study 4 are both labeled Figure 8. The latter should be renumbered as Figure 9.
- [Abstract and Section 2.1.2] The abstract refers to GPT-4, while the experiments used ChatGPT o1. The same inconsistency appears in the caption of Figure 3, which mentions GPT-4 rather than ChatGPT o1. Please clarify which model versions were used.
- [Section 2.3.1] The participants section for Study 3 reports 32 participants, but the breakdown of 24 females and 9 males sums to 33. Please verify the demographic counts.
- [References] The reference "Authors. 2022. Anonymized" is a placeholder self-citation. If anonymity is required, consider replacing it with a more standard anonymization placeholder or removing it; if the paper is under review, this is acceptable, but it should be flagged as a reminder for the final version.
Circularity Check
No significant circularity: the paper is a comparative empirical study, not a derivation with fitted parameters or self-citation chains.
full rationale
The paper reports four human-LLM comparison studies and reports correlations and percentage agreements. It contains no equation whose output is an input by construction, no parameter fitted to a subset and then renamed a prediction, and no invoked uniqueness theorem from prior author work. The repeated statement that 'the prompt presented to the system was exactly the one presented to human subjects' is a procedural claim that is unsupported by prompt text, but it is not an internal circularity: the paper does not define the LLM output in terms of the human result. Likewise, 'after some trials we settled on ChatGPT o1' raises a legitimate concern about model selection and post-hoc tuning, but the paper does not state that the model was chosen by fitting to the human data, so flagging that as circularity would require speculation about author intent, which the review rules forbid. The conceptual point that an LLM trained on human-produced text cannot independently validate human judgments is a limitation on evidentiary value, not a circular derivation within the paper. Therefore the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Human participants' judgments are the ground truth against which LLM outputs should be measured.
- domain assumption The prompts given to ChatGPT o1 were identical to the instructions given to human participants.
- standard math Spearman rank correlation between mean human ratings and LLM outputs is an appropriate measure of convergence.
Cite this review
Pith. "Pith review of Using AI to replicate human experimental results: a motion study." pith.science (2026). https://pith.science/paper/WK22WG72
@misc{pith2026250710342,
author = {Pith},
title = {Pith review of: Using AI to replicate human experimental results: a motion study},
year = {2026},
howpublished = {\url{https://pith.science/paper/WK22WG72}},
note = {Machine review of arXiv:2507.10342}
}
read the original abstract
This paper explores the potential of large language models (LLMs) as reliable analytical tools in linguistic research, focusing on the emergence of affective meanings in temporal expressions involving manner-of-motion verbs. While LLMs like GPT-4 have shown promise across a range of tasks, their ability to replicate nuanced human judgements remains under scrutiny. We conducted four psycholinguistic studies (on emergent meanings, valence shifts, verb choice in emotional contexts, and sentence-emoji associations) first with human participants and then replicated the same tasks using an LLM. Results across all studies show a striking convergence between human and AI responses, with statistical analyses (e.g., Spearman's rho = .73-.96) indicating strong correlations in both rating patterns and categorical choices. While minor divergences were observed in some cases, these did not alter the overall interpretative outcomes. These findings offer compelling evidence that LLMs can augment traditional human-based experimentation, enabling broader-scale studies without compromising interpretative validity. This convergence not only strengthens the empirical foundation of prior human-based findings but also opens possibilities for hypothesis generation and data expansion through AI. Ultimately, our study supports the use of LLMs as credible and informative collaborators in linguistic inquiry.
Figures
Reference graph
Works this paper leans on
-
[6]
Computational Linguistics 48(1)
Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics 48(1). 207–219. https://doi.org/10.1162/coli_a_00422 Crosthwaite, Peter
-
[8]
Spearman correlation of emoji choices of human participants and AI system In this study, we asked the AI system to justify its answers; the system accordingly provided a brief description of each of the emojis in its justification. In this way, we were able to detect that the system had difficulties understanding two of the emojis; when leaving out the va...
work page 2022
-
[9]
Behavior Research Methods, 4974–4981
The impact of ChatGPT on human data collection: A case study involving typicality norming data. Behavior Research Methods, 4974–4981. https://doi.org/10.3758/s13428-023-02235-w Louwerse, Max
-
[10]
Topics in Cognitive Science 10(3)
Knowing the meaning of a word by the linguistic and perceptual company it keeps. Topics in Cognitive Science 10(3). 573–589. https://doi.org/10.1111/tops.12349 Louwerse, Max
-
[11]
arXiv preprint arXiv:2304.06588
ChatGPT-4 outperforms experts and crowd workers in annotating political Twitter messages with zero-shot learning. arXiv preprint arXiv:2304.06588. Torrent, Thiago, Thomas Hoffmann, Arthur Lorenzi Almeida & Mark Turner
-
[13]
https://doi.org/10.1162/opmi_a_00144
723–738. https://doi.org/10.1162/opmi_a_00144
-
[56]
https://doi.org/10.3758/s13428-024-02337-z Trott, Sean
6082–6100. https://doi.org/10.3758/s13428-024-02337-z Trott, Sean. 2024b. Large language models and the wisdom of small crowds. Open Mind
-
[553]
doi: https://doi.org/10.1038/d41586-018-01023-3 Törnberg, Petter
399–401. doi: https://doi.org/10.1038/d41586-018-01023-3 Törnberg, Petter
Show all 13 references
-
[2018]
2023; Törnberg 2023; Trott 2024b) that have argued in this sense
and joins other works (e.g., Gilardi et al. 2023; Törnberg 2023; Trott 2024b) that have argued in this sense. These results open up a thrilling possibility: what if instead of using 15 sentences as stimuli, we used 200? There are only two logical possibilities: either the resu...
2023
-
[2022]
arXiv preprint arXiv:2208.10264
Using large language models to simulate multiple humans and replicate human subject studies. arXiv preprint arXiv:2208.10264. Alzahrani, Alaa
-
[2023]
arXiv preprint arXiv:2303.15056
ChatGPT outperforms crowd-workers for text-annotation tasks. arXiv preprint arXiv:2303.15056. Heyman, Tom & Geert Heyman
-
[2025]
Heliyon 11(2)
The acceptability and validity of AI-generated psycholinguistic stimuli. Heliyon 11(2). https://doi.org/10.1016/j.heliyon.2025.e42083. Authors
2025 doi
-
[5596]
here” becomes “now
maria-del-rosario.illan-castillo@cnrs.fr 2 Universidad de Murcia jvalen@um.es Abstract: This paper explores the potential of large language models (LLMs) as reliable analytical tools in linguistic research, focusing on the emergence of affective meanings in temporal expression...
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.