Pith. sign in

REVIEW 4 major objections 5 minor 13 references

Using AI to replicate human experimental results: a motion study

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper claims that a large language model given the same tasks as human participants reaches the same interpretive conclusions across four psycholinguistic experiments on time and motion verbs.

desk verdict A useful but under-specified proof-of-concept: the authors show LLM-human convergence on four psycholinguistic tasks, but the missing prompts and model selection details make the central substitution claim untestable as reported. read the letter →

arxiv 2507.10342 v1 pith:WK22WG72 submitted 2025-07-14 cs.CL

classification cs.CL
keywords largelanguagemodelsreplicationpsycholinguisticsmotionaffectivemeaningtimeconceptualizationmannerofverbs
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a large language model can stand in for human participants in psycholinguistic studies of how time is described with manner-of-motion verbs. It reports four experiments—emergent affective meanings, valence shifts in metaphorical contexts, forced-choice verb selection, and sentence-emoji associations—each run first with human volunteers and then replicated with the same instructions given to the LLM. Across all four, the LLM's ratings and choices correlated strongly with human responses (Spearman's $\rho$ from .73 to .96, with exact agreement on forced-choice tasks), and the qualitative conclusions were the same. The authors take this as evidence that LLM responses can augment or replace human data collection for such studies, enabling larger stimulus sets and new hypothesis generation without changing the interpretive outcome.

What carries the argument

The central object is the comparison between two response distributions: human participants recruited through online platforms and the LLM ChatGPT o1 prompted with what the authors state was the exact instruction text given to humans. The statistical machinery is rank-order correlation (Spearman's $\rho$) between human and LLM ratings across items, supplemented by item-by-item difference heatmaps and percentage agreement on categorical choices. The underlying assumption is that if the LLM's relative ordering and qualitative pattern match human ordering on the same prompts, then the LLM captures the interpretive regularities at issue—here, the mapping of motion-verb speed onto affective evaluations of time.

What would settle it

Rerun the four experiments with the authors' exact prompt text on a new set of manner-of-motion verbs and with an open-weight language model; if the Spearman correlations fall below 0.5 or the qualitative valence-polarization pattern reverses, the claim that LLM responses replicate human interpretive outcomes for these tasks would be falsified.

Watch

Extended reading notes

Core claim

The central claim is that, for the restricted domain of affective meanings triggered by manner-of-motion verbs in temporal expressions, the tested LLM (ChatGPT o1) reproduces human experimental outcomes closely enough that the same conclusions would be drawn if only the LLM data were used. In Study 1, ratings of nine semantic dimensions across fifteen sentences correlated with human ratings at $\rho = .73$, with most item-by-item deviations below one point on a seven-point scale. In Study 2, valence ratings in literal versus metaphorical contexts correlated at $\rho = .96$ and showed the same polarizing pattern: fast verbs became more positive and slow verbs more negative in temporal settings. In Study 3, the LLM's forced-choice verb selections matched human choices exactly, and its probability estimates fell within 10–26 percentage points of human choice percentages. In Study 4, emoji choices matched human choices 100% of the time, with a Spearman correlation of .69 that rose to .73 when two emojis the model misread were excluded. The paper concludes that the LLM's answers are fully compatible with human data and that the divergences, where present, do not change the interpretation.

Load-bearing premise

The conclusion depends on the assumption that the prompt given to the LLM was exactly equivalent to the instructions given to human participants, so that any differences in responses come from the model rather than from a different statement of the task.

Editorial extensions

If this is right

  • If the replication holds, psycholinguistic stimulus sets can be scaled from dozens of sentences to hundreds, allowing broader empirical coverage without new human data collection.
  • LLM-generated responses can serve as a hypothesis-generation tool, suggesting new affective meanings that researchers then validate with targeted human studies.
  • The convergence supports a 'convergent evidence' interpretation, strengthening the empirical foundation of the original human-based findings on manner-of-motion verbs.
  • The success of an LLM on these tasks indicates that the affective information is largely encoded in language statistics, not only in embodied experience.
  • Researchers can use LLM probability outputs to approximate the strength of human preferences in forced-choice tasks, as the paper demonstrates with the fill-in-the-blank and emoji studies.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the convergence generalizes beyond the tested verb set, the same method could be used as a cheap pretest before running costly human experiments, flagging stimuli that need human validation.
  • The dependence on prompt wording is untested; a natural extension would be to vary the instruction phrasing and measure how much the LLM-human agreement shifts, since replication results may be sensitive to framing.
  • The authors' speculation about the verb 'race' being skewed by noun frequency could be tested directly with probing methods or frequency-controlled stimuli, providing a concrete next step.
  • A stronger test would be to run the same replication in a different language or with a non-GPT model; convergence there would argue for a language-general phenomenon rather than a model-specific artifact.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports four psycholinguistic studies on affective meanings of manner-of-motion verbs in temporal expressions, each conducted with human participants and then replicated using ChatGPT o1. The studies cover emergent meaning ratings, valence shifts in literal vs. metaphorical contexts, forced-choice verb selection, and sentence-emoji associations. The authors report strong Spearman correlations between human and AI responses (rho = .73-.96) and conclude that the same conclusions would have been reached if the LLM had replaced the human pool, arguing that LLMs can augment or substitute human experimentation in linguistic research. The central claim is that LLM-based replication can preserve interpretative validity across these four tasks.

Significance. If the central claim were established, the paper would make a useful contribution to a live methodological debate about using LLMs as replacements for human participants in psycholinguistic research. The comparative design across four distinct tasks is a strength, as is the inclusion of full stimulus lists in the appendices. The paper also attempts a promising extension by asking the AI to provide justifications in Study 4. However, the current evidence is largely descriptive, and the lack of prompt transparency and model-selection disclosure makes the substitution claim difficult to evaluate. The paper is likely to interest researchers working on LLM-based experimentation and on the empirical grounding of linguistic semantics.

major comments (4)
  1. [Sections 2.1.2, 2.2.2, 2.3.2, 2.4.2] The statement that "the prompt presented to the system was exactly the one presented to human subjects" is not verifiable because no prompt text is included or deposited. Human participants used sliders, a "Not apply" option, and emoji displays, which cannot be reproduced verbatim as plain-text prompts. The transformation of the human instrument into an LLM prompt is a consequential design choice, and small wording changes can alter LLM ratings. Please provide the exact prompts for all four studies, along with a description of how the human interface was converted into text and how the "Not apply" option and emoji choices were handled.
  2. [Section 2.1.2] The sentence "After some trials we settled on ChatGPT o1 for this task" indicates undisclosed model selection. If the model and prompt were chosen after inspecting human data, the reported correlations reflect an exploratory fit rather than a pre-specified replication. Please disclose the full selection history, including models tried, the criteria used, and whether any human data were used to make the selection. Alternatively, pre-register the protocol for future replications.
  3. [Sections 2.1.2, 2.2.2, 2.4.2] The statistical evidence consists mainly of Spearman's rho values (0.73, 0.96, 0.69) without confidence intervals, p-values, or a measure of human between-subject variance. The claim in Section 2.1.2 that differences are "non-significant" is unsupported because no test is named. A stronger analysis would compare the human data distribution with the AI point estimates, for example using mixed-effects models with participant random effects, or at least report bootstrap intervals around the correlations and an equivalence test for the substitution claim.
  4. [Section 2.4.2] The post hoc exclusion of emojis "the system did not fully understand" is ad hoc; reporting that the correlation rises from 0.69 to 0.73 after exclusion does not establish the main claim because the exclusion criterion is not pre-specified and could inflate the apparent agreement. Please report results with and without exclusion, justify the criterion, or conduct a sensitivity analysis across different exclusion rules.
minor comments (5)
  1. [Appendices 3 and 4] The appendix headers are transposed: Appendix 3 is titled "Stimuli for study 4: Fill in the blanks with path and manner verbs" and Appendix 4 is titled "Stimuli for study 3: Sentence-emoji association." The cross-references in the text should be corrected accordingly.
  2. [Figures 8 and 9] Figure 8 appears twice: the sentence-emoji choice figure and the Spearman correlation figure for Study 4 are both labeled Figure 8. The latter should be renumbered as Figure 9.
  3. [Abstract and Section 2.1.2] The abstract refers to GPT-4, while the experiments used ChatGPT o1. The same inconsistency appears in the caption of Figure 3, which mentions GPT-4 rather than ChatGPT o1. Please clarify which model versions were used.
  4. [Section 2.3.1] The participants section for Study 3 reports 32 participants, but the breakdown of 24 females and 9 males sums to 33. Please verify the demographic counts.
  5. [References] The reference "Authors. 2022. Anonymized" is a placeholder self-citation. If anonymity is required, consider replacing it with a more standard anonymization placeholder or removing it; if the paper is under review, this is acceptable, but it should be flagged as a reminder for the final version.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a comparative empirical study, not a derivation with fitted parameters or self-citation chains.

full rationale

The paper reports four human-LLM comparison studies and reports correlations and percentage agreements. It contains no equation whose output is an input by construction, no parameter fitted to a subset and then renamed a prediction, and no invoked uniqueness theorem from prior author work. The repeated statement that 'the prompt presented to the system was exactly the one presented to human subjects' is a procedural claim that is unsupported by prompt text, but it is not an internal circularity: the paper does not define the LLM output in terms of the human result. Likewise, 'after some trials we settled on ChatGPT o1' raises a legitimate concern about model selection and post-hoc tuning, but the paper does not state that the model was chosen by fitting to the human data, so flagging that as circularity would require speculation about author intent, which the review rules forbid. The conceptual point that an LLM trained on human-produced text cannot independently validate human judgments is a limitation on evidentiary value, not a circular derivation within the paper. Therefore the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters are fitted in the paper; the central evidence is descriptive correlations. The main unstated premises are the comparability of human and LLM prompts and the sufficiency of rank correlations as evidence of replication.

assumptions (3)
  • domain assumption Human participants' judgments are the ground truth against which LLM outputs should be measured.
    The entire comparison assumes human responses are the benchmark. Stated in Section 1 and used throughout.
  • domain assumption The prompts given to ChatGPT o1 were identical to the instructions given to human participants.
    Repeated in Sections 2.1.2, 2.2.2, 2.3.2, 2.4.2. The exact prompt text is never included.
  • standard math Spearman rank correlation between mean human ratings and LLM outputs is an appropriate measure of convergence.
    Used in Sections 2.1.2, 2.2.2, 2.4.2 without justification or confidence intervals.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Using AI to replicate human experimental results: a motion study." pith.science (2026). https://pith.science/paper/WK22WG72

@misc{pith2026250710342,
  author       = {Pith},
  title        = {Pith review of: Using AI to replicate human experimental results: a motion study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WK22WG72}},
  note         = {Machine review of arXiv:2507.10342}
}
read the original abstract

This paper explores the potential of large language models (LLMs) as reliable analytical tools in linguistic research, focusing on the emergence of affective meanings in temporal expressions involving manner-of-motion verbs. While LLMs like GPT-4 have shown promise across a range of tasks, their ability to replicate nuanced human judgements remains under scrutiny. We conducted four psycholinguistic studies (on emergent meanings, valence shifts, verb choice in emotional contexts, and sentence-emoji associations) first with human participants and then replicated the same tasks using an LLM. Results across all studies show a striking convergence between human and AI responses, with statistical analyses (e.g., Spearman's rho = .73-.96) indicating strong correlations in both rating patterns and categorical choices. While minor divergences were observed in some cases, these did not alter the overall interpretative outcomes. These findings offer compelling evidence that LLMs can augment traditional human-based experimentation, enabling broader-scale studies without compromising interpretative validity. This convergence not only strengthens the empirical foundation of prior human-based findings but also opens possibilities for hypothesis generation and data expansion through AI. Ultimately, our study supports the use of LLMs as credible and informative collaborators in linguistic inquiry.

Figures

Figures reproduced from arXiv: 2507.10342 by the authors.

Figure 1
Figure 1. Heatmap with differences in responses of human participants and AI system averaged by verb [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

13 extracted references · 7 canonical work pages

  1. [6]

    Computational Linguistics 48(1)

    Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics 48(1). 207–219. https://doi.org/10.1162/coli_a_00422 Crosthwaite, Peter

  2. [8]

    fill-in-the-blanks

    Spearman correlation of emoji choices of human participants and AI system In this study, we asked the AI system to justify its answers; the system accordingly provided a brief description of each of the emojis in its justification. In this way, we were able to detect that the system had difficulties understanding two of the emojis; when leaving out the va...

  3. [9]

    Behavior Research Methods, 4974–4981

    The impact of ChatGPT on human data collection: A case study involving typicality norming data. Behavior Research Methods, 4974–4981. https://doi.org/10.3758/s13428-023-02235-w Louwerse, Max

  4. [10]

    Topics in Cognitive Science 10(3)

    Knowing the meaning of a word by the linguistic and perceptual company it keeps. Topics in Cognitive Science 10(3). 573–589. https://doi.org/10.1111/tops.12349 Louwerse, Max

  5. [11]

    arXiv preprint arXiv:2304.06588

    ChatGPT-4 outperforms experts and crowd workers in annotating political Twitter messages with zero-shot learning. arXiv preprint arXiv:2304.06588. Torrent, Thiago, Thomas Hoffmann, Arthur Lorenzi Almeida & Mark Turner

  6. [13]

    https://doi.org/10.1162/opmi_a_00144

    723–738. https://doi.org/10.1162/opmi_a_00144

  7. [56]

    https://doi.org/10.3758/s13428-024-02337-z Trott, Sean

    6082–6100. https://doi.org/10.3758/s13428-024-02337-z Trott, Sean. 2024b. Large language models and the wisdom of small crowds. Open Mind

  8. [553]

    doi: https://doi.org/10.1038/d41586-018-01023-3 Törnberg, Petter

    399–401. doi: https://doi.org/10.1038/d41586-018-01023-3 Törnberg, Petter

Show all 13 references
  1. [2018]

    2023; Törnberg 2023; Trott 2024b) that have argued in this sense

    and joins other works (e.g., Gilardi et al. 2023; Törnberg 2023; Trott 2024b) that have argued in this sense. These results open up a thrilling possibility: what if instead of using 15 sentences as stimuli, we used 200? There are only two logical possibilities: either the resu...

  2. [2022]

    arXiv preprint arXiv:2208.10264

    Using large language models to simulate multiple humans and replicate human subject studies. arXiv preprint arXiv:2208.10264. Alzahrani, Alaa

  3. [2023]

    arXiv preprint arXiv:2303.15056

    ChatGPT outperforms crowd-workers for text-annotation tasks. arXiv preprint arXiv:2303.15056. Heyman, Tom & Geert Heyman

  4. [2025]

    Heliyon 11(2)

    The acceptability and validity of AI-generated psycholinguistic stimuli. Heliyon 11(2). https://doi.org/10.1016/j.heliyon.2025.e42083. Authors

  5. [5596]

    here” becomes “now

    maria-del-rosario.illan-castillo@cnrs.fr 2 Universidad de Murcia jvalen@um.es Abstract: This paper explores the potential of large language models (LLMs) as reliable analytical tools in linguistic research, focusing on the emergence of affective meanings in temporal expression...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.