{"id":"f18a4272-8425-4b3c-bfa6-ca1a80ff5dc8","arxiv_id":"2502.10266","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"In two replicated linguistics experiments, GPT-4o-mini's zero-shot responses matched or beat published human performance, but the study lacks statistical validation and relies on only two tasks.","lead":"Can large language models replace human crowd workers in linguistics? This paper replicated two human experiments with GPT-4o-mini and found the model's choices often matched or exceeded human performance, while also revealing cases where it over-performs. The work is a pilot study, not proof that LLMs can stand in for human informants.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'LLM-informant' independence assumption (Sec. 3.3) is unestablished: repeated API calls share weights, prompt, and decoding, so the reported 'outperforms humans' means likely reflect a single homogeneous model beating a heterogeneous human average, not a valid synthetic crowd sample.","rationale":"The paper's central claim is that GPT-4o-mini outperforms humans in all tested conditions and can therefore serve as synthetic crowd workers. The reader's weakest_assumption targets the independence of API calls; I agree that this is the most load-bearing assumption because the entire comparison of model means to human means is only meaningful for the synthetic-crowd claim if the runs represent a sample of informants. The claim 'outperforms' is a point-estimate comparison and could be true even for a single model, but the paper's inference to 'LLMs as crowd workers' requires the runs to supply the between-participant variation that human crowds provide. The paper provides no evidence for this and in fact Appendix B shows very few errors across runs, consistent with near-deterministic responses. A direct test of per-item unanimity would settle the matter. I do not raise the separate issue of missing statistical tests as the primary concern because, even with error bars, the independence/representativeness failure would still undermine the central interpretative claim. The paper is honest about its pilot scale and limitations, so the appropriate verdict remains CONDITIONAL (as the reader concluded), not REJECT.","tokens_in":15893,"tokens_out":7429,"duration_ms":74400,"concrete_test":"Using the paper's released code, run the Lombard_21 zero-shot pipeline 68 times (or Cruz_23 34 times) with identical prompts and default API sampling, and record the per-item response across all runs. Compute per-item entropy of the model's choice distribution and the proportion of items on which all runs agree. Compare this to the within-item choice distribution reported for human participants (e.g., a binomial with the human proportion). If most items are unanimous across runs (e.g., >90% agreement) while human within-item distributions are mixed, the 'LLM-informants' are not independent draws from a population, invalidating the synthetic-crowd interpretation and the claim that repeated runs simulate several informants.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 asserts that because each GPT-4o-mini API call is a separate run, 'the answers are unrelated between LLM-informants and contamination is prevented.' This is the sole justification for treating n repeated calls as n synthetic human participants. But statistical independence of draws from a human-like population does not follow from separate API calls: every 'informant' shares the same model weights, prompt template, decoding defaults, and training corpus, so the runs are correlated realizations of one stochastic process (and near-deterministic at low temperature). The headline claim (Section 6) that 'GPT-4o-mini outperforms human participants in both case studies across all conditions tested' is computed as the mean over those runs versus published human means. In a forced-choice task with a majority response, a homogeneous model that always selects the majority option will mechanically beat the human mean—because the human mean is diluted by inter-participant variation—while carrying zero between-informant variance, which is precisely the variation that empirical linguistic studies aim to measure (e.g., Cruz_23 gender congruency, Lombard_21 detection rates). The paper's own discussion of model 'over-performance' (Section 6) and its concession that 'LLMs do not consistently serve as suitable substitutes' (Section 7) undercut the interpretation but do not acknowledge that the independence assumption itself is unsupported. If the assumption fails, the n=34/n=68 'LLM-informants' collapse to effectively one agent; the two cross-validation runs are not independent replications, and the outperformance numbers are not evidence about human-LLM comparability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using OpenAI's GPT-4o-mini as a synthetic crowd worker in empirical linguistics. It replicates two forced-choice studies—Cruz (2023) on gender assignment in Spanish–English code-switching and Lombard et al. (2021) on neologism detection in French—by zero-shot prompting the model with participant profiles and task instructions. Each API call is treated as one 'LLM-informant,' and the model's aggregate response rates (gender congruency and detection accuracy) are compared with the published human means. The paper reports that GPT-4o-mini outperforms human informants in all experimental conditions, then adds two follow-up experiments using role prompting and chain-of-thought prompting to address the model's poor performance on filler items. The authors argue that the proposed pipeline is adaptable, accessible to non-programmers, and points toward a possible interdisciplinary opening between NLP and the humanities.","tokens_in":16122,"tokens_out":5505,"duration_ms":57739,"significance":"If substantiated, the claim that a zero-shot LLM pipeline can reproduce or exceed human performance in linguistically motivated forced-choice tasks would be a valuable proof of concept for synthetic crowdsourcing, particularly for Romance languages that are under-represented in this literature. The paper has concrete strengths: the code is released, the prompts are documented in full in Appendix C, the two case studies go beyond simple labeling tasks, and the authors explicitly discuss the risk of model over-performance. However, the headline result currently rests on descriptive mean comparisons without uncertainty quantification, and the central assumption that separate API calls simulate independent human informants is unexamined. The significance of the result is therefore not yet established; the paper is better read as a promising pilot than as evidence that LLMs can replace human participants.","major_comments":[{"comment":"The central claim that 'GPT-4o-mini outperforms human informants in all experimental conditions tested' is not statistically supported. Figures 3 and 4 plot single mean values for humans and the model with no error bars, confidence intervals, significance tests, or effect sizes, and the text in Sections 4.2 and 5.2 reports only aggregate M-human and M-GPT values. The statement in Section 3.3 that the pipeline 'is repeated twice per replicated study to ensure cross-validation and statistical relevance' does not make n=2 runs a basis for inference. The authors should report human variability from the original studies wherever possible, compute bootstrap or other confidence intervals for the LLM means, and use tests that account for item and participant clustering before making the outperformance claim.","section":"Section 6; Figures 3 and 4"},{"comment":"The load-bearing assumption that separate GPT-4o-mini API calls constitute independent and diverse 'LLM-informants' is unsupported. All runs share the same model weights, prompt template, decoding defaults, and training data, so the repeated calls are correlated realizations of one stochastic process rather than independent draws from a human-like population. If the model is near-deterministic at the default temperature, the variance across synthetic informants will vastly underestimate human population heterogeneity, making the comparison one between a single homogeneous model response and a heterogeneous human average. The paper should either provide evidence of between-call response diversity (temperature settings, response entropy, item-level variance across runs) or reframe the claims as model-level behavior rather than simulated crowds.","section":"Section 3.3"},{"comment":"The chain-of-thought follow-up is presented as demonstrating 'higher alignment to human performance,' but the CoT prompt was constructed after observing the baseline's poor filler accuracy, and its two examples encode the desired response pattern ('non' for a known-word sentence). Because the same test data were used to select and tune this prompt, the reported filler improvement is a post hoc fit rather than an independent evaluation of CoT prompting as a general strategy. This finding should be labeled exploratory, and the authors should either validate CoT on a held-out dataset or report the prompt-selection process explicitly as a limitation.","section":"Section 5.2.2; Appendix C"},{"comment":"The second replication discards reaction time as 'not a reliable parameter' for the model, but reaction time was one of the two dependent measures in Lombard et al. (2021). Reporting only detection accuracy means the LLM pipeline does not actually reproduce the full experimental outcome of the original study, only a subset of it. The paper should state this explicitly as a limitation and narrow the claim of 'replication' accordingly.","section":"Section 5 and Section 5.2"}],"minor_comments":[{"comment":"The heading 'GPT-4-mini as crowd worker' is inconsistent with the model name used throughout the rest of the paper, GPT-4o-mini; please correct it.","section":"Section 3.1 heading"},{"comment":"The word 'overperfom' is a typo for 'overperform'; please proofread the manuscript for such errors.","section":"Section 7"},{"comment":"The French prompt cells contain English translations inside the same table cell (e.g., 'partecipating'), which could be mistaken for part of the actual prompt; the translations should be separated into a distinct column or footnote.","section":"Appendix C"},{"comment":"The numeric labels in Figure 4 are not clearly keyed to the four conditions and two groups in the caption; please add a legend or restructure the figure so readers can map each value to a condition.","section":"Figures 3 and 4"},{"comment":"Several references are incomplete or formatted as author-year citations without full bibliographic entries (e.g., Ahmadabadi et al., Carvalho, Sheehan et al., Kocoń et al.); please complete these entries.","section":"References"},{"comment":"The phrase 'to ensure cross-validation and statistical relevance' is misleading for two repeated runs of the same pipeline; 'cross-validation' is a model-selection procedure, and two runs do not provide statistical relevance. Please rephrase this sentence.","section":"Section 3.3"},{"comment":"The comparison of task duration (72 seconds for the model vs. 25 minutes for humans) should clarify that the model time is API wall-clock time, not a psycholinguistic measure, so that readers do not interpret this as evidence about human-like processing speed.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is best understood as a proof-of-concept with promising, well-documented materials, but the headline claim is overstated relative to the evidence. The major revision should focus on adding uncertainty quantification and addressing the independence assumption; if the authors cannot provide evidence of between-call diversity, they should reframe the contribution as a demonstration of a model-level pipeline rather than a synthetic crowd. I do not see the issues as irreparable, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know two things about this paper. First, it is a genuinely useful pilot: it takes two real empirical linguistics studies, reproduces their forced-choice tasks with zero-shot GPT-4o-mini, and discloses all prompts and code. That is exactly the kind of transparent exploratory work that helps map where LLMs might fit in data collection. Second, the headline claim that the model 'outperforms human informants in all experimental conditions' is not backed by any statistical inference, and the deeper issue is that the paper treats each API call as an independent 'informant' without supporting that assumption.\n\nWhat is new is applying LLM-as-participant to linguistic tasks that go beyond annotation: gender assignment in code-switching and neologism detection. The paper does a good job of describing the pipeline, showing the prompts, and honestly discussing the model's over-performance and its failure on filler items. The Chain-of-Thought follow-up is a sensible attempt to address the filler problem. The limitations section is candid about the small scale.\n\nThe problems are real but concentrated. First, no confidence intervals, error bars, or significance tests anywhere. The reader is asked to accept that means like 0.88 vs 0.55 indicate superiority. Second, and more load-bearing, the independence assumption in Section 3.3. The paper says that because each run is a separate API call, the answers are unrelated and contamination is prevented. But all runs share the same model weights, prompt template, and decoding defaults. They are correlated draws from a narrow stochastic process, not independent samples from a human-like population. A homogeneous model that always picks the majority option will mechanically beat a human mean diluted by inter-individual variation, while carrying none of the variance the original experiments aim to measure. The paper's own discussion of 'over-performance' concedes this but does not connect it back to the independence assumption. The post-hoc exclusion of 'velours' is also a minor red flag.\n\nTo be fair, the stress-test's claim that the runs collapse to 'one agent' is overstated—temperature should produce some spread—but the core objection stands. The abstract and finding list say 'outperforms', while Section 6 says LLMs 'do not consistently serve as suitable substitutes.' That internal tension needs resolving.\n\nBottom line: the paper deserves a serious referee, because the question is timely, the methods transparent, and the flaws fixable. A revision should add inferential statistics, interrogate the independence assumption, and reframe the conclusion as 'comparable in some conditions, with important caveats' rather than 'outperforms.' I would not cite the current results as evidence of human-LLM equivalence, but I would cite it as a transparent pilot and a cautionary example of how not to count API calls as participants.\n\nRecommendation: send it to review, with the expectation of heavy revision.","headline":"A useful, transparent pilot undermined by an unsupported independence assumption and absent inferential statistics; worth referee effort, but the 'outperforms humans' claim needs reframing.","tokens_in":16706,"tokens_out":3181,"would_cite":true,"duration_ms":30956,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that GPT-4o-mini, prompted with zero-shot instructions, can replace human informants in forced-choice linguistic experiments, matching or exceeding human accuracy in two replicated studies.","keywords":["large language models","crowdsourcing","empirical linguistics","zero-shot prompting","chain-of-thought prompting","GPT-4o-mini","gender assignment","neologism detection"],"falsifier":"Repeat the exact zero-shot prompt on the same items many times and compare the response distribution with the between-participant distribution in the original human data; if the model is near-deterministic or its run-to-run variance is far smaller than human variance, the 'LLM-informant' assumption fails and the claimed outperformance is not evidence about human–model equivalence.","tokens_in":15591,"feed_emoji":"🤖","tokens_out":8766,"duration_ms":76192,"temperature":0.7,"pith_summary":"The paper asks whether a large language model can serve as the crowd worker of empirical linguistics, taking over data elicitation normally done by human participants. It replicates two forced-choice tasks—Spanish–English gender assignment in code-switched sentences and French neologism detection—using GPT-4o-mini with zero-shot prompts. The paper's headline finding is that the model outperforms human informants in all experimental conditions tested, completing the tasks in about a minute per informant instead of the 20–50 minutes reported for people. It presents this as evidence that LLM-based synthetic informants are a viable, low-cost, accessible complement or substitute to human crowdsourcing, with prompt design (notably Chain-of-Thought) as a way to steer the model toward human-like performance.","feed_headline":"GPT-4o-mini outperforms human informants in two linguistics tasks","feed_subtitle":"Reproducing two forced-choice experiments with zero-shot prompts suggests LLMs can stand in for human participants.","key_machinery":"The machinery is a replication wrapper around the model: it reads the original stimuli and filler sentences, sends each item as a zero-shot prompt to GPT-4o-mini, treats each separate API run as one 'LLM-informant,' repeats the run as many times as there were human participants in the original study, and scores the synthetic answers against the original human baseline. The system prompt impersonates a participant profile—for instance, a Spanish–English bilingual or a native French speaker—and demands a forced binary choice. Two follow-up modifications isolate the effect of prompt design: a strengthened system-role instruction and Chain-of-Thought prompting with worked examples.","core_discovery":"The paper's central claim is that GPT-4o-mini, queried once per 'LLM-informant' with the same zero-shot prompt, produces responses that align with—and on the main comparison exceed—human performance on two published linguistics experiments. Its stated headline finding is that the model outperforms human informants in all experimental conditions tested. In the gender-assignment study, the model reproduces the human pattern of gender congruency—semantic gender for human-denoting nouns, a masculine default for inanimate nouns, and facilitation from morphological cues—with higher congruency scores than humans. In the neologism-detection study, the model reaches a mean accuracy of 0.985 against 0.91 for humans, although it also shows a yes-bias on filler items; a Chain-of-Thought follow-up raises filler accuracy from 0.77 to 0.99, and the paper notes that in the CoT condition humans slightly outperform the model on one condition (0.92 vs 0.91) while the two scores are close. The paper concludes that the model is a viable basis for synthetic crowdsourcing, while cautioning that over-performance can make it a poor substitute when the research question targets human-specific variation.","pith_inferences":["The paper leaves implicit that its 'LLM-informants' are not independent in a statistical sense: every run shares the same model weights, prompt template, and training data, so run-to-run variance likely underestimates human population variance and the reported outperformance may reflect a single frozen response tendency rather than population-level competence.","Because the model's high accuracy can erase the between-participant differences that linguistic experiments are designed to measure, a more cautious use would be as a pilot or sensitivity-analysis tool—testing whether stimuli are ambiguous—rather than as a full replacement for human participants.","The filler error pattern suggests a general yes-bias in instruction-tuned models; this is testable across other models and tasks, and CoT prompting may be a general corrective whenever negative responses are required.","The replication design could be extended to open-source models and to more Romance languages to separate genuine linguistic generalization from memorized training-data patterns, a comparison the paper names as future work."],"forward_implications":["A zero-shot LLM pipeline can reproduce at least two kinds of forced-choice linguistics experiments, a discourse-completion task and a metalinguistic judgment task, without human participants.","In gender assignment, the model replicates the main human tendencies (semantic gender, masculine default, morphological cues), suggesting synthetic responses could pre-test stimuli or generate hypotheses before human data collection.","Chain-of-Thought prompting corrects the model's yes-bias on filler items, raising filler accuracy from 0.77 to 0.99, so prompt design can tune the model between raw accuracy and human-like performance.","Reaction times recorded from LLM runs are not a meaningful cognitive measure, so replications involving response-time comparisons with humans need a different evaluation method.","The framework is lightweight and reusable by non-programmers, extending the potential use of LLMs beyond NLP annotation to humanities research workflows."],"supporting_citations":[{"why":"Supplies the original Spanish–English code-switching experiment, its stimuli, and the human gender-congruency baseline the replication is compared against.","marker":"[Cruz]"},{"why":"Supplies the original French neologism-detection task, its stimuli, and the human accuracy and reaction-time baseline the replication is judged on.","marker":"[Lombard et al.]"},{"why":"Provides the Chain-of-Thought prompting method used in the second follow-up experiment to improve filler-item performance.","marker":"[Wei et al.]"},{"why":"Evidence that a GPT model with zero-shot learning outperforms experts and crowd workers in annotation, motivating the claim that LLMs can replace human annotators.","marker":"[Törnberg]"},{"why":"Shows ChatGPT outperforming crowd-workers on text-annotation tasks, a cornerstone for transferring the result to linguistic tasks.","marker":"[Gilardi et al.]"},{"why":"Provides the idea that language models can simulate human sub-populations, the theoretical bridge to treating LLMs as synthetic informants.","marker":"[Argyle et al., 2023]"},{"why":"Shows LLMs can replicate existing crowdsourcing pipelines, the direct methodological precedent for this replication design.","marker":"[Wu et al.]"}],"fun_headline_variants":["LLMs beat humans in two linguistics elicitation tasks","GPT-4o-mini superior to human informants in linguistic studies","Can LLMs replace human subjects in linguistics? Yes, for some tasks","AI outperforms humans in forced-choice linguistics experiments","GPT-4o-mini outdoes human participants in two linguistic tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result stands on treating each separate GPT-4o-mini API call as an independent human-like informant; if repeated calls are not independent draws from a participant population, then the reported alignment scores do not measure human–model comparability.","fun_headline_variants_meta":{"raw":{"variants":["LLMs beat humans in two linguistics elicitation tasks","GPT-4o-mini superior to human informants in linguistic studies","Can LLMs replace human subjects in linguistics? Yes, for some tasks","AI outperforms humans in forced-choice linguistics experiments","GPT-4o-mini outdoes human participants in two linguistic tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000326,"raw_usage":{"total_tokens":1877,"prompt_tokens":1050,"completion_tokens":827,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":666,"completion_tokens_details":{"reasoning_tokens":740}},"tokens_in":666,"tokens_out":827,"duration_ms":7496,"temperature":1.0,"reasoning_tokens":740,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T18:42:03.717886+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the exact zero-shot prompt on the same items many times and compare the response distribution with the between-participant distribution in the original human data; if the model is near-deterministic or its run-to-run variance is far smaller than human variance, the 'LLM-informant' assumption fails and the claimed outperformance is not evidence about human–model equivalence.","supporting_citations":[],"review_version":1}