Across 34 models and 45 configurations, no LLM individually or in ensemble reproduced the lexical diversity or retrieval structure of human phonemic fluency, with the best model producing fewer than half the unique words.
Number of responses Human participants produced an average of 16.89 correct responses within the one-minute time con- straint (SD= 4.84 )
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can LLMs Simulate Human Behavioral Variability? A Case Study in the Phonemic Fluency Task
Across 34 models and 45 configurations, no LLM individually or in ensemble reproduced the lexical diversity or retrieval structure of human phonemic fluency, with the best model producing fewer than half the unique words.