No LLM-prompt pair among 11 models and 4 prompts aligns with average NAEP student performance across math and reading in grades 4, 8, and 12.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can LLMs Reliably Simulate Real Students' Abilities in Mathematics and Reading Comprehension?
No LLM-prompt pair among 11 models and 4 prompts aligns with average NAEP student performance across math and reading in grades 4, 8, and 12.