REVIEW 5 major objections 9 minor 5 cited by
Embracing AI in Education: Understanding the Surge in Large Language Model Use by Secondary Students
T0 review · 5 major / 9 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A survey of 306 U.S. secondary students finds that 70 percent have used large language models, a higher share than the 43 percent of young adults who have used ChatGPT.
desk verdict Honest but non-representative survey of 306 US secondary students; the 70% LLM-use headline is a sample statistic, not a population estimate. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a 306-respondent survey of 6th through 12th graders recruited through a compensated online panel, a web form shared via tutoring programs and social media, and in-person outreach in California. The survey asks about weekly LLM usage frequency, subjects used, school policy strictness, ethical concern, and open-ended impressions, and the analysis is descriptive: percentages, averages, and simple cross-tabulations rather than statistical modeling. The load-bearing comparison is the paper's 71 percent usage figure measured against the 43 percent young-adult ChatGPT rate from a cited February poll, and the emphasized result is the consistency of usage across grades 7 to 12.
What would settle it
A nationally representative survey of U.S. middle and high school students that asks the same 'ever used an LLM?' question, uses a documented sampling frame and weights, and finds a usage rate below 70 percent or a rate that climbs sharply from 7th to 12th grade would falsify the paper's central claim.
Extended reading notes
Core claim
The paper's central claim is that LLM adoption among secondary students is not a fringe behavior but a majority behavior: 71 percent of survey respondents said they had used an LLM at least once, with 9 percent using one daily, and the proportion is roughly equal across grades 7 through 12. The authors take this as evidence of a 'surge' that outpaces the 43 percent ChatGPT usage rate reported for young adults, and they find that usage continues despite an average school policy strictness of 3.48 out of 5 and an average student ethical concern of 2.83 out of 5. On demographics, the paper claims an 80.2 percent usage rate in the five highest-ranked technology states versus 64.3 percent in other states, a 76.7 percent rate for private school students versus 71.3 percent for public school students, and disproportionately frequent use among the 17 respondents with paid subscriptions. The authors conclude that restriction has not worked, and they recommend fine-tuning models for educational content, building AI tutors with personalized learning paths, and creating free 'AI classrooms' to equalize access.
Load-bearing premise
The load-bearing assumption is that the 306 students who responded are representative of U.S. middle and high school students generally, even though the sample came from a compensated online panel, a convenience web form, and in-person outreach concentrated in one state with no response rates or weighting described.
Editorial extensions
If this is right
- If usage is already this common, school policies aimed at outright prohibition are unlikely to change behavior and should shift toward teaching appropriate use.
- Since usage is consistent across grades 7 to 12, AI literacy and assignment design should address the whole secondary span, not just upper high school.
- The reported gap between students in top-technology states and other states implies that unequal access to LLMs may widen educational disparities, making free AI tutoring a plausible equity intervention.
- Students' emphasis on accurate, coherent, complex answers over conversational features suggests that education-focused fine-tuning should prioritize factual reliability over chat ability.
Reading between the lines
- Beyond the paper's survey, the flat grade trend hints that LLM adoption starts before middle school and saturates by 7th grade; a longitudinal cohort study would separate this from a simple age effect.
- Because the survey asked about 'ever used,' the 70 percent figure may overstate routine reliance on LLMs for homework; a diary-based or assignment-log study would establish how often and on what tasks students actually depend on them.
- A natural next analysis would reweight the sample by state population and control for family income and school type, which would clarify whether the top-tech-state gap is really about technology access or about wealth.
- The paper's proposed AI classrooms imply that the bottleneck is not student willingness but model quality, content grounding, and infrastructure cost; a pilot comparing a fine-tuned subject tutor against a general chatbot in an under-resourced district would test that.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports on a survey of 306 U.S. middle and high school students about their use of large language models (LLMs). Respondents were recruited through a compensated Centiment panel, a Google Form distributed through the authors' tutoring and social-media networks, and in-person outreach in California. The central claims are that 70-71% of respondents have used LLMs at least once, that this rate exceeds the 43% ChatGPT usage reported for 18-29-year-olds in a February 2024 Pew poll, and that usage is roughly constant across grades 7-12. Secondary results describe subject-specific use (writing, math, history, foreign languages), student perceptions of usefulness and hallucination, perceived school-policy strictness and ethical concerns, a 15-percentage-point gap between students in the five top-ranked tech states and the rest, and a smaller private/public school gap. The paper closes with proposals for subject-fine-tuned educational models, AI tutors, and AI classrooms, and a brief limitations paragraph acknowledging internet-access bias.
Significance. The topic is timely, and the paper makes a genuinely useful descriptive contribution if its claims are read at the level of the sample: it is one of the largest published surveys of LLM use by secondary students (n=306 vs. 24-76 in the studies it tabulates), it is transparent about its recruitment channels, and it releases anonymized data to a public repository. There are no fitted parameters or model-derived quantities, so no circularity concern arises; the empirical claims are measured directly from self-reports, and the Pew comparison is an external benchmark rather than a derived quantity. The main significance risk is external validity: because the sample is non-probability, compensated, and mode-mixed, the headline prevalence and grade-invariance findings cannot be extrapolated to U.S. secondary students as the abstract currently implies. If the authors reframe to sample-level claims and quantify uncertainty, the paper would be a solid empirical baseline for educators and developers; as written, the strength of the central claim exceeds what the design supports.
major comments (5)
- [Abstract; Section 2; Section 3] The headline claim that '70% of students have utilized LLMs' is phrased as a population-level fact, but the data come entirely from a non-probability convenience sample: a compensated Centiment panel, a Google Form distributed through the authors' own tutoring programs and social media, and in-person outreach in California. The paper reports no response rates, sampling frame, eligibility screening, or post-stratification weights, and the Limitations paragraph itself concedes internet-access bias. As written, the 70% figure is a raw sample statistic, not an estimate of the population parameter. Please reframe all prevalence statements as sample-level claims ('in our sample, 70% of respondents reported...') or add a weighting or bounding analysis that would justify population inference; the abstract's wording overstates what the design can support.
- [Figure 1; Table 2] The claim that usage 'remains consistent across 7th to 12th grade' is presented as a substantive finding, but no trend test, equivalence test, or confidence intervals accompany Figure 1, and Table 2 shows grade subsamples ranging from n=17 (7th grade) to n=99 (9th grade), with n=3 in 6th grade. A 70% usage rate in the 7th-grade cell carries a 95% confidence interval of roughly ±22 percentage points, so the observed flatness is statistically indistinguishable from lack of power. The stress-test concern here lands: please report per-grade confidence intervals and a formal trend or equivalence analysis before the invariance claim is made.
- [Section 3, technologically advanced regions] The 80.2% vs. 64.3% gap between students in the five top-ranked tech states and students elsewhere is confounded with recruitment mode: California, one of the five states, is the site of the in-person surveys conducted through the first author's school and tutoring connections, while the contrast group depends on the online channels. No adjustment is made for mode of administration or geography. The difference should be described as a between-group difference in a self-selected sample with this confound named, not as evidence of a 'tech savviness' effect.
- [Section 3, Pew benchmark; Abstract] The comparison with the 43% young-adult ChatGPT figure [10] is not a controlled benchmark: Pew used a probability-based panel, asked specifically about ChatGPT rather than any LLM, surveyed 18-29-year-olds rather than secondary students, and fielded the poll in February 2024 rather than November 2024. The abstract's claim that student usage is 'higher than the usage percentage among young adults' should therefore be presented as an informal contrast with these caveats, or removed from the abstract.
- [Section 2 (Data Collection)] The survey collected self-reports from minors, including 6th- and 7th-grade students who may be younger than 13, yet the manuscript reports no IRB approval, parental or guardian consent procedure, or child assent statement, and it does not say how the 80-cent compensation was delivered or approved for minors. For research involving children, this is a required reporting item; please add the ethics and consent information or state the applicable exemption.
minor comments (9)
- [Section 3; Abstract] The usage prevalence is given as 70% in the abstract and 71% in Section 3; please report one consistent value and state the rounding convention.
- [Section 3, ethical concerns paragraph] The sentence '2% of the students who rated themselves as a 5 on ethical concern reported using LLMs at least once a week' is ambiguous: the 2% could refer to all respondents or to the subset who rated 5; please state the denominator.
- [Section 1] The claim that 'access to chatgpt.com is very easy with no registration requirement' is inaccurate, since ChatGPT requires an account; the intended point about low access barriers can be made without this assertion.
- [Figure 4; Section 3] The subject-usage percentages (for example, 28% for math and 57% for writing) do not indicate whether the denominator is all respondents or only LLM users; state the denominator in the caption or text.
- [Section 2] The paper states that responses are available in a public GitHub repository but gives no URL; include the link so that the data-sharing claim can be verified.
- [Limitations paragraph] The limitations paragraph acknowledges internet-access bias but not non-response bias, mode-of-administration differences, or the heavy California concentration of in-person responses; expanding the paragraph would help readers calibrate the prevalence and comparison claims.
- [Table 2; Figure 3] Sixth-grade respondents (n=3) are included in the reported 306 total, while the consistency claim covers grades 7-12; state explicitly whether the 6th-grade responses enter the overall usage figure and Figure 3.
- [Section 3, hallucination citation] The claim that ChatGPT 'hallucinates the most when questions are inputted consecutively' is attributed to a single small study [2]; a broader or more rigorous citation, or a softening of the generalization, would be more appropriate.
- [Figure 5] The comparison of the policy-strictness average (3.48) with the ethical-concern average (2.83) is described as 'usage outweighs ethical concerns,' but the two numbers measure different constructs (perceived school policy vs. personal ethics) on a single-item scale; the interpretation should be labeled accordingly.
Circularity Check
No significant circularity: the survey's prevalence and subgroup claims are direct empirical measurements with no fitted parameters or self-referential derivation.
full rationale
The paper makes no model-derived predictions and contains no equations, fitted parameters, or derived quantities that reduce to the survey inputs. The 70% usage figure is a direct tabulation of self-reported responses (Section 3: '71% of the correspondents have used LLMs at least once'), and the grade-level comparison in Figure 1 is a direct cross-tabulation of the same responses. The comparison to 43% of young adults is an external Pew Research benchmark cited from [10], not a number computed from the present survey, so it cannot be circular. Subgroup analyses such as the 80.2% vs. 64.3% tech-state gap and the private/public 76.7% vs. 71.3% gap are also direct subgroup summaries. The only input-dependent step is the paper's own limitations passage, which explicitly concedes that respondents all had internet access and that the study 'may not encompass all viewpoints'; acknowledging this sampling bias is the opposite of hiding a circularity. No self-citations are load-bearing: the authors' prior work is not invoked to justify the empirical claims. The central quantity is measured, not derived, so there is no circularity to report.
Assumptions & free parameters
assumptions (2)
- domain assumption Survey respondents accurately self-report their LLM use, school policies, and ethical concerns.
- domain assumption The recruited sample represents the target population of US secondary students.
Cite this review
Pith. "Pith review of Embracing AI in Education: Understanding the Surge in Large Language Model Use by Secondary Students." pith.science (2026). https://pith.science/paper/NZGFP35D
@misc{pith2026241118708,
author = {Pith},
title = {Pith review of: Embracing AI in Education: Understanding the Surge in Large Language Model Use by Secondary Students},
year = {2026},
howpublished = {\url{https://pith.science/paper/NZGFP35D}},
note = {Machine review of arXiv:2411.18708}
}
read the original abstract
The impressive essay writing and problem-solving capabilities of large language models (LLMs) like OpenAI's ChatGPT have opened up new avenues in education. Our goal is to gain insights into the widespread use of LLMs among secondary students to inform their future development. Despite school restrictions, our survey of over 300 middle and high school students revealed that a remarkable 70% of students have utilized LLMs, higher than the usage percentage among young adults, and this percentage remains consistent across 7th to 12th grade. Students also reported using LLMs for multiple subjects, including language arts, history, and math assignments, but expressed mixed thoughts on their effectiveness due to occasional hallucinations in historical contexts and incorrect answers for lack of rigorous reasoning. The survey feedback called for LLMs better adapted for students, and also raised questions to developers and educators on how to help students from underserved communities leverage LLMs' capabilities for equal access to advanced education resources. We propose a few ideas to address such issues, including subject-specific models, personalized learning, and AI classrooms.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 5 Pith papers
-
Multilingual Performance Biases of Large Language Models in Education
LLMs are less reliable at tutoring, feedback, and misconception detection in lower-resource languages, and English prompts usually work as well as or better than prompts in the target language.
-
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
Expert annotators who frequently use LLMs for writing detected AI-generated non-fiction articles with 99.3% accuracy by majority vote, outperforming all but one commercial detector.
-
When LLM Essays Outscore Student Essays: What a Korean Writing Rubric Rewards and Where Readers Disagree
Rubric-guided calibration lifted three Korean-language experts' majority-vote accuracy at detecting LLM essays from 60% to 90% on disjoint 30-essay sets (10/10 on a 10-essay subset), with agreement rising from Fleiss ...
-
Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning
A 14B RL-trained model reports 96.2% on an internal Chinese K-12 benchmark, about 15x lower serving cost than DeepSeek-R1, with three new training tricks.
-
Large Language Models Will Change The Way Children Think About Technology And Impact Every Interaction Paradigm
Children's early experiences with LLMs will shift their expectations for all technology, moving interaction design from commands and menus toward conversational, context-aware systems.
Reference graph
Works this paper leans on
-
[10]
Americans’ use of chatgpt is ticking up, but few trust its election information
Colleen McClain. Americans’ use of chatgpt is ticking up, but few trust its election information. https: //pewrsr.ch/3VyXbJR, Mar 2024
work page 2024
-
[1]
The opportunities and challenges of chatgpt in education
Ibrahim Adeshola and Adeola Praise Adepoju. The opportunities and challenges of chatgpt in education. Interactive Learning Environments, pages 1–14, 2023
work page 2023
-
[2]
Hallucinations in chatgpt: An unreliable tool for learning
Zakia Ahmad, Wahid Kaiser, and Sifatur Rahim. Hallucinations in chatgpt: An unreliable tool for learning. Rupkatha Journal on Interdisciplinary Studies in Humanities, 15, 12 2023
work page 2023
-
[3]
David Baidoo-anu and Leticia Owusu Ansah. Education in the era of generative artificial intelligence (ai): Understanding the potential benefits of chatgpt in promoting teaching and learning. Journal of AI, 7:52–62, 2023
work page 2023
-
[4]
Iris: An ai-driven virtual tutor for computer science education
Patrick Bassner, Eduard Frankford, and Stephan Krusche. Iris: An ai-driven virtual tutor for computer science education. Innovation and Technology in Computer Science Education, 1, May 2024
work page 2024
-
[5]
Yasmine Belghith, Atefeh Mahdavi Goloujeh, Brian Magerko, Duri Long, Tom Mcklin, and Jessica Roberts. Testing, socializing, exploring: Characterizing middle schoolers’ approaches to and conceptions of chatgpt. In Proceedings of the CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY , USA, 2024. Association for Computing Machinery
work page 2024
- [6]
-
[7]
Mathematical capabilities of chatgpt
Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Petersen, and Julius Berner. Mathematical capabilities of chatgpt. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems , volume 36, pages 27699–27744. Curran Associates, Inc., 2023
work page 2023
Show all 21 references
-
[8]
Chatgpt for good? on opportunities and challenges of large language models for education
Enkelejda Kasneci, Kathrin Sessler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Sail...
2023
-
[9]
Research on motivation and behavior of chatgpt use in middle school students
Yalin Li, Jinyuan Chen, Haowen Zhou, Haoyang Yuan, and Ruolin Yang. Research on motivation and behavior of chatgpt use in middle school students. International Journal of New Developments in Education, 5, 2023
2023
-
[11]
State technology and science index
Milken Institute. State technology and science index. https://statetechandscience.org, 2022
2022
-
[12]
Terms of use
OpenAI. Terms of use. https://openai.com/policies/terms-of-use/, 2023
2023
-
[13]
OpenAI. Chatgpt. https://chat.openai.com, 2024
2024
-
[14]
Openai for education
OpenAI. Openai for education. https://openai.com/index/introducing-chatgpt-edu/, 2024
2024
- [15]
-
[16]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
-
[17]
Mostafizer Rahman and Yutaka Watanobe
Md. Mostafizer Rahman and Yutaka Watanobe. Chatgpt for education and research: Opportunities, threats, and strategies. Applied Sciences, 13(9), 2023
2023
-
[18]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry, Lepikhin, Timothy Lillicrap, Jean baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew Dai, Katie Millican, Ethan D...
-
[19]
Improving student learning with hybrid human-ai tutoring: A three-study quasi-experimental investigation
Danielle R Thomas, Jionghao Lin, Erin Gatz, Ashish Gurung, Shivang Gupta, Kole Norberg, Stephen E Fancsali, Vincent Aleven, Lee Branstetter, Emma Brunskill, and Kenneth R Koedinger. Improving student learning with hybrid human-ai tutoring: A three-study quasi-experimental inve...
2024
-
[20]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language...
- [21]
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.