Pith. sign in

REVIEW 5 major objections 9 minor 5 cited by

Embracing AI in Education: Understanding the Surge in Large Language Model Use by Secondary Students

T0 review · 5 major / 9 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A survey of 306 U.S. secondary students finds that 70 percent have used large language models, a higher share than the 43 percent of young adults who have used ChatGPT.

desk verdict Honest but non-representative survey of 306 US secondary students; the 70% LLM-use headline is a sample statistic, not a population estimate. read the letter →

arxiv 2411.18708 v1 pith:NZGFP35D submitted 2024-11-27 cs.HC cs.AI

classification cs.HCcs.AI
keywords largelanguagemodelssecondaryeducationstudentsurveyChatGPTAIindigitaldivideschoolpoliciesmiddle
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that large language models have already become a common homework tool for U.S. middle and high school students. In a survey of 306 students across 43 states, over 70 percent reported using an LLM at least once, and the share stays nearly flat from 7th to 12th grade. The authors contrast this with the 43 percent of 18-to-29-year-olds who have used ChatGPT in a February poll, and they report that usage persists even where schools have restrictive policies. They also document subject-by-subject use, a usage gap between students in top-ranked technology states and other states, and student demand for more accurate and coherent model responses. The practical stakes are that educators and developers should treat LLM use by secondary students as the norm and design curricula and models around it.

What carries the argument

The central object is a 306-respondent survey of 6th through 12th graders recruited through a compensated online panel, a web form shared via tutoring programs and social media, and in-person outreach in California. The survey asks about weekly LLM usage frequency, subjects used, school policy strictness, ethical concern, and open-ended impressions, and the analysis is descriptive: percentages, averages, and simple cross-tabulations rather than statistical modeling. The load-bearing comparison is the paper's 71 percent usage figure measured against the 43 percent young-adult ChatGPT rate from a cited February poll, and the emphasized result is the consistency of usage across grades 7 to 12.

What would settle it

A nationally representative survey of U.S. middle and high school students that asks the same 'ever used an LLM?' question, uses a documented sampling frame and weights, and finds a usage rate below 70 percent or a rate that climbs sharply from 7th to 12th grade would falsify the paper's central claim.

Watch

Extended reading notes

Core claim

The paper's central claim is that LLM adoption among secondary students is not a fringe behavior but a majority behavior: 71 percent of survey respondents said they had used an LLM at least once, with 9 percent using one daily, and the proportion is roughly equal across grades 7 through 12. The authors take this as evidence of a 'surge' that outpaces the 43 percent ChatGPT usage rate reported for young adults, and they find that usage continues despite an average school policy strictness of 3.48 out of 5 and an average student ethical concern of 2.83 out of 5. On demographics, the paper claims an 80.2 percent usage rate in the five highest-ranked technology states versus 64.3 percent in other states, a 76.7 percent rate for private school students versus 71.3 percent for public school students, and disproportionately frequent use among the 17 respondents with paid subscriptions. The authors conclude that restriction has not worked, and they recommend fine-tuning models for educational content, building AI tutors with personalized learning paths, and creating free 'AI classrooms' to equalize access.

Load-bearing premise

The load-bearing assumption is that the 306 students who responded are representative of U.S. middle and high school students generally, even though the sample came from a compensated online panel, a convenience web form, and in-person outreach concentrated in one state with no response rates or weighting described.

Editorial extensions

If this is right

  • If usage is already this common, school policies aimed at outright prohibition are unlikely to change behavior and should shift toward teaching appropriate use.
  • Since usage is consistent across grades 7 to 12, AI literacy and assignment design should address the whole secondary span, not just upper high school.
  • The reported gap between students in top-technology states and other states implies that unequal access to LLMs may widen educational disparities, making free AI tutoring a plausible equity intervention.
  • Students' emphasis on accurate, coherent, complex answers over conversational features suggests that education-focused fine-tuning should prioritize factual reliability over chat ability.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's survey, the flat grade trend hints that LLM adoption starts before middle school and saturates by 7th grade; a longitudinal cohort study would separate this from a simple age effect.
  • Because the survey asked about 'ever used,' the 70 percent figure may overstate routine reliance on LLMs for homework; a diary-based or assignment-log study would establish how often and on what tasks students actually depend on them.
  • A natural next analysis would reweight the sample by state population and control for family income and school type, which would clarify whether the top-tech-state gap is really about technology access or about wealth.
  • The paper's proposed AI classrooms imply that the bottleneck is not student willingness but model quality, content grounding, and infrastructure cost; a pilot comparing a fine-tuned subject tutor against a general chatbot in an under-resourced district would test that.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 9 minor

Summary. The paper reports on a survey of 306 U.S. middle and high school students about their use of large language models (LLMs). Respondents were recruited through a compensated Centiment panel, a Google Form distributed through the authors' tutoring and social-media networks, and in-person outreach in California. The central claims are that 70-71% of respondents have used LLMs at least once, that this rate exceeds the 43% ChatGPT usage reported for 18-29-year-olds in a February 2024 Pew poll, and that usage is roughly constant across grades 7-12. Secondary results describe subject-specific use (writing, math, history, foreign languages), student perceptions of usefulness and hallucination, perceived school-policy strictness and ethical concerns, a 15-percentage-point gap between students in the five top-ranked tech states and the rest, and a smaller private/public school gap. The paper closes with proposals for subject-fine-tuned educational models, AI tutors, and AI classrooms, and a brief limitations paragraph acknowledging internet-access bias.

Significance. The topic is timely, and the paper makes a genuinely useful descriptive contribution if its claims are read at the level of the sample: it is one of the largest published surveys of LLM use by secondary students (n=306 vs. 24-76 in the studies it tabulates), it is transparent about its recruitment channels, and it releases anonymized data to a public repository. There are no fitted parameters or model-derived quantities, so no circularity concern arises; the empirical claims are measured directly from self-reports, and the Pew comparison is an external benchmark rather than a derived quantity. The main significance risk is external validity: because the sample is non-probability, compensated, and mode-mixed, the headline prevalence and grade-invariance findings cannot be extrapolated to U.S. secondary students as the abstract currently implies. If the authors reframe to sample-level claims and quantify uncertainty, the paper would be a solid empirical baseline for educators and developers; as written, the strength of the central claim exceeds what the design supports.

major comments (5)
  1. [Abstract; Section 2; Section 3] The headline claim that '70% of students have utilized LLMs' is phrased as a population-level fact, but the data come entirely from a non-probability convenience sample: a compensated Centiment panel, a Google Form distributed through the authors' own tutoring programs and social media, and in-person outreach in California. The paper reports no response rates, sampling frame, eligibility screening, or post-stratification weights, and the Limitations paragraph itself concedes internet-access bias. As written, the 70% figure is a raw sample statistic, not an estimate of the population parameter. Please reframe all prevalence statements as sample-level claims ('in our sample, 70% of respondents reported...') or add a weighting or bounding analysis that would justify population inference; the abstract's wording overstates what the design can support.
  2. [Figure 1; Table 2] The claim that usage 'remains consistent across 7th to 12th grade' is presented as a substantive finding, but no trend test, equivalence test, or confidence intervals accompany Figure 1, and Table 2 shows grade subsamples ranging from n=17 (7th grade) to n=99 (9th grade), with n=3 in 6th grade. A 70% usage rate in the 7th-grade cell carries a 95% confidence interval of roughly ±22 percentage points, so the observed flatness is statistically indistinguishable from lack of power. The stress-test concern here lands: please report per-grade confidence intervals and a formal trend or equivalence analysis before the invariance claim is made.
  3. [Section 3, technologically advanced regions] The 80.2% vs. 64.3% gap between students in the five top-ranked tech states and students elsewhere is confounded with recruitment mode: California, one of the five states, is the site of the in-person surveys conducted through the first author's school and tutoring connections, while the contrast group depends on the online channels. No adjustment is made for mode of administration or geography. The difference should be described as a between-group difference in a self-selected sample with this confound named, not as evidence of a 'tech savviness' effect.
  4. [Section 3, Pew benchmark; Abstract] The comparison with the 43% young-adult ChatGPT figure [10] is not a controlled benchmark: Pew used a probability-based panel, asked specifically about ChatGPT rather than any LLM, surveyed 18-29-year-olds rather than secondary students, and fielded the poll in February 2024 rather than November 2024. The abstract's claim that student usage is 'higher than the usage percentage among young adults' should therefore be presented as an informal contrast with these caveats, or removed from the abstract.
  5. [Section 2 (Data Collection)] The survey collected self-reports from minors, including 6th- and 7th-grade students who may be younger than 13, yet the manuscript reports no IRB approval, parental or guardian consent procedure, or child assent statement, and it does not say how the 80-cent compensation was delivered or approved for minors. For research involving children, this is a required reporting item; please add the ethics and consent information or state the applicable exemption.
minor comments (9)
  1. [Section 3; Abstract] The usage prevalence is given as 70% in the abstract and 71% in Section 3; please report one consistent value and state the rounding convention.
  2. [Section 3, ethical concerns paragraph] The sentence '2% of the students who rated themselves as a 5 on ethical concern reported using LLMs at least once a week' is ambiguous: the 2% could refer to all respondents or to the subset who rated 5; please state the denominator.
  3. [Section 1] The claim that 'access to chatgpt.com is very easy with no registration requirement' is inaccurate, since ChatGPT requires an account; the intended point about low access barriers can be made without this assertion.
  4. [Figure 4; Section 3] The subject-usage percentages (for example, 28% for math and 57% for writing) do not indicate whether the denominator is all respondents or only LLM users; state the denominator in the caption or text.
  5. [Section 2] The paper states that responses are available in a public GitHub repository but gives no URL; include the link so that the data-sharing claim can be verified.
  6. [Limitations paragraph] The limitations paragraph acknowledges internet-access bias but not non-response bias, mode-of-administration differences, or the heavy California concentration of in-person responses; expanding the paragraph would help readers calibrate the prevalence and comparison claims.
  7. [Table 2; Figure 3] Sixth-grade respondents (n=3) are included in the reported 306 total, while the consistency claim covers grades 7-12; state explicitly whether the 6th-grade responses enter the overall usage figure and Figure 3.
  8. [Section 3, hallucination citation] The claim that ChatGPT 'hallucinates the most when questions are inputted consecutively' is attributed to a single small study [2]; a broader or more rigorous citation, or a softening of the generalization, would be more appropriate.
  9. [Figure 5] The comparison of the policy-strictness average (3.48) with the ethical-concern average (2.83) is described as 'usage outweighs ethical concerns,' but the two numbers measure different constructs (perceived school policy vs. personal ethics) on a single-item scale; the interpretation should be labeled accordingly.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's prevalence and subgroup claims are direct empirical measurements with no fitted parameters or self-referential derivation.

full rationale

The paper makes no model-derived predictions and contains no equations, fitted parameters, or derived quantities that reduce to the survey inputs. The 70% usage figure is a direct tabulation of self-reported responses (Section 3: '71% of the correspondents have used LLMs at least once'), and the grade-level comparison in Figure 1 is a direct cross-tabulation of the same responses. The comparison to 43% of young adults is an external Pew Research benchmark cited from [10], not a number computed from the present survey, so it cannot be circular. Subgroup analyses such as the 80.2% vs. 64.3% tech-state gap and the private/public 76.7% vs. 71.3% gap are also direct subgroup summaries. The only input-dependent step is the paper's own limitations passage, which explicitly concedes that respondents all had internet access and that the study 'may not encompass all viewpoints'; acknowledging this sampling bias is the opposite of hiding a circularity. No self-citations are load-bearing: the authors' prior work is not invoked to justify the empirical claims. The central quantity is measured, not derived, so there is no circularity to report.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters or invented entities appear because the paper makes no model-based derivation. The central claims rest on two domain assumptions about self-report validity and sample representativeness, both only partially acknowledged in the Limitations section.

assumptions (2)
  • domain assumption Survey respondents accurately self-report their LLM use, school policies, and ethical concerns.
    All prevalence and attitude claims rest on self-reported answers to an online form; no validation, response-rate, or attention checks are reported. Section 2 Data Collection.
  • domain assumption The recruited sample represents the target population of US secondary students.
    The headline '70% of students have utilized LLMs' interprets sample responses as informative about students generally; only accessibility bias is noted in Section 4 Limitations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Embracing AI in Education: Understanding the Surge in Large Language Model Use by Secondary Students." pith.science (2026). https://pith.science/paper/NZGFP35D

@misc{pith2026241118708,
  author       = {Pith},
  title        = {Pith review of: Embracing AI in Education: Understanding the Surge in Large Language Model Use by Secondary Students},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NZGFP35D}},
  note         = {Machine review of arXiv:2411.18708}
}
read the original abstract

The impressive essay writing and problem-solving capabilities of large language models (LLMs) like OpenAI's ChatGPT have opened up new avenues in education. Our goal is to gain insights into the widespread use of LLMs among secondary students to inform their future development. Despite school restrictions, our survey of over 300 middle and high school students revealed that a remarkable 70% of students have utilized LLMs, higher than the usage percentage among young adults, and this percentage remains consistent across 7th to 12th grade. Students also reported using LLMs for multiple subjects, including language arts, history, and math assignments, but expressed mixed thoughts on their effectiveness due to occasional hallucinations in historical contexts and incorrect answers for lack of rigorous reasoning. The survey feedback called for LLMs better adapted for students, and also raised questions to developers and educators on how to help students from underserved communities leverage LLMs' capabilities for equal access to advanced education resources. We propose a few ideas to address such issues, including subject-specific models, personalized learning, and AI classrooms.

Figures

Figures reproduced from arXiv: 2411.18708 by the authors.

Figure 1
Figure 1. LLM usage remained consistent among 7th-12th graders. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. We received responses from students in 43 states. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Over 70% of students have used LLMs. LLMs are used in all school subjects. Previous research reveals that LLMs exhibit inferior mathe￾matical abilities compared to addressing open-ended requests in English [7]. Despite this, our survey shows 28% students still used LLMs for the math subject. They also use LLMs in various other subjects, including foreign languages and history classes ( [PITH_FULL_IMAGE:figures/full… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: LLMs are used in all school subjects. Students share various thoughts regarding LLM impact. Responses to an open-ended question regarding the impact of LLMs on students’ academic performance were diverse, despite the common usage. The following are a few student respon…
Figure 5
Figure 5. Figure 5: figure 5. Interestingly, we found that 2% of the students who rated themselves as a 5 on ethical concern [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multilingual Performance Biases of Large Language Models in Education

    cs.CL 2025-04 conditional novelty 6.0 of 10

    LLMs are less reliable at tutoring, feedback, and misconception detection in lower-resource languages, and English prompts usually work as well as or better than prompts in the target language.

  2. People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Expert annotators who frequently use LLMs for writing detected AI-generated non-fiction articles with 99.3% accuracy by majority vote, outperforming all but one commercial detector.

  3. When LLM Essays Outscore Student Essays: What a Korean Writing Rubric Rewards and Where Readers Disagree

    cs.CL 2026-01 conditional novelty 5.0 of 10

    Rubric-guided calibration lifted three Korean-language experts' majority-vote accuracy at detecting LLM essays from 60% to 90% on disjoint 30-essay sets (10/10 on a 10-essay subset), with agreement rising from Fleiss ...

  4. Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A 14B RL-trained model reports 96.2% on an internal Chinese K-12 benchmark, about 15x lower serving cost than DeepSeek-R1, with three new training tricks.

  5. Large Language Models Will Change The Way Children Think About Technology And Impact Every Interaction Paradigm

    cs.HC 2025-04 conditional novelty 3.0 of 10

    Children's early experiences with LLMs will shift their expectations for all technology, moving interaction design from commands and menus toward conversational, context-aware systems.

Reference graph

Works this paper leans on

21 extracted references · 16 canonical work pages · cited by 5 Pith papers

  1. [10]

    Americans’ use of chatgpt is ticking up, but few trust its election information

    Colleen McClain. Americans’ use of chatgpt is ticking up, but few trust its election information. https: //pewrsr.ch/3VyXbJR, Mar 2024

  2. [1]

    The opportunities and challenges of chatgpt in education

    Ibrahim Adeshola and Adeola Praise Adepoju. The opportunities and challenges of chatgpt in education. Interactive Learning Environments, pages 1–14, 2023

  3. [2]

    Hallucinations in chatgpt: An unreliable tool for learning

    Zakia Ahmad, Wahid Kaiser, and Sifatur Rahim. Hallucinations in chatgpt: An unreliable tool for learning. Rupkatha Journal on Interdisciplinary Studies in Humanities, 15, 12 2023

  4. [3]

    Education in the era of generative artificial intelligence (ai): Understanding the potential benefits of chatgpt in promoting teaching and learning

    David Baidoo-anu and Leticia Owusu Ansah. Education in the era of generative artificial intelligence (ai): Understanding the potential benefits of chatgpt in promoting teaching and learning. Journal of AI, 7:52–62, 2023

  5. [4]

    Iris: An ai-driven virtual tutor for computer science education

    Patrick Bassner, Eduard Frankford, and Stephan Krusche. Iris: An ai-driven virtual tutor for computer science education. Innovation and Technology in Computer Science Education, 1, May 2024

  6. [5]

    Testing, socializing, exploring: Characterizing middle schoolers’ approaches to and conceptions of chatgpt

    Yasmine Belghith, Atefeh Mahdavi Goloujeh, Brian Magerko, Duri Long, Tom Mcklin, and Jessica Roberts. Testing, socializing, exploring: Characterizing middle schoolers’ approaches to and conceptions of chatgpt. In Proceedings of the CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY , USA, 2024. Association for Computing Machinery

  7. [6]

    Survey with centiment

    Centiment. Survey with centiment. https://www.centiment.co/

  8. [7]

    Mathematical capabilities of chatgpt

    Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Petersen, and Julius Berner. Mathematical capabilities of chatgpt. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems , volume 36, pages 27699–27744. Curran Associates, Inc., 2023

Show all 21 references
  1. [8]

    Chatgpt for good? on opportunities and challenges of large language models for education

    Enkelejda Kasneci, Kathrin Sessler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Sail...

  2. [9]

    Research on motivation and behavior of chatgpt use in middle school students

    Yalin Li, Jinyuan Chen, Haowen Zhou, Haoyang Yuan, and Ruolin Yang. Research on motivation and behavior of chatgpt use in middle school students. International Journal of New Developments in Education, 5, 2023

  3. [11]

    State technology and science index

    Milken Institute. State technology and science index. https://statetechandscience.org, 2022

  4. [12]

    Terms of use

    OpenAI. Terms of use. https://openai.com/policies/terms-of-use/, 2023

  5. [13]

    OpenAI. Chatgpt. https://chat.openai.com, 2024

  6. [14]

    Openai for education

    OpenAI. Openai for education. https://openai.com/index/introducing-chatgpt-edu/, 2024

  7. [15]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...

  8. [16]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...

  9. [17]

    Mostafizer Rahman and Yutaka Watanobe

    Md. Mostafizer Rahman and Yutaka Watanobe. Chatgpt for education and research: Opportunities, threats, and strategies. Applied Sciences, 13(9), 2023

  10. [18]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

    Gemini Team, Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry, Lepikhin, Timothy Lillicrap, Jean baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew Dai, Katie Millican, Ethan D...

  11. [19]

    Improving student learning with hybrid human-ai tutoring: A three-study quasi-experimental investigation

    Danielle R Thomas, Jionghao Lin, Erin Gatz, Ashish Gurung, Shivang Gupta, Kole Norberg, Stephen E Fancsali, Vincent Aleven, Lee Branstetter, Emma Brunskill, and Kenneth R Koedinger. Improving student learning with hybrid human-ai tutoring: A three-study quasi-experimental inve...

  12. [20]

    Llama: Open and efficient foundation language models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language...

  13. [21]

    Yu, and Qingsong Wen

    Shen Wang, Tianlong Xu, Hang Li, Chaoli Zhang, Joleen Liang, Jiliang Tang, Philip S. Yu, and Qingsong Wen. Large language models for education: A survey and outlook. https://doi.org/10.48550/arXiv. 2403.18105, 2024. 8

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.