{"id":"4863c7a4-93e3-4e4d-a3a2-b6eeceb2f0e0","arxiv_id":"2411.18708","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 306-student US survey found over 70% of respondents reported using LLMs like ChatGPT for schoolwork, with use reported in every grade and subject area.","lead":"A survey of 306 US middle and high school students reports that over 70% have used large language models like ChatGPT, despite many school restrictions. The finding matters because it suggests secondary students already rely on AI for schoolwork and that access differences by state, school type, and paid subscriptions may widen educational gaps.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 70% prevalence claim is not established for the population because the convenience sample lacks a sampling frame, response rates, and weighting; the grade-invariance sub-claim is also underpowered.","rationale":"The reader's weakest-assumption identification is correct and is the most load-bearing: every headline number in the abstract is a proportion computed from a sample whose selection mechanism is unknown and likely skewed toward tech-savvy, affluent, or already-engaged students. The internal descriptive findings (e.g., subject-use patterns, policy perceptions) are still informative and the authors are transparent about accessibility bias, which is why I do not recommend REJECT. However, the paper cannot claim a population-level '70% of students' or 'higher than young adults' without either a probability sample or a weighting/benchmarking exercise. The grade-invariance claim also needs error bars and a trend test. The reader's CONDITIONAL verdict with requests for recruitment details, margins of error, and data availability is appropriate; my check sharpens the condition: reweight and see if the estimate survives. I therefore recommend UNCHANGED.","tokens_in":5495,"tokens_out":4665,"duration_ms":42362,"concrete_test":"Obtain the promised de-identified raw responses from the GitHub repository and reweight the sample to match US secondary-school population margins for grade, state, and school type (using NCES Common Core of Data). Recompute the overall LLM-usage proportion with these post-stratification weights. If the weighted estimate differs from 70% by more than 10 percentage points, the headline is a sample artifact rather than a population fact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that 70% of secondary students have used LLMs, with usage consistent across grades 7–12—rests entirely on a non-probability convenience sample described in Section 2. Participants were recruited through a compensated Centiment panel, a Google Form distributed via the authors' tutoring and social-media networks, and in-person outreach in California (where the first author's school is located). No response rates, sampling frame, eligibility screening, or post-stratification weights are reported, and the paper's own limitations section concedes internet-access bias. The grade-invariance sub-claim is even less secure: the 7th-grade subsample has n=17 (Table 2), so a 70% usage rate carries a 95% confidence interval of roughly ±22 percentage points; no significance or trend test is presented for Figure 1. The comparison to 43% of young adults (Pew, Feb 2024) is not a controlled benchmark: it uses a different survey mode, a different question, and a different time period. Thus the headline prevalence is a raw sample statistic, not an estimate of the population parameter, and the grade-consistency claim is indistinguishable from a lack of statistical power.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on a survey of 306 U.S. middle and high school students about their use of large language models (LLMs). Respondents were recruited through a compensated Centiment panel, a Google Form distributed through the authors' tutoring and social-media networks, and in-person outreach in California. The central claims are that 70-71% of respondents have used LLMs at least once, that this rate exceeds the 43% ChatGPT usage reported for 18-29-year-olds in a February 2024 Pew poll, and that usage is roughly constant across grades 7-12. Secondary results describe subject-specific use (writing, math, history, foreign languages), student perceptions of usefulness and hallucination, perceived school-policy strictness and ethical concerns, a 15-percentage-point gap between students in the five top-ranked tech states and the rest, and a smaller private/public school gap. The paper closes with proposals for subject-fine-tuned educational models, AI tutors, and AI classrooms, and a brief limitations paragraph acknowledging internet-access bias.","tokens_in":5666,"tokens_out":9934,"duration_ms":79681,"significance":"The topic is timely, and the paper makes a genuinely useful descriptive contribution if its claims are read at the level of the sample: it is one of the largest published surveys of LLM use by secondary students (n=306 vs. 24-76 in the studies it tabulates), it is transparent about its recruitment channels, and it releases anonymized data to a public repository. There are no fitted parameters or model-derived quantities, so no circularity concern arises; the empirical claims are measured directly from self-reports, and the Pew comparison is an external benchmark rather than a derived quantity. The main significance risk is external validity: because the sample is non-probability, compensated, and mode-mixed, the headline prevalence and grade-invariance findings cannot be extrapolated to U.S. secondary students as the abstract currently implies. If the authors reframe to sample-level claims and quantify uncertainty, the paper would be a solid empirical baseline for educators and developers; as written, the strength of the central claim exceeds what the design supports.","major_comments":[{"comment":"The headline claim that '70% of students have utilized LLMs' is phrased as a population-level fact, but the data come entirely from a non-probability convenience sample: a compensated Centiment panel, a Google Form distributed through the authors' own tutoring programs and social media, and in-person outreach in California. The paper reports no response rates, sampling frame, eligibility screening, or post-stratification weights, and the Limitations paragraph itself concedes internet-access bias. As written, the 70% figure is a raw sample statistic, not an estimate of the population parameter. Please reframe all prevalence statements as sample-level claims ('in our sample, 70% of respondents reported...') or add a weighting or bounding analysis that would justify population inference; the abstract's wording overstates what the design can support.","section":"Abstract; Section 2; Section 3"},{"comment":"The claim that usage 'remains consistent across 7th to 12th grade' is presented as a substantive finding, but no trend test, equivalence test, or confidence intervals accompany Figure 1, and Table 2 shows grade subsamples ranging from n=17 (7th grade) to n=99 (9th grade), with n=3 in 6th grade. A 70% usage rate in the 7th-grade cell carries a 95% confidence interval of roughly ±22 percentage points, so the observed flatness is statistically indistinguishable from lack of power. The stress-test concern here lands: please report per-grade confidence intervals and a formal trend or equivalence analysis before the invariance claim is made.","section":"Figure 1; Table 2"},{"comment":"The 80.2% vs. 64.3% gap between students in the five top-ranked tech states and students elsewhere is confounded with recruitment mode: California, one of the five states, is the site of the in-person surveys conducted through the first author's school and tutoring connections, while the contrast group depends on the online channels. No adjustment is made for mode of administration or geography. The difference should be described as a between-group difference in a self-selected sample with this confound named, not as evidence of a 'tech savviness' effect.","section":"Section 3, technologically advanced regions"},{"comment":"The comparison with the 43% young-adult ChatGPT figure [10] is not a controlled benchmark: Pew used a probability-based panel, asked specifically about ChatGPT rather than any LLM, surveyed 18-29-year-olds rather than secondary students, and fielded the poll in February 2024 rather than November 2024. The abstract's claim that student usage is 'higher than the usage percentage among young adults' should therefore be presented as an informal contrast with these caveats, or removed from the abstract.","section":"Section 3, Pew benchmark; Abstract"},{"comment":"The survey collected self-reports from minors, including 6th- and 7th-grade students who may be younger than 13, yet the manuscript reports no IRB approval, parental or guardian consent procedure, or child assent statement, and it does not say how the 80-cent compensation was delivered or approved for minors. For research involving children, this is a required reporting item; please add the ethics and consent information or state the applicable exemption.","section":"Section 2 (Data Collection)"}],"minor_comments":[{"comment":"The usage prevalence is given as 70% in the abstract and 71% in Section 3; please report one consistent value and state the rounding convention.","section":"Section 3; Abstract"},{"comment":"The sentence '2% of the students who rated themselves as a 5 on ethical concern reported using LLMs at least once a week' is ambiguous: the 2% could refer to all respondents or to the subset who rated 5; please state the denominator.","section":"Section 3, ethical concerns paragraph"},{"comment":"The claim that 'access to chatgpt.com is very easy with no registration requirement' is inaccurate, since ChatGPT requires an account; the intended point about low access barriers can be made without this assertion.","section":"Section 1"},{"comment":"The subject-usage percentages (for example, 28% for math and 57% for writing) do not indicate whether the denominator is all respondents or only LLM users; state the denominator in the caption or text.","section":"Figure 4; Section 3"},{"comment":"The paper states that responses are available in a public GitHub repository but gives no URL; include the link so that the data-sharing claim can be verified.","section":"Section 2"},{"comment":"The limitations paragraph acknowledges internet-access bias but not non-response bias, mode-of-administration differences, or the heavy California concentration of in-person responses; expanding the paragraph would help readers calibrate the prevalence and comparison claims.","section":"Limitations paragraph"},{"comment":"Sixth-grade respondents (n=3) are included in the reported 306 total, while the consistency claim covers grades 7-12; state explicitly whether the 6th-grade responses enter the overall usage figure and Figure 3.","section":"Table 2; Figure 3"},{"comment":"The claim that ChatGPT 'hallucinates the most when questions are inputted consecutively' is attributed to a single small study [2]; a broader or more rigorous citation, or a softening of the generalization, would be more appropriate.","section":"Section 3, hallucination citation"},{"comment":"The comparison of the policy-strictness average (3.48) with the ethical-concern average (2.83) is described as 'usage outweighs ethical concerns,' but the two numbers measure different constructs (perceived school policy vs. personal ethics) on a single-item scale; the interpretation should be labeled accordingly.","section":"Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The main gap between this manuscript and a publishable descriptive report is the distance between the sample-level statistics and the population-level language in the abstract; a quantitative reviewer will likely ask for confidence intervals and a weaker framing. For the editor: the California in-person samples were collected through the first author's school and tutoring connections, which the manuscript only indirectly signals through the affiliation; a direct disclosure would be cleaner, and the absence of any IRB/consent statement for a study of minors should be resolved before publication. The data-release practice is a genuine strength and should be preserved. The paper's fit with a cs.HC venue is reasonable if it is reframed as a descriptive survey plus design ideas rather than as a population prevalence study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague —\n\nThis is a short paper with one useful new dataset and a lot of interpretive caution that doesn't fully compensate for the sampling problems.\n\nWhat’s new: it surveys 306 middle and high school students across 43 states—more than any cited prior study—and reports that 70% of them have used an LLM. The grade-consistency observation (7th through 12th, no upward slope) is genuinely interesting, as is the subject breakdown (math still gets used by 28%, even though LLMs are known to be weak there). The authors are also transparent about their limitations, including the internet-access bias, and they don't pretend the sample is national.\n\nThe soft spots are real and load-bearing. The sample is a convenience sample: a compensated Centiment panel, a Google Form spread through the author’s tutoring and social networks, and in-person outreach in California. There is no sampling frame, no response rate, no post-stratification weighting, and no confidence intervals. So the 70% figure is a raw sample statistic. The comparison to Pew’s 43% young-adult figure is apples-to-oranges—different mode, different question, different time. The grade-invariance sub-claim is underpowered: 7th grade has n=17, so that bar is ±22 points; Figure 1 without error bars reads as absence of evidence, not evidence of consistency.\n\nAlso, the paper says responses are on a public GitHub repository, but there’s no link, and the repository isn’t named. That's an easy fix, but right now the data isn't actually available.\n\nThe proposals in Section 4 (fine-tuned models, AI tutors, AI classrooms) are reasonable but speculative and not evaluated. They read as discussion, not contribution.\n\nWould a serious referee spend time on it? Yes. The dataset is new and could be a useful baseline, and the grade-consistency finding, if it survives better sampling, is a real anomaly worth explaining. But it needs major revision: proper reporting of recruitment, margins of error, significance tests on the grade comparison, and an actual link to the data. The best version of this paper is a descriptive report with the population claims removed or heavily qualified.\n\nFor a reading group: maybe, if you're interested in AI-in-education survey methods. I wouldn't cite the 70% figure in my own work without a strong caveat.\n\nRecommendation: send it to peer review, but only with the expectation of substantial revision. The core data deserve closer study; the current claims overstate their reach.","headline":"Honest but non-representative survey of 306 US secondary students; the 70% LLM-use headline is a sample statistic, not a population estimate.","tokens_in":6206,"tokens_out":2179,"would_cite":false,"duration_ms":18406,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of 306 U.S. secondary students finds that 70 percent have used large language models, a higher share than the 43 percent of young adults who have used ChatGPT.","keywords":["large language models","secondary education","student survey","ChatGPT","AI in education","digital divide","school AI policies","middle school"],"falsifier":"A nationally representative survey of U.S. middle and high school students that asks the same 'ever used an LLM?' question, uses a documented sampling frame and weights, and finds a usage rate below 70 percent or a rate that climbs sharply from 7th to 12th grade would falsify the paper's central claim.","tokens_in":5266,"feed_emoji":"🎒","tokens_out":7226,"duration_ms":61377,"temperature":0.7,"pith_summary":"The paper tries to establish that large language models have already become a common homework tool for U.S. middle and high school students. In a survey of 306 students across 43 states, over 70 percent reported using an LLM at least once, and the share stays nearly flat from 7th to 12th grade. The authors contrast this with the 43 percent of 18-to-29-year-olds who have used ChatGPT in a February poll, and they report that usage persists even where schools have restrictive policies. They also document subject-by-subject use, a usage gap between students in top-ranked technology states and other states, and student demand for more accurate and coherent model responses. The practical stakes are that educators and developers should treat LLM use by secondary students as the norm and design curricula and models around it.","feed_headline":"70% of U.S. secondary students have tried AI chatbots, survey finds","feed_subtitle":"Use held steady from 7th to 12th grade, outpacing the 43% young-adult ChatGPT rate.","key_machinery":"The central object is a 306-respondent survey of 6th through 12th graders recruited through a compensated online panel, a web form shared via tutoring programs and social media, and in-person outreach in California. The survey asks about weekly LLM usage frequency, subjects used, school policy strictness, ethical concern, and open-ended impressions, and the analysis is descriptive: percentages, averages, and simple cross-tabulations rather than statistical modeling. The load-bearing comparison is the paper's 71 percent usage figure measured against the 43 percent young-adult ChatGPT rate from a cited February poll, and the emphasized result is the consistency of usage across grades 7 to 12.","core_discovery":"The paper's central claim is that LLM adoption among secondary students is not a fringe behavior but a majority behavior: 71 percent of survey respondents said they had used an LLM at least once, with 9 percent using one daily, and the proportion is roughly equal across grades 7 through 12. The authors take this as evidence of a 'surge' that outpaces the 43 percent ChatGPT usage rate reported for young adults, and they find that usage continues despite an average school policy strictness of 3.48 out of 5 and an average student ethical concern of 2.83 out of 5. On demographics, the paper claims an 80.2 percent usage rate in the five highest-ranked technology states versus 64.3 percent in other states, a 76.7 percent rate for private school students versus 71.3 percent for public school students, and disproportionately frequent use among the 17 respondents with paid subscriptions. The authors conclude that restriction has not worked, and they recommend fine-tuning models for educational content, building AI tutors with personalized learning paths, and creating free 'AI classrooms' to equalize access.","pith_inferences":["Beyond the paper's survey, the flat grade trend hints that LLM adoption starts before middle school and saturates by 7th grade; a longitudinal cohort study would separate this from a simple age effect.","Because the survey asked about 'ever used,' the 70 percent figure may overstate routine reliance on LLMs for homework; a diary-based or assignment-log study would establish how often and on what tasks students actually depend on them.","A natural next analysis would reweight the sample by state population and control for family income and school type, which would clarify whether the top-tech-state gap is really about technology access or about wealth.","The paper's proposed AI classrooms imply that the bottleneck is not student willingness but model quality, content grounding, and infrastructure cost; a pilot comparing a fine-tuned subject tutor against a general chatbot in an under-resourced district would test that."],"forward_implications":["If usage is already this common, school policies aimed at outright prohibition are unlikely to change behavior and should shift toward teaching appropriate use.","Since usage is consistent across grades 7 to 12, AI literacy and assignment design should address the whole secondary span, not just upper high school.","The reported gap between students in top-technology states and other states implies that unequal access to LLMs may widen educational disparities, making free AI tutoring a plausible equity intervention.","Students' emphasis on accurate, coherent, complex answers over conversational features suggests that education-focused fine-tuning should prioritize factual reliability over chat ability."],"supporting_citations":[{"why":"Supplies the 43 percent young-adult ChatGPT usage figure that the paper compares against its own 71 percent student usage rate.","marker":"[10]"},{"why":"Provides the compensated online survey panel through which a substantial portion of the 306 respondents were recruited.","marker":"[6]"},{"why":"Supplies the state technology rankings used to split respondents into top-five technology states versus other states for the 80.2 percent versus 64.3 percent usage comparison.","marker":"[11]"}],"fun_headline_variants":["71% of teens have tried AI chatbots, survey finds","AI chatbot use among students tops adult rates","Despite school bans, 71% of students use AI chatbots","Teens outpace adults in AI chatbot use, survey shows","Majority of teens use AI chatbots for schoolwork"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the 306 students who responded are representative of U.S. middle and high school students generally, even though the sample came from a compensated online panel, a convenience web form, and in-person outreach concentrated in one state with no response rates or weighting described.","fun_headline_variants_meta":{"raw":{"variants":["71% of teens have tried AI chatbots, survey finds","AI chatbot use among students tops adult rates","Despite school bans, 71% of students use AI chatbots","Teens outpace adults in AI chatbot use, survey shows","Majority of teens use AI chatbots for schoolwork"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000944,"raw_usage":{"total_tokens":4041,"prompt_tokens":963,"completion_tokens":3078,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":2999}},"tokens_in":579,"tokens_out":3078,"duration_ms":114071,"temperature":1.0,"reasoning_tokens":2999,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:56:39.598661+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A nationally representative survey of U.S. middle and high school students that asks the same 'ever used an LLM?' question, uses a documented sampling frame and weights, and finds a usage rate below 70 percent or a rate that climbs sharply from 7th to 12th grade would falsify the paper's central claim.","supporting_citations":[{"cited_title":"Americans’ use of chatgpt is ticking up, but few trust its election information","cited_arxiv_id":null,"evidence_quote":"Supplies the 43 percent young-adult ChatGPT usage figure that the paper compares against its own 71 percent student usage rate."},{"cited_title":"Survey with centiment","cited_arxiv_id":null,"evidence_quote":"Provides the compensated online survey panel through which a substantial portion of the 306 respondents were recruited."},{"cited_title":"State technology and science index","cited_arxiv_id":null,"evidence_quote":"Supplies the state technology rankings used to split respondents into top-five technology states versus other states for the 80.2 percent versus 64.3 percent usage comparison."}],"review_version":1}