{"id":"63674603-d000-448f-9c88-f7006cdf936b","arxiv_id":"2509.21890","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a graduate data science course, self-reported technical expertise predicted homework grades even with equal access to an LLM assistant, while self-rated AI familiarity and communication skills did not.","lead":"This study tracked 36 students using Google's Gemini AI in a data science course and found that students with stronger technical backgrounds still got better grades, even though everyone had the same AI tool. It also shows that novices and experts use AI differently, with experts prompting more clearly and strategically, which suggests AI training should go beyond quick demonstrations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on a marginal p-value from a small, collinear self-report model; robustness checks and a no-AI baseline are needed before 'technical expertise remains a predictor' can be accepted.","rationale":"I chose the regression robustness over the LLM-annotation validity because the paper's headline claim is specifically the RQ1 result. The reader's annotation concern is valid and material for the behavioral and pedagogical conclusions, but it is downstream of the central quantitative claim. The paper deserves credit for open-sourcing its scripts, transparently reporting the exact p-value, and for the HW0 contrast, which actually provides a useful boundary condition. My concern is not that the authors are wrong but that the evidence bar for the headline is high and the single p=.041 result is thin: small n, collinear self-reports, post-hoc composite formation, and no baseline control. The proposed reanalysis would either confirm the result across specifications or show it is fragile. The verdict remains CONDITIONAL because the paper is transparent and its claims are modest, but the stated conditions should include this robustness check in addition to annotation validation.","tokens_in":26491,"tokens_out":6537,"duration_ms":65031,"concrete_test":"Recompute Table 2 with the open-sourced data (github.com/mqo00/dspm) under three specifications: (1) each expertise composite entered alone and in all pairwise combinations, with VIFs and the predictor correlation matrix reported; (2) leave-one-participant-out refits and 1,000 bootstrap resamples; (3) HW0 no-AI grade added as a fixed covariate. If the technical coefficient loses significance (p>0.05) in any common specification, or if controlling for HW0 baseline removes its effect, the central claim should be weakened to 'technical expertise predicts grades generally, not specifically under LLM access.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption for the headline RQ1 claim is not the LLM annotation (which matters for RQ2/RQ3) but the stability of the Table 2 regression. With only 36 participants and three self-report composites, the technical-expertise effect has p=.041 (beta=6.09, SE=2.86), a single marginal result with no reported multicollinearity diagnostics, no leave-one-out analysis, and no pre-specified composite construction. Section 4.1 reports that the underlying items are correlated (rho=.52 for Python/data science and .72 for data science/toolkit), and Section 8.2 concedes that self-reported expertise is miscalibrated. The composites were formed after inspecting correlations, so the model's unique-variance partition is fragile. The paper's own HW0 analysis in Section 8.2 found that the expert-novice grade gap disappeared when LLMs were used under time pressure (Mann-Whitney p>.05), a direct boundary condition absent from the abstract's blanket claim. Finally, the model does not include a no-AI baseline (HW0 without LLM), so the result cannot distinguish 'LLM access fails to close the gap' from 'stable individual differences predict grades with or without LLMs.' A robustness reanalysis is the decisive check.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Ma et al. present a mixed-methods classroom study of 36 graduate students performing Python-based data analysis in Google Colab with the embedded Gemini assistant. They ask whether three self-reported expertise dimensions—technical, LLM, and communication—predict homework grades (RQ1); whether expertise shapes AI-use behaviors in interaction logs (RQ2); and which AI-use skills improve after lightweight demonstration and extended time (RQ3). The quantitative centerpiece is a linear mixed-effects model (Table 2) in which technical expertise significantly predicts normalized homework grades (beta=6.09, SE=2.86, p=.041) while LLM and communication expertise do not. Using a codebook of AI-use behaviors and an LLM-based annotation pipeline that segments 7,315 log events into episodes and steps, the authors report that experts write clearer prompts, use AI more strategically, and succeed more often in AI-assisted coding steps, while novices use AI reactively when challenged. A missed-opportunity analysis and surveys ground a four-dimensional taxonomy of AI-use competencies (conceptual, procedural, metacognitive, dispositional), and pre/post-demonstration comparisons suggest prompt-quality behaviors improve while evaluative behaviors lag. The paper also reports a counterbalanced HW0 comparison in which the expertise gap is significant without AI (p=.027) but not with AI (p>.05) under time pressure.","tokens_in":26787,"tokens_out":15612,"duration_ms":123330,"significance":"The study's raw material is a genuine asset: grades with substantial inter-rater reliability (kappa=0.71, §3.2), interaction logs, surveys, think-aloud recordings, a counterbalanced no-AI baseline in HW0, an open-source analysis pipeline, and complete annotation prompts and survey instruments in the appendices. If the headline result survives robustness checks, the paper would provide valuable naturalistic evidence on AI-assisted data science learning and a reusable behavior-grounded codebook for studying AI-use skills. The four-dimensional AI-literacy taxonomy and the 'less is more' scaffolding recommendation for evaluative feedback are actionable. However, the current evidence does not yet match the force of the claims: the RQ1 result is a single marginal p-value from a small, collinear, self-report-based model with post hoc composites, and the only between-condition comparison in the paper (HW0) points to a boundary condition—the gap disappeared under time pressure with AI—that is absent from the abstract. The RQ2/RQ3 findings depend on LLM annotation of moderate accuracy (70.9% on the key label) validated on 10% of steps by one author. The contribution is real but provisional.","major_comments":[{"comment":"The headline RQ1 result in Table 2 is not yet robust enough to carry the paper's central claim. The technical-expertise effect is marginal (beta=6.09, SE=2.86, p=.041) in a model with 36 participants; the predictors are self-reported composites constructed post hoc after inspecting item correlations (§4.1), and their constituent dimensions are themselves correlated (rho=.52 and .72, §4.1). Section 8.2 concedes that self-perceived expertise is subject to known miscalibration (Kruger and Dunning, ref. [24]). No multicollinearity diagnostics, leave-one-out analysis, or alternative operationalizations of the predictors are reported, and the 'not communication skills' part of the claim is weakened by the restricted variance of communication scores (median 4/5, acknowledged in §5). I request VIF or inter-composite correlations and a leave-one-participant-out check on the technical coefficient.","section":"§5, Table 2; §4.1; §8.2"},{"comment":"As specified, the Table 2 model cannot distinguish 'LLM access fails to close the gap' from 'stable individual differences predict grades with or without LLMs,' because it contains no no-AI baseline. The authors do possess such a baseline: HW0 counterbalanced tasks with and without AI, and §8.2 reports that the expertise gap was significant without AI (Mann–Whitney p=.027) but not with AI (p>.05) under time pressure. This is a direct boundary condition on the abstract's unqualified claim that technical experience 'remains a significant predictor of success,' and it should appear in the RQ1 analysis or the conclusion. I recommend a condition-by-expertise interaction analysis on HW0, and/or including HW0 no-AI task grades as a covariate in the Table 2 model.","section":"§5 vs. §8.2 (HW0 baseline)"},{"comment":"The RQ2 and RQ3 behavioral findings (Figures 4–9, the expert–novice comparisons, and the missed-opportunity ratios) rest on an LLM-based annotation whose validation is thin. Section 6.1.2 reports overall accuracy of 75%, and specifically 70.9% for the AI-use behavior labels, checked by a single author on 10% of steps; no per-code accuracy, confusion matrix, human inter-rater reliability, or label base rates are reported. For rare codes such as ai_breakdown_intent (0.65%) and decompose_task_in_prompt (0.31%), 70.9% overall accuracy could correspond to poor per-class utility, and because both the assistant under study and the annotator are LLMs, shared failure modes are a credible risk. I ask for per-code precision and recall, oversampling of rare codes in the validation sample, and a second human annotator on a shared subset with a reported kappa.","section":"§6.1.2; Figures 4–9"},{"comment":"Several load-bearing comparisons in RQ2/RQ3 are quantitative claims presented without inferential statistics: expert vs. novice success on AI-assisted coding steps (90% vs. 79%, Figure 6A), the challenge-driven vs. proactive AI-use contrast (Figure 7), and the pre/post-demonstration changes in appropriate-AI-use ratio (Figure 8). The data are clustered (steps within episodes within students), so nominal percentages overstate precision. Section 7.1's caveat that the Figure 8 changes were 'not fully mastered or statistically compared' is to the authors' credit, but the surrounding narrative in §7 and the abstract nevertheless treats the skill-acquisition findings as results. I recommend mixed-effects logistic regressions on step success and on the used/missed labels, or a consistently exploratory framing in the abstract and conclusion.","section":"§6.2, §7.1 (Figures 6–8)"}],"minor_comments":[{"comment":"Section 4.2 reports 16,315 collected events, Section 6.1.1 analyzes 7,315 events after exclusions, and the abstract says 'over 7,000 raw log events'; the relationship between these numbers should be stated at first mention.","section":"§4.2, §6.1.1, abstract"},{"comment":"Section 6.1.1 excludes logs shorter than the median as 'disengaged'; this rule conflates disengagement with fast, efficient task completion and should be justified or tested as a sensitivity analysis.","section":"§6.1.1"},{"comment":"In Table 3, the decompose_task_in_prompt row mixes a positive definition with a negative example ('[copied output] without intent would be too vague'); clarify whether the example column intentionally contains counterexamples.","section":"Table 3"},{"comment":"Figures 6–8 present expert–novice and pre–post differences as point estimates without error bars or confidence intervals; adding them would materially aid interpretation.","section":"Figures 6–8"},{"comment":"In Table 2, the LLM predictor row is formatted as '−3.052.69'; it should read '−3.05 (2.69)', and Section 4.1 contains a typo ('tookit' for 'toolkit').","section":"Table 2; §4.1"},{"comment":"The expert/novice median split (§4.1) of skewed self-report scores discards information; reporting the RQ2 comparisons against the continuous technical score would strengthen the behavioral claims.","section":"§4.1"},{"comment":"Section 6.1.1 excludes HW3–4 logs for comparability with HW0, but HW0 used a different dataset and 15-minute time limits; one sentence clarifying what is held constant across the compared assignments would help.","section":"§6.1.1"}],"recommendation":"major_revision","confidential_remarks":"Two items are decisive for the revision: (i) re-estimate RQ1 with multicollinearity diagnostics, leave-one-out checks, and a no-AI baseline (HW0 condition-by-expertise interaction), and (ii) substantially strengthen the annotation validation. If the robustness reanalysis does not support the technical-expertise coefficient, the paper should be reframed around the behavioral findings and the HW0 boundary condition rather than the headline gap claim. The HW0 counterbalanced design is the strongest causal piece in the manuscript and is currently buried in the limitations section; I would advise the authors to promote it. The paper is a good fit for an HCI or learning-analytics venue, and the open-sourced pipeline and appendices are genuine assets."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this paper. First, it contributes a genuinely useful resource: a naturalistic dataset of 36 graduate students working with an embedded LLM in Colab, with 16k logged events, screen recordings, and open-sourced processing and annotation scripts. Second, the headline claim—that technical expertise, not AI familiarity or communication skill, predicts grades—is real but weaker than the abstract suggests: the key regression has p=.041, the expertise composites were formed after looking at the data, and the paper's own HW0 comparison shows the expert-novice gap disappeared when LLMs were used under time pressure.\n\nThe strongest contribution is the behavior-grounded annotation scheme (episodes, steps, AI-use behaviors) and the LLM-based scaling of that coding. The RQ2 findings—experts prompt more strategically, provide context, use AI proactively; novices use AI reactively for debugging—are concrete and useful for anyone designing AI-literacy instruction. The knowledge-dimension framework for RQ3 is also a reasonable way to organize pedagogical targets. The authors are transparent about many limitations: small sample, miscalibrated self-reports, correlated dimensions, external tool use, modest annotation accuracy.\n\nThe soft spots are real and worth testing. The RQ1 regression is fragile: 36 students, three collinear self-report composites, no leave-one-out or alternative-specification checks, p=.041. The stress-test note is right that HW0 provides a boundary condition that the abstract ignores. A no-AI baseline doesn't fully exist for HW1-4, so the result cannot distinguish 'LLM access fails to close a stable gap' from 'stable individual differences predict grades regardless of LLMs.' The LLM annotation has 70.9% accuracy on AI-use behaviors from a single-author validation on 10% of steps; RQ2/RQ3 distributions could shift with systematic bias. The pre-post demonstration comparison is descriptive and confounded with time-on-task.\n\nStill, this is a serious paper. The dataset alone is a contribution, and the method for turning noisy notebook logs into semantic episodes is worth building on. The authors' honesty about limitations earns trust. A careful referee should push for robustness analyses, a clear presentation of the HW0 condition as a boundary on the claims, and a larger human-coded validation set. The paper deserves peer review rather than a desk reject.","headline":"Valuable dataset and annotation method, but the abstract overstates a fragile regression; deserves peer review with robustness checks.","tokens_in":27237,"tokens_out":2690,"would_cite":true,"duration_ms":23413,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Equal access to an in-notebook LLM assistant does not erase the expertise gap: in a graduate data-science class, technical skill was the only significant predictor of grades, and logs show experts prompt strategically while novices lean…","keywords":["large language models","programmatic data analysis","AI literacy","expertise gap","human-AI interaction","computational notebooks","behavioral log analysis","data science education"],"falsifier":"A concrete test: have two independent human coders, blind to student expertise, hand-annotate all 1,483 steps of the homework logs using the paper's codebook, then re-examine whether experts still show higher prompt-quality ratios, lower challenge-driven AI use, and more missed evaluation opportunities. If those expert-novice differences shrink or disappear at human-level annotation accuracy rather than the reported 75%, the behavioral account is an artifact of the automated annotator rather than a genuine expertise effect.","tokens_in":26225,"feed_emoji":"📊","tokens_out":14571,"duration_ms":108841,"temperature":0.7,"pith_summary":"This paper claims that giving students equal access to an LLM assistant inside a computational notebook does not level the playing field: self-reported technical expertise still significantly predicts homework grades, while AI familiarity and communication skill do not. Using logs from 36 graduate students completing Python data-analysis assignments in Google Colab with the built-in Gemini assistant, the authors show that experts use the LLM more strategically — clearer prompts, more context, more proactive exploration — while novices rely on it reactively to debug or explain errors. The authors introduce a behavior-grounded codebook that segments thousands of logged events into episodes and steps, annotated by an LLM, to identify where and why effective AI use breaks down. They argue that lightweight demonstrations improve surface skills like prompt clarity, but deeper evaluation, metacognitive, and dispositional skills require scaffolded instruction. If the claim is right, AI-literacy training must go beyond prompt tips and teach when to delegate, how to verify, and how to persist.","feed_headline":"Equal AI access leaves the expertise gap intact","feed_subtitle":"Technical skill, not AI familiarity, predicted who earned the higher homework grades.","key_machinery":"The paper's central analytical device is a behavior-grounded codebook for AI use, built from thematic analysis of logs and screen recordings, that segments raw log events into intent-driven episodes and four steps modeled on Norman's Stages of Action: forming intent, expressing input, understanding output, and assessing output. Each step is annotated for challenges, step success or failure, which AI-use behaviors appeared (improve, code, explain, evaluate) and their quality (clear instruction, provided context, decomposed task), plus missed opportunities where AI could have helped. The annotation is scaled up by an LLM that applied the codebook automatically, because manual annotation of a single session took over two hours. This machinery is what lets the authors move from outcomes (grades) to process (how students actually interact with the assistant), and it is what generates the claims about which skills improve after demonstration and which remain bottlenecks.","core_discovery":"The central claim is that in realistic, LLM-supported programmatic data analysis, technical expertise remains a significant predictor of success even when every student has identical access to an AI assistant. In a linear mixed-effects model with participant and assignment random intercepts, self-reported technical expertise significantly predicted normalized homework grades ($\\beta = 6.09$, $p = .041$), while LLM expertise ($p = .266$) and communication skills ($p = .381$) did not. Log analysis shows that experts and novices use AI at similar overall frequencies, but experts write clearer prompts, provide more context, decompose tasks, and ask for explanations mainly when they hit genuine obstacles; novices turn to AI reactively and struggle at the stages of forming intent, expressing input, interpreting output, and assessing results. The paper further claims that a 90-minute live-coding demonstration plus extended time-on-task improved prompt-quality behaviors by roughly 30% in the appropriate-use ratio, yet left evaluation behaviors (critiquing outputs, reading long explanations) below half of their opportunities, indicating those skills require guided practice and feedback rather than demonstration alone.","pith_inferences":["A testable extension the paper does not run: randomly assign a 'verify your output' checklist scaffold to half of the novices and measure whether the evaluation-skill bottleneck closes faster than it did from demonstration alone.","The paper's own observation that the course tasks were not 'LLM-hard' suggests a boundary condition: on harder, open-ended tasks where current models struggle, prior technical expertise may matter even more, so the size of the expertise gap is probably not a constant.","A methodological corollary: with human annotation of a single student log taking over two hours, the LLM-annotation route is what makes log-scale behavioral studies feasible, but its reported 75% accuracy means future replications should report uncertainty bounds on behavior frequencies.","A reporting caveat the paper itself gives: the appropriate-use ratios for rare behaviors (improve-prompt, decompose-task, edit-partial-code) rest on one or two log instances per student, so those post-demo gains should be read as suggestive rather than measured."],"forward_implications":["Instructors should treat equal tool access as necessary but not sufficient: in this classroom, giving every student the same notebook-embedded assistant left the performance gap intact, so effective AI-literacy training must target the behaviors that actually separate users.","The episode-step codebook gives teachers and tool designers a diagnostic vocabulary: a failure can be located in forming intent, expressing input, understanding output, or assessing results, and each failure type points to a different AI-use skill to teach.","Lightweight instruction has a measurable but bounded payoff: after one live-coding demonstration, prompt-clarity and context-providing behaviors increased in the logs, but evaluation behaviors stayed below half of their available opportunities.","AI-use competence is not a single skill: the paper's four knowledge dimensions imply that curricula should combine conceptual facts about what a given assistant can see and do, procedural practice in phrasing requests, metacognitive training in when to delegate, and support for persisting through errors.","Assistants embedded in notebooks should make their data access legible: recurring hallucination loops stem from students assuming the model can see the dataframe, so surfacing exactly what context the model has should reduce an entire failure class."],"supporting_citations":[{"why":"prior result that experts benefit more than novices from LLM companions; the expert-novice gap this study extends to notebook-based data analysis.","marker":"[5]"},{"why":"Norman's Stages of Action, the cognitive framework behind the paper's four-step episode annotation (intent, input, understand, assess).","marker":"[46]"},{"why":"the mixed-effects regression implementation behind the RQ1 model that identifies technical expertise as the only significant predictor.","marker":"[2]"},{"why":"prior work on requirement-driven prompting whose conclusion that prompt tips alone do not teach effective use underlies the pedagogical findings.","marker":"[30]"},{"why":"the finding that self-reported expertise is miscalibrated, which the paper invokes as a caveat on its self-report predictors.","marker":"[24]"},{"why":"the knowledge-dimensions taxonomy (conceptual, procedural, metacognitive, dispositional) into which the paper sorts AI-use skills.","marker":"[23]"},{"why":"evidence that beginning programmers do not learn prompting from code feedback alone, a premise behind the demonstration-versus-practice distinction.","marker":"[36]"},{"why":"the established notebook-logging toolkit cited as a more robust alternative to the paper's custom Colab instrumentation.","marker":"[49]"}],"fun_headline_variants":["Technical skill, not AI savvy, predicts LLM success","AI access alone won't close the data science gap","In AI classrooms, technical know-how still wins","Why AI tools favor the already technical","LLMs in class: experts still outperform, says study"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The behavioral findings rest on the assumption that the automated annotation of student logs, which matched a single human coder's review of 10% of steps in about 75% of cases (71% for AI-use behaviors), correctly identifies what students actually did with the assistant.","fun_headline_variants_meta":{"raw":{"variants":["Technical skill, not AI savvy, predicts LLM success","AI access alone won't close the data science gap","In AI classrooms, technical know-how still wins","Why AI tools favor the already technical","LLMs in class: experts still outperform, says study"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000479,"raw_usage":{"total_tokens":2358,"prompt_tokens":916,"completion_tokens":1442,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":1368}},"tokens_in":532,"tokens_out":1442,"duration_ms":11323,"temperature":1.0,"reasoning_tokens":1368,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:44:18.598741+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: have two independent human coders, blind to student expertise, hand-annotate all 1,483 steps of the homework logs using the paper's codebook, then re-examine whether experts still show higher prompt-quality ratios, lower challenge-driven AI use, and more missed evaluation opportunities. If those expert-novice differences shrink or disappear at human-level annotation accuracy rather than the reported 75%, the behavioral account is an artifact of the automated annotator rather than a genuine expertise effect.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Norman's Stages of Action, the cognitive framework behind the paper's four-step episode annotation (intent, input, understand, assess)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the established notebook-logging toolkit cited as a more robust alternative to the paper's custom Colab instrumentation."}],"review_version":1}