{"id":"9d00f189-4d0f-4bc0-9af7-1b451c05b71c","arxiv_id":"1908.06193","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Students often switched tabs while taking online physics assessments, but the effects on average class scores were small, so the online results remain broadly comparable to paper-based tests.","lead":"This study tracked whether physics students copied text, printed pages, or switched away from browser tabs while taking online concept tests at five universities, and matched those actions to their scores. The purpose is to check whether online tests are still trustworthy and comparable to the old paper versions, since students can more easily look up answers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that copying had negligible impact on class averages rests on a removal analysis that conflates selection with behavior; copyers score 0.45 SD higher, so the 1% drop is the expected selection effect.","rationale":"The reader's conditional verdict is reasonable, but my stress-test identifies a different load-bearing issue than the reader's weakest_assumption. The abstract's causal claim ('resulted in') is not supported by the analysis used to justify it. The removal-based average comparison in Sec. IV.C is a selection difference, not a treatment effect. This is not a pedantic point: the paper's own data show copyers are 0.45 SD above non-copyers, so the 1% drop is exactly what would be expected from removing high-scoring students regardless of any benefit from copying. The item-level analysis is a strength and suggests copying does help on copied items, but the paper does not use it to calculate the population-average counterfactual. A focused reanalysis could settle the question. Because the paper's descriptive findings (prevalence, timing, and the within-student item comparison) remain valuable and the flaw is fixable, I would keep the reader's CONDITIONAL verdict rather than reject; the condition should include the reanalysis proposed above.","tokens_in":12096,"tokens_out":8188,"duration_ms":87045,"concrete_test":"Re-estimate the class-average copying effect from the item-level data rather than the removal analysis: for each of the 147 introductory students with copy events, compute the within-student difference in success probability between items they copied and items they did not copy, residualized for item difficulty; then aggregate as (mean per-student gain) × (mean fraction of items copied by copyers) × (147/1287). If the implied class-average shift is below about 0.3% and matches the reported mean per-item shift, the central claim survives; if it exceeds about 1% or differs materially from the removal estimate, the 'negligible impact' conclusion is not supported by the analysis as reported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that no targeted online behavior 'resulted in a significant change in the population's average performance.' For copying, the quantitative support in Sec. IV.C is the 1% drop in the introductory average after removing all students with copy events (Cohen's d = 0.05, p = 0.2). That estimate conflates treatment with selection. If f is the fraction of students with copy events (147/1287 ≈ 0.114) and μ_copy, μ_non are the mean z-scores of copyers and non-copyers, the removal difference is f(μ_copy − μ_non). The paper reports μ_copy − μ_non = 0.45 SD, so removing copyers lowers the average by roughly 0.05 SD even if copying raises nobody's score. The item-level analysis (Table IV; copied questions correct 77% vs 58%, within-student z-difference 0.44) provides better evidence of a per-item benefit, but it is not converted into a class-average counterfactual. Thus the stated basis for 'negligible impact on class average' cannot distinguish 'copying inflates scores' from 'higher-scoring students copy.' For focus events, no analogous removal or counterfactual analysis is reported at all; the positive 0.16 correlation for introductory students is consistent with either distraction or resource use.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an observational study of student behavior during online administration of four research-based physics assessments (FMCE, BEMA, CUE, QMCA) at eight institutions, using embedded JavaScript to record copy, print, and browser-focus events. The authors find that browser focus loss is common, while copying and printing are rare, and that correlations between these behaviors and student scores differ between introductory and upper-division populations. They conclude that none of the monitored behaviors materially changes the population average score, so online RBAs remain comparable to paper-based implementations.","tokens_in":12318,"tokens_out":3878,"duration_ms":38717,"significance":"If the central claim holds, the study gives practical reassurance for online RBA administration and extends prior work by Bonham with direct browser telemetry across a multi-institution sample. Strengths include the use of JavaScript rather than self-report, explicit attention to both introductory and upper-division populations, and a large valid-response sample (N=1287 introductory, N=308 upper-division). The paper also provides useful descriptive data on the prevalence and duration of focus-loss events. However, the key quantitative support for the 'negligible impact on class average' claim is currently flawed and needs revision before the central conclusion can be accepted.","major_comments":[{"comment":"The removal analysis that supports the 'negligible impact on class average' conclusion conflates selection with behavior. Removing the 147 introductory students with copy events lowers the class average by approximately f(μ_copy − μ_non) ≈ 0.114 × 0.45 ≈ 0.05 SD simply because copyers are higher-scoring, even if copying has no causal effect on any student's score. The reported 1% drop with d=0.05 is therefore exactly what the selection effect predicts, and cannot distinguish 'copying inflates scores' from 'higher-achieving students copy.' The item-level analysis in Table IV and the within-student z-difference of 0.44 provide stronger evidence of a per-item benefit, but these estimates are not converted into a class-average counterfactual (for example, using the fraction of items copied and the per-item benefit). Without such a calculation, the central claim is not supported by the reported statistics.","section":"Sec. IV.C"},{"comment":"No removal or counterfactual analysis is reported for browser focus events, despite the claim in Sec. V that disengagement 'does not appear to negatively impact their performance.' The Spearman correlation of r=0.16 for introductory students is positive and statistically significant; it is consistent with resource use, with distraction, or with selection. The paper should report the difference in mean scores with and without focus events (analogous to the copy-event removal) or otherwise bound the effect of focus loss on the class average before concluding that focus behavior is harmless.","section":"Sec. IV.B and Sec. V"},{"comment":"The Discussion states that 'the improvement to students scores is, on average small (roughly a third of a standard deviation of improvement in average score)', which is inconsistent with the reported 1% drop (d=0.05) from the copy-event removal analysis. The 0.44 SD within-student z-difference is a per-copier, per-item benefit, not a class-average effect. This conflation should be corrected, and the Discussion should state the class-average effect size explicitly and consistently for both the per-item and whole-test comparisons.","section":"Sec. V"}],"minor_comments":[{"comment":"The text 'coping or highlighted text' appears to be a typo for 'copying or highlighted text'.","section":"Sec. II"},{"comment":"The phrases 'loosing browser focus' and 'deferentially include' should read 'losing browser focus' and 'differentially include'.","section":"Sec. V"},{"comment":"The chi-squared test on Table IV pools student-question responses, which are not independent; reporting a student-level or mixed-effects analysis would strengthen the per-item benefit claim beyond the already-provided within-student z-score comparison.","section":"Table IV and Sec. IV.C"},{"comment":"The historical paper-based comparisons are neither concurrent nor matched on cohort composition; the paper should more explicitly state how this limits the 5% drop interpretation, even if the drop is attributed mainly to participation differences.","section":"Sec. IV and Sec. V"},{"comment":"The copy-to-search classification uses a 5-second window; a brief robustness check with different windows (e.g., 2 s or 10 s) would clarify how sensitive the prevalence estimates are to this threshold.","section":"Sec. IV.C"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a physics education research journal and addresses a practically important question. The data collection is a real strength, but the central quantitative argument needs reanalysis rather than new data. I would be willing to review a revised version that provides a proper counterfactual for the copy events and a clearer treatment of focus events."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look, and worth a referee. What's actually new here is the item-level copy data and the population difference: introductory students who copy text score higher, upper-division copiers score lower, and the within-student z-difference of 0.44 on copied items is direct evidence that the copy-to-search sequence pays off for some students. That extends Bonham's earlier work and gives the field a much sharper picture of what happens when low-stakes conceptual assessments go online. The data collection is careful—JavaScript telemetry, time-stamped events, large multi-institution sample, high participation rates—and the authors are appropriately cautious about what browser-level events can and cannot show. They also flag their own limitations, which is more than many papers in this space do.\n\nThe stress-test note is right about the removal analysis. Removing all copyers from the introductory sample and seeing a 1% drop is mathematically what you'd expect from dropping a higher-scoring subgroup: 0.114 × 0.45 SD ≈ 0.05 SD. So that analysis does not distinguish \"copying inflates scores\" from \"higher-scoring students copy.\" That's a genuine flaw in the stated basis for the central claim. The better evidence is the per-item within-student analysis, but the paper never converts it into an explicit class-average counterfactual. My guess is the conclusion survives the fix—copyers are only ~10% and they copy a small fraction of items—but the authors should show that rather than lean on the removal arithmetic.\n\nThe other soft spots are minor: the historical paper controls aren't concurrent, the 5% drop is attributed to participation without a direct test, and there are many comparisons without correction. None of these threaten the main empirical contribution. The proxy thresholds (4-second hidden, 5-second copy-to-search) are reasonable and the paper acknowledges they're imperfect.\n\nThis paper is primarily for PER researchers and assessment practitioners who need to know whether online administration of concept inventories is defensible. It deserves serious peer review; the analysis needs revision, but the empirical work is solid and the question is important. I'd take it with revisions and ask specifically for a corrected class-average impact estimate that separates selection from behavior.","headline":"Solid, useful empirical study of online RBA behavior with a real flaw in the removal analysis (selection vs. effect), but the practical conclusion is probably right and deserves referee time.","tokens_in":12854,"tokens_out":1957,"would_cite":true,"duration_ms":22530,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Online physics assessments stay comparable to paper at the class level, despite tab-switching and copying.","keywords":["online assessment","research-based assessments","browser telemetry","test security","concept inventories","physics education research","student cheating behavior","assessment comparability"],"falsifier":"A replication with matched cohorts—sections randomly assigned to online or paper administration in the same term, with screen recordings of a sample of online students—would settle it: if the copy-then-hide pattern does not correspond to actual web searches, or if the online-minus-paper score gap remains when participation is matched, the comparability claim would be undermined.","tokens_in":11861,"feed_emoji":"📊","tokens_out":9913,"duration_ms":93522,"temperature":0.7,"pith_summary":"Online research-based assessments (RBAs) are standardized conceptual tests that physics instructors give before and after instruction. This paper asks whether moving them from paper-and-pencil, in-class administration to web-based, out-of-class administration changes what the scores mean. Using JavaScript embedded in four RBAs across introductory and upper-division courses, the authors logged students' copy commands, print commands, and browser tab-switches. They found that a majority of students engaged in at least one of these behaviors—most commonly clicking away from the assessment for under a minute—but that the behaviors had little aggregate effect: removing all students who copied text shifted the introductory class average by only about 1%. The paper's claim is that online RBAs, at current levels of these behaviors, remain interpretable at the group level and comparable to paper-based implementations.","feed_headline":"Copying and tab-switching don't skew online physics scores","feed_subtitle":"Class averages stayed within about one percent of paper versions even though most students clicked away at least once.","key_machinery":"The load-bearing mechanism is a set of JavaScript event listeners embedded in the online survey platform, which detect browser-level actions: print commands, copy commands, and changes in browser focus. A focus loss is counted as a hidden event only if the assessment tab remains hidden for more than four seconds; a copy event is flagged as a probable web search if a hidden focus event follows within five seconds. The copy-then-hide pairing is the operational bridge between raw telemetry and the paper's interpretation of why some students score higher on copied questions. The rest of the analysis consists of within-course z-scores to pool data across institutions, non-parametric rank comparisons, rank correlations, and chi-squared tests for item-level contingency.","core_discovery":"The central claim is that browser-level behaviors unique to online administration—printing, copying text, and losing browser focus—did not, in this sample, change population-average scores enough to threaten interpretation or comparison with paper versions. Loss of focus was common (46% of introductory and 52% of upper-division students hid the tab at least once), but most absences lasted under a minute. Copying was rarer, involving about one in ten students, and three-quarters of copy events were followed within five seconds by a hidden-tab event, the pattern the paper interprets as searching the web. Copying correlated with higher scores for introductory students and lower scores for upper-division students, and introductory students were more likely to get exactly the question right from which they had copied text. Yet removing all copyers from the introductory data changed the course average by about 1% (effect size d = 0.05), and the paper concludes that the aggregate effect is negligible, with the only notable online-versus-paper gap being a roughly 5% lower average that it attributes to broader participation from lower-performing students.","pith_inferences":["A natural extension would be a controlled comparison of the same online RBA under open-browser and lockdown-browser conditions; if the copy-then-hide benefit disappears when navigation is blocked, the paper's inferred search mechanism is confirmed.","Because the telemetry records only browser-level actions, the reported copy prevalence is a lower bound on outside-resource use; students consulting a phone or another computer would look like non-copyers, so the true 'clean' comparison group may be smaller than it appears.","The item-level result implies a testable threshold: as solution websites spread, the per-question copying benefit should grow, and monitoring that benefit over time could serve as an early-warning signal for when an online RBA's group average is no longer comparable to its paper baseline.","A difference-in-differences reanalysis of the item-level data—comparing a student's copied versus uncopied questions of similar difficulty, rather than comparing copyers with non-copyers—would test whether the score gain is caused by the web search or by selection of which students copy."],"forward_implications":["Instructors can assign these RBAs online outside of class and still treat the resulting course mean as a group-level measure of learning, gaining back class time and centralizing the scoring.","The roughly 5% drop in online scores relative to paper appears to reflect who shows up, not how they behave: online administration reaches more of the lower-performing tail, which is a sampling advantage rather than a validity threat.","Copying text predicts higher item-level success only when answers are findable online, as they already are for the introductory instruments used here; for the upper-division instruments, copyers scored lower, suggesting the searches did not pay off.","The conclusion is conditional on copying staying rare; if the fraction of students who copy grows, the same per-student gains would eventually move the class average and the comparability claim would need revisiting.","Periodic searches for posted prompts and solutions are needed because the security of an RBA can erode over time as it becomes older and more widely used."],"supporting_citations":[{"why":"Supplies the JavaScript event-detection method and the prior prevalence findings this study replicates and extends.","marker":"[18]"},{"why":"Establishes the prior result that online and in-class concept inventory scores are statistically comparable when participation is similar.","marker":"[13]"},{"why":"Documents lower online participation rates and the best-practice reminders and credit that can close the gap.","marker":"[12]"},{"why":"Provides an earlier web-versus-paper comparison finding no score difference, a key baseline for comparability.","marker":"[6]"},{"why":"Frames the study by warning that online and paper equivalence must be demonstrated empirically, not assumed.","marker":"[16]"},{"why":"Earlier proposal of an online administration and analysis model that this study's implementation builds on.","marker":"[5]"}],"fun_headline_variants":["Online physics scores resist copy-paste and tab-hopping","Most students click away, but physics averages hold","Copy and tab tricks don't move online physics scores","Browser clicks fail to skew online physics test results","Physics scores stable despite students' online detours"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on treating a browser tab hidden for more than four seconds and a copy event followed within five seconds by a hidden tab as reliable evidence of outside-resource use, and on treating earlier paper-based scores from the same courses as a fair baseline despite different cohorts and conditions.","fun_headline_variants_meta":{"raw":{"variants":["Online physics scores resist copy-paste and tab-hopping","Most students click away, but physics averages hold","Copy and tab tricks don't move online physics scores","Browser clicks fail to skew online physics test results","Physics scores stable despite students' online detours"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1421,"prompt_tokens":1054,"completion_tokens":367,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":670,"completion_tokens_details":{"reasoning_tokens":293}},"tokens_in":670,"tokens_out":367,"duration_ms":4768,"temperature":1.0,"reasoning_tokens":293,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:52:57.894318+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication with matched cohorts—sections randomly assigned to online or paper administration in the same term, with screen recordings of a sample of online students—would settle it: if the copy-then-hide pattern does not correspond to actual web searches, or if the online-minus-paper score gap remains when participation is matched, the comparability claim would be undermined.","supporting_citations":[{"cited_title":"Reliability, compliance, and security in web-based course assessments,","cited_arxiv_id":null,"evidence_quote":"Supplies the JavaScript event-detection method and the prior prevalence findings this study replicates and extends."},{"cited_title":"Performance on in- class vs. online administration of concept inventories,","cited_arxiv_id":null,"evidence_quote":"Establishes the prior result that online and in-class concept inventory scores are statistically comparable when participation is similar."},{"cited_title":"Participation rates of in-class vs. online administration of low-stakes research-based assessments,","cited_arxiv_id":null,"evidence_quote":"Documents lower online participation rates and the best-practice reminders and credit that can close the gap."},{"cited_title":"Standardized test- ing in physics via the world wide web,","cited_arxiv_id":null,"evidence_quote":"Provides an earlier web-versus-paper comparison finding no score difference, a key baseline for comparability."},{"cited_title":"The equivalence of paper-and-pencil and computer-based testing,","cited_arxiv_id":null,"evidence_quote":"Frames the study by warning that online and paper equivalence must be demonstrated empirically, not assumed."},{"cited_title":"Alternative model for adminis- tration and analysis of research-based assessments,","cited_arxiv_id":null,"evidence_quote":"Earlier proposal of an online administration and analysis model that this study's implementation builds on."}],"review_version":1}