Pith. sign in

REVIEW 3 major objections 5 minor 23 references

Investigating students' behavior and performance in online conceptual assessment

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Online physics assessments stay comparable to paper at the class level, despite tab-switching and copying.

desk verdict Solid, useful empirical study of online RBA behavior with a real flaw in the removal analysis (selection vs. effect), but the practical conclusion is probably right and deserves referee time. read the letter →

arxiv 1908.06193 v1 pith:KIOHGRWZ submitted 2019-08-16 physics.ed-ph

classification physics.ed-ph
keywords onlineassessmentresearch-basedassessmentsbrowsertelemetrytestsecurityconceptinventoriesphysicseducationresearchstudentcheatingbehaviorcomparability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Online research-based assessments (RBAs) are standardized conceptual tests that physics instructors give before and after instruction. This paper asks whether moving them from paper-and-pencil, in-class administration to web-based, out-of-class administration changes what the scores mean. Using JavaScript embedded in four RBAs across introductory and upper-division courses, the authors logged students' copy commands, print commands, and browser tab-switches. They found that a majority of students engaged in at least one of these behaviors—most commonly clicking away from the assessment for under a minute—but that the behaviors had little aggregate effect: removing all students who copied text shifted the introductory class average by only about 1%. The paper's claim is that online RBAs, at current levels of these behaviors, remain interpretable at the group level and comparable to paper-based implementations.

What carries the argument

The load-bearing mechanism is a set of JavaScript event listeners embedded in the online survey platform, which detect browser-level actions: print commands, copy commands, and changes in browser focus. A focus loss is counted as a hidden event only if the assessment tab remains hidden for more than four seconds; a copy event is flagged as a probable web search if a hidden focus event follows within five seconds. The copy-then-hide pairing is the operational bridge between raw telemetry and the paper's interpretation of why some students score higher on copied questions. The rest of the analysis consists of within-course z-scores to pool data across institutions, non-parametric rank comparisons, rank correlations, and chi-squared tests for item-level contingency.

What would settle it

A replication with matched cohorts—sections randomly assigned to online or paper administration in the same term, with screen recordings of a sample of online students—would settle it: if the copy-then-hide pattern does not correspond to actual web searches, or if the online-minus-paper score gap remains when participation is matched, the comparability claim would be undermined.

Watch

Extended reading notes

Core claim

The central claim is that browser-level behaviors unique to online administration—printing, copying text, and losing browser focus—did not, in this sample, change population-average scores enough to threaten interpretation or comparison with paper versions. Loss of focus was common (46% of introductory and 52% of upper-division students hid the tab at least once), but most absences lasted under a minute. Copying was rarer, involving about one in ten students, and three-quarters of copy events were followed within five seconds by a hidden-tab event, the pattern the paper interprets as searching the web. Copying correlated with higher scores for introductory students and lower scores for upper-division students, and introductory students were more likely to get exactly the question right from which they had copied text. Yet removing all copyers from the introductory data changed the course average by about 1% (effect size d = 0.05), and the paper concludes that the aggregate effect is negligible, with the only notable online-versus-paper gap being a roughly 5% lower average that it attributes to broader participation from lower-performing students.

Load-bearing premise

The argument depends on treating a browser tab hidden for more than four seconds and a copy event followed within five seconds by a hidden tab as reliable evidence of outside-resource use, and on treating earlier paper-based scores from the same courses as a fair baseline despite different cohorts and conditions.

Editorial extensions

If this is right

  • Instructors can assign these RBAs online outside of class and still treat the resulting course mean as a group-level measure of learning, gaining back class time and centralizing the scoring.
  • The roughly 5% drop in online scores relative to paper appears to reflect who shows up, not how they behave: online administration reaches more of the lower-performing tail, which is a sampling advantage rather than a validity threat.
  • Copying text predicts higher item-level success only when answers are findable online, as they already are for the introductory instruments used here; for the upper-division instruments, copyers scored lower, suggesting the searches did not pay off.
  • The conclusion is conditional on copying staying rare; if the fraction of students who copy grows, the same per-student gains would eventually move the class average and the comparability claim would need revisiting.
  • Periodic searches for posted prompts and solutions are needed because the security of an RBA can erode over time as it becomes older and more widely used.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension would be a controlled comparison of the same online RBA under open-browser and lockdown-browser conditions; if the copy-then-hide benefit disappears when navigation is blocked, the paper's inferred search mechanism is confirmed.
  • Because the telemetry records only browser-level actions, the reported copy prevalence is a lower bound on outside-resource use; students consulting a phone or another computer would look like non-copyers, so the true 'clean' comparison group may be smaller than it appears.
  • The item-level result implies a testable threshold: as solution websites spread, the per-question copying benefit should grow, and monitoring that benefit over time could serve as an early-warning signal for when an online RBA's group average is no longer comparable to its paper baseline.
  • A difference-in-differences reanalysis of the item-level data—comparing a student's copied versus uncopied questions of similar difficulty, rather than comparing copyers with non-copyers—would test whether the score gain is caused by the web search or by selection of which students copy.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports an observational study of student behavior during online administration of four research-based physics assessments (FMCE, BEMA, CUE, QMCA) at eight institutions, using embedded JavaScript to record copy, print, and browser-focus events. The authors find that browser focus loss is common, while copying and printing are rare, and that correlations between these behaviors and student scores differ between introductory and upper-division populations. They conclude that none of the monitored behaviors materially changes the population average score, so online RBAs remain comparable to paper-based implementations.

Significance. If the central claim holds, the study gives practical reassurance for online RBA administration and extends prior work by Bonham with direct browser telemetry across a multi-institution sample. Strengths include the use of JavaScript rather than self-report, explicit attention to both introductory and upper-division populations, and a large valid-response sample (N=1287 introductory, N=308 upper-division). The paper also provides useful descriptive data on the prevalence and duration of focus-loss events. However, the key quantitative support for the 'negligible impact on class average' claim is currently flawed and needs revision before the central conclusion can be accepted.

major comments (3)
  1. [Sec. IV.C] The removal analysis that supports the 'negligible impact on class average' conclusion conflates selection with behavior. Removing the 147 introductory students with copy events lowers the class average by approximately f(μ_copy − μ_non) ≈ 0.114 × 0.45 ≈ 0.05 SD simply because copyers are higher-scoring, even if copying has no causal effect on any student's score. The reported 1% drop with d=0.05 is therefore exactly what the selection effect predicts, and cannot distinguish 'copying inflates scores' from 'higher-achieving students copy.' The item-level analysis in Table IV and the within-student z-difference of 0.44 provide stronger evidence of a per-item benefit, but these estimates are not converted into a class-average counterfactual (for example, using the fraction of items copied and the per-item benefit). Without such a calculation, the central claim is not supported by the reported statistics.
  2. [Sec. IV.B and Sec. V] No removal or counterfactual analysis is reported for browser focus events, despite the claim in Sec. V that disengagement 'does not appear to negatively impact their performance.' The Spearman correlation of r=0.16 for introductory students is positive and statistically significant; it is consistent with resource use, with distraction, or with selection. The paper should report the difference in mean scores with and without focus events (analogous to the copy-event removal) or otherwise bound the effect of focus loss on the class average before concluding that focus behavior is harmless.
  3. [Sec. V] The Discussion states that 'the improvement to students scores is, on average small (roughly a third of a standard deviation of improvement in average score)', which is inconsistent with the reported 1% drop (d=0.05) from the copy-event removal analysis. The 0.44 SD within-student z-difference is a per-copier, per-item benefit, not a class-average effect. This conflation should be corrected, and the Discussion should state the class-average effect size explicitly and consistently for both the per-item and whole-test comparisons.
minor comments (5)
  1. [Sec. II] The text 'coping or highlighted text' appears to be a typo for 'copying or highlighted text'.
  2. [Sec. V] The phrases 'loosing browser focus' and 'deferentially include' should read 'losing browser focus' and 'differentially include'.
  3. [Table IV and Sec. IV.C] The chi-squared test on Table IV pools student-question responses, which are not independent; reporting a student-level or mixed-effects analysis would strengthen the per-item benefit claim beyond the already-provided within-student z-score comparison.
  4. [Sec. IV and Sec. V] The historical paper-based comparisons are neither concurrent nor matched on cohort composition; the paper should more explicitly state how this limits the 5% drop interpretation, even if the drop is attributed mainly to participation differences.
  5. [Sec. IV.C] The copy-to-search classification uses a 5-second window; a brief robustness check with different windows (e.g., 2 s or 10 s) would clarify how sensitive the prevalence estimates are to this threshold.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central comparability claim is supported by direct telemetry and score comparisons, not by a fitted parameter or self-referential definition.

full rationale

The paper's central claim is that students' online behaviors (copying, printing, loss of browser focus) did not change population average performance enough to threaten interpretation or comparison to paper RBAs. This claim is supported by direct measurements: JavaScript recorded actual browser events, and the analysis compares score distributions of students with and without those events. No parameter is fitted from the outcome and then renamed as a prediction; the removal analysis is a straightforward arithmetic comparison of the full sample to the subset without copy events. The operational definitions (hidden focus >4 s; copy followed within 5 s by hidden focus) are explicit measurement choices rather than equations whose output is the conclusion. The paper also draws on prior work on online/paper comparability (e.g., LASSO and MacIsaac) and on RBA validation, including some self-citations, but those are external empirical results about assessment equivalence and instrument validity; the present paper's inference about behavior impact does not reduce to those citations. The selection-versus-treatment confound in the copy-event removal analysis is a possible internal-validity limitation, not a circularity: the reported 1% average drop is computed, not assumed, and the paper additionally reports item-level contingency data. Under the stated standard of flagging circularity only when the derivation is equivalent to its inputs by construction, no circular step is present.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The central claim relies on three hand-chosen thresholds that define the measured behaviors, plus three domain assumptions about what the telemetry means and how the comparison baseline is constructed. There are no newly postulated physical or conceptual entities; the 'copy-to-search' pattern is an interpretive construct built from existing measurements.

free parameters (3)
  • Hidden focus threshold = 4 seconds
    A browser focus loss is logged as 'hidden' only if the RBA tab is not visible 4 seconds after the event. This hand-chosen cutoff determines the prevalence and duration of focus-loss events, a core dependent measure.
  • Copy-to-search window = 5 seconds
    A copy event is counted as consistent with web searching only if followed within 5 seconds by a sustained hidden focus event. This choice shapes the finding that three-quarters of copy events fit the searching pattern.
  • Sustained disengagement cutoff = 5 minutes
    Focus events longer than 5 minutes are used to characterize long disengagement. This analytic cutoff affects the narrative about how long students left the assessment.
assumptions (3)
  • domain assumption Browser-level focus loss is a valid proxy for student disengagement from the assessment.
    The focus-event analysis interprets a hidden browser tab as the student not working on the RBA, but the student could be reading another tab relevant to the test, taking notes, or checking messages. Invoked in Sec. III and Sec. IV B.
  • domain assumption Historical paper-based administrations in the same courses are a valid baseline for online comparison.
    Sec. V compares online scores to historical in-class scores from the same courses and instructors, assuming cohort and term effects are negligible, with no concurrent control group.
  • domain assumption RBAs are intended for low-stakes group-level interpretation, not individual scores.
    The paper relies on this framing, citing Engelhardt, to argue that small impacts on individual scores do not threaten group-level interpretation. Used in Sec. IV C to contextualize the copying results.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Investigating students' behavior and performance in online conceptual assessment." pith.science (2026). https://pith.science/paper/KIOHGRWZ

@misc{pith2026190806193,
  author       = {Pith},
  title        = {Pith review of: Investigating students' behavior and performance in online conceptual assessment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KIOHGRWZ}},
  note         = {Machine review of arXiv:1908.06193}
}
read the original abstract

Historically, the implementation of research-based assessments (RBAs) has been a driver of educational change within physics and helped motivate adoption of interactive engagement pedagogies. Until recently, RBAs were given to students exclusively on paper and in-class; however, this approach has important drawbacks including decentralized data collection and the need to sacrifice class time. Recently, some RBAs have been moved to online platforms to address these limitations. Yet, online RBAs present new concerns such as student participation rates, test security, and students' use of outside resources. Here, we report on a study addressing these concerns in both upper-division and lower-division undergraduate physics courses. We gave RBAs to courses at five institutions; the RBAs were hosted online and featured embedded JavaScript code which collected information on students' behaviors (e.g., copying text, printing). With these data, we examine the prevalence of these behaviors, and their correlation with students' scores, to determine if online and paper-based RBAs are comparable. We find that browser loss of focus is the most common online behavior while copying and printing events were rarer. We found that correlations between these behaviors and student performance varied significantly between introductory and upper-division student populations, particularly with respect to the impact of students copying text in order to utilize internet resources. However, while the majority of students engaged in one or more of the targeted online behaviors, we found that, for our sample, none of these behaviors resulted in a significant change in the population's average performance that would threaten our ability to interpret this performance or compare it to paper-based implementations of the RBA.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

23 extracted references · 23 canonical work pages

  1. [1]

    Resource letter rbai-1: research-based assessmen t instruments in physics and astronomy,

    Adrian Madsen, Sarah B McKagan, and Eleanor C Sayre, “Resource letter rbai-1: research-based assessmen t instruments in physics and astronomy,” American Jour- nal of Physics 85, 245–264 (2017)

  2. [2]

    The student-centered activities for large enrollmen t undergraduate programs (scale-up) project,

    Robert J Beichner, Jeffery M Saul, David S Abbott, Jeanne J Morse, Duane Deardorff, Rhett J Allain, Scott W Bonham, Melissa H Dancy, and John S Ris- ley, “The student-centered activities for large enrollmen t undergraduate programs (scale-up) project,” Research- based reform of university physics 1, 2–39 (2007)

  3. [3]

    Toward a modeling theory of physics instruction,

    David Hestenes, “Toward a modeling theory of physics instruction,” American journal of physics 55, 440–454 (1987)

  4. [4]

    Peer instruction: Ten years of experience and results,

    Catherine H Crouch and Eric Mazur, “Peer instruction: Ten years of experience and results,” American journal of physics 69, 970–977 (2001)

  5. [5]

    Alternative model for adminis- tration and analysis of research-based assessments,

    Bethany R. Wilcox, Benjamin M. Zwickl, Robert D. Hobbs, John M. Aiken, Nathan M. Welch, and H. J. Lewandowski, “Alternative model for adminis- tration and analysis of research-based assessments,” 9 Phys. Rev. Phys. Educ. Res. 12, 010139 (2016)

  6. [6]

    Standardized test- ing in physics via the world wide web,

    Dan MacIsaac, Rebecca Pollard Cole, David M Cole, Laura McCullough, and Jim Maxka, “Standardized test- ing in physics via the world wide web,” Electronic Journal of Science Education 6 (2002)

  7. [7]

    Learning assistant supported student outcomes (lasso) study initial findings,

    Ben Van Dusen, Laurie Langdon, and Valerie Otero, “Learning assistant supported student outcomes (lasso) study initial findings,” in Physics Education Research Conference 2015 , PER Conference (College Park, MD,

  8. [8]

    https://www.physport.org/assessments/, (2015)

Show all 23 references
  1. [9]

    Quantifying crit- ical thinking: Development and validation of the physics lab inventory of critical thinking,

    Cole Walsh, Katherine N. Quinn, C. Wie- man, and N. G. Holmes, “Quantifying crit- ical thinking: Development and validation of the physics lab inventory of critical thinking,” Phys. Rev. Phys. Educ. Res. 15, 010135 (2019)

  2. [10]

    Force concept inventory,

    David Hestenes, Malcolm Wells, and Gregg Swackhamer, “Force concept inventory,” The physics teacher 30, 141– 158 (1992)

  3. [11]

    Surveying students conceptual knowledge of electricity and magnetism,

    David P Maloney, Thomas L OKuma, Curtis J Hieggelke, and Alan Van Heuvelen, “Surveying students conceptual knowledge of electricity and magnetism,” American Jour- nal of Physics 69, S12–S23 (2001)

  4. [12]

    Participation rates of in-class vs. online administration of low-stakes research-based assessments,

    Manher Jariwala, Jayson Nissen, Xochith Herrera, Eleanor Close, and Ben Van Dusen, “Participation rates of in-class vs. online administration of low-stakes research-based assessments,” in Physics Education Re- search Conference 2017 , PER Conference (Cincinnati, OH, 2017) pp. 196–199

  5. [13]

    Performance on in- class vs. online administration of concept inventories,

    Jayson Nissen, Manher Jariwala, Xochith Herrera, Eleanor Close, and Ben Van Dusen, “Performance on in- class vs. online administration of concept inventories,” i n Physics Education Research Conference 2017 , PER Con- ference (Cincinnati, OH, 2017) pp. 272–275

  6. [14]

    A comparison of computer-administered and written tests,

    David Zandvliet and Pierce Farragher, “A comparison of computer-administered and written tests,” Journal of Re- search on Computing in Education 29, 423–438 (1997)

  7. [15]

    Post-graduate student perfor - mance in supervised in-class vs.unsupervised online mul- tiple choice tests: implications for cheating and test secu - rity,

    Richard K Ladyshewsky, “Post-graduate student perfor - mance in supervised in-class vs.unsupervised online mul- tiple choice tests: implications for cheating and test secu - rity,” Assessment & Evaluation in Higher Education 40, 883–897 (2015)

  8. [16]

    The equivalence of paper-and-pencil and computer-based testing,

    Alan C Bugbee Jr, “The equivalence of paper-and-pencil and computer-based testing,” Journal of research on com- puting in education 28, 282–299 (1996)

  9. [17]

    Cheating on tests: Prevalence, detection, and implications for online testing,

    Walter M Haney and Michael J Clarke, “Cheating on tests: Prevalence, detection, and implications for online testing,” in Psychology of academic cheating (Elsevier,

  10. [18]

    Reliability, compliance, and security in web-based course assessments,

    Scott Bonham, “Reliability, compliance, and security in web-based course assessments,” Phys. Rev. ST Phys. Educ. Res. 4, 010106 (2008)

  11. [19]

    Quantum mechanics concept assess- ment: Development and validation study,

    Homeyra R. Sadaghiani and Steven J. Pol- lock, “Quantum mechanics concept assess- ment: Development and validation study,” Phys. Rev. ST Phys. Educ. Res. 11, 010110 (2015)

  12. [20]

    Val- idation and analysis of the coupled multiple re- sponse colorado upper-division electrostatics diagnosti c,

    Bethany R. Wilcox and Steven J. Pollock, “Val- idation and analysis of the coupled multiple re- sponse colorado upper-division electrostatics diagnosti c,” Phys. Rev. ST Phys. Educ. Res. 11, 020130 (2015)

  13. [21]

    Assessing student learning of newtons laws: The force and motion conceptual evaluation and the evaluation of active learn- ing laboratory and lecture curricula,

    Ronald K Thornton and David R Sokoloff, “Assessing student learning of newtons laws: The force and motion conceptual evaluation and the evaluation of active learn- ing laboratory and lecture curricula,” american Journal of Physics 66, 338–352 (1998)

  14. [22]

    Evaluating an elec- tricity and magnetism assessment tool: Brief electricity and magnetism assessment,

    Lin Ding, Ruth Chabay, Bruce Sherwood, and Robert Beichner, “Evaluating an elec- tricity and magnetism assessment tool: Brief electricity and magnetism assessment,” Phys. Rev. ST Phys. Educ. Res. 2, 010105 (2006)

  15. [23]

    An introduction to classical test t he- ory as applied to conceptual multiple-choice tests,

    Paula Engelhardt, “An introduction to classical test t he- ory as applied to conceptual multiple-choice tests,” in Getting Started in PER , Vol. 2 (2009)

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.