{"id":"cafb2cfc-d6f1-4e08-9941-f31c3d8d36e1","arxiv_id":"2411.13382","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Students in a first-year mechanics course perceived weekly smartphone-based experimental tasks as comparable to standard recitation tasks and more positive than programming tasks on several affective and learning-perception measures.","lead":"This paper evaluates a semester-long introduction of smartphone-based physics experiments as weekly homework tasks in a large first-year mechanics course, comparing student perceptions of these tasks with programming tasks and traditional problem sets. It finds that students rated the smartphone experiments positively, generally similar to standard tasks and higher than programming tasks on several engagement measures.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Exp-over-Pro advantage may hinge on the outlier task Pro 1; excluding it could erase the significant differences, so the comparative claim is not yet robust.","rationale":"I agree with the reader's conditional verdict. The paper's feasibility claim—that weekly smartphone experiments can be implemented and are generally well-received—is well supported by high participation, positive ratings, and transparent reporting, including open data and explicit acknowledgment of limitations. The comparative claim is the weak point, and the most concrete threat is the asymmetry between nine Exp tasks and three Pro tasks, one of which is a clear outlier. The paper acknowledges this possibility but does not test robustness to removing Pro 1. The Rec comparison is also confounded by timing and single retrospective measurement, but the abstract's wording for Rec is already modest ('comparable to, or only partly below'), so the Pro comparison carries more weight in the headline claim. I therefore identify the Pro 1 sensitivity analysis as the single check that would settle whether the 'tend to outperform' claim is genuine. Since the reader already assigned CONDITIONAL, my analysis does not change the verdict; it specifies the test that the condition should require.","tokens_in":31745,"tokens_out":5209,"duration_ms":57238,"concrete_test":"Re-run all Exp-vs-Pro comparisons from Sec. IV.A (overall rating, time, goal clarity, feasibility at home, use of technologies) and Sec. IV.B (curiosity, interest, reference to reality, disciplinary authenticity, experience of competence) with Pro 1 excluded, leaving Pro 2 and Pro 3 as the Pro sample. If the significant Exp>Pro results (goal clarity, reference to reality, experience of competence) cease to be significant or reverse, the abstract's 'tend to outperform' should be qualified as dependent on one specific task; if they persist, the claim is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparative claim—that Exp tasks 'tend to outperform' Pro tasks—rests on aggregate comparisons of nine Exp tasks against only three Pro tasks. Of those three, Pro 1 was rated too difficult by 43% of students and had insufficient instructions by 25% (Fig. 9), and the open-text analysis attributes most negative Pro feedback to Pro 1 (29% of complexity statements, 28% of instruction statements, Table IV). Because the Pro sample is so small, a single flawed task strongly influences the mean Pro scores used in the Mann-Whitney-U tests (Sec. IV.A) and in the matched affective comparisons (Sec. IV.B). The paper itself notes that 'the too-difficult task Pro 1 likely played a significant role' (Sec. V.B.1), but no sensitivity analysis is reported. Similarly, the Rec comparison depends on a single retrospective rating collected days before the exam (Sec. V.D), which may inflate Rec scores; however, the abstract's 'comparable to, or only partly below' Rec is already hedged, so the Pro comparison is the more load-bearing issue. Without excluding Pro 1 or modeling task as a random effect, the claimed Exp-over-Pro advantage cannot be distinguished from the effect of one poorly calibrated programming task.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports an exploratory field study of weekly smartphone-based experimental tasks (Exp) in a first-year introductory mechanics course at RWTH Aachen, benchmarked against three Python programming tasks (Pro) and long-established standard recitation tasks (Rec). Nine Exp tasks and three Pro tasks were embedded in weekly exercise sheets, and data come from 14 online surveys: twelve weekly task surveys with participation ranging from 188 to 41 students, plus two comparative surveys with 108 and 78 respondents. The analyses address students' perceptions of learning (overall rating, goal clarity, feasibility at home, time spent, difficulty, open-text feedback) and affective responses (curiosity, interest, authenticity, experience during tasks, perceived effectiveness), together with a predictor analysis (RQ3). The main conclusions are that Exp tasks were generally well received, that Exp tasks outperformed Pro tasks on a subset of perception and affective variables (goal clarity, reference to reality, experience of competence), and that Exp tasks were comparable to or slightly below Rec tasks on most affective variables, with a higher experience of competence for Exp tasks.","tokens_in":31935,"tokens_out":5059,"duration_ms":48788,"significance":"If the comparative results are accepted, this is a useful contribution to the sparse evaluation literature on smartphone experiments in university physics: it provides a semester-long implementation, openly available data, detailed instruments with factor analyses, and a three-way comparison that goes beyond proof-of-concept. The paper's explicit treatment of limitations and its disclosure of the phyphox developer conflict of interest are strengths. However, the Exp-vs-Pro comparative claim, which is central to the abstract, is less robust than the descriptive feasibility claim, and the paper's own limitations section acknowledges the main threats. The recommendation therefore depends on whether the comparative claims can be reanalyzed or appropriately hedged.","major_comments":[{"comment":"The abstract's statement that the experimental tasks 'tend to outperform the programming tasks in terms of perceptions of learning with the tasks and affective responses' is stronger than the results in Table VI support. Of the twelve Exp-vs-Pro comparisons listed, only goal clarity, reference to reality, and experience of competence are statistically significant after the corrections reported; overall task rating, feasibility at home, time spent, use of technologies, curiosity, interest, disciplinary authenticity, and the short curiosity/interest and autonomy/creativity scales are not. In particular, Table X shows curiosity with Bonferroni-corrected p = 0.061 and interest with p = 0.63, so the word 'outperform' should be reserved for the specific subset of scales on which the difference was significant.","section":"Abstract; Sec. IV.A; Table VI"},{"comment":"The Exp-vs-Pro aggregate comparisons are not robust to the influence of Pro 1, and no sensitivity analysis is reported. Pro 1 is one of only three programming tasks; 43% of students rated it too difficult and 25% reported insufficient instructions (Fig. 9), and the open-text analysis attributes 29% of negative complexity statements and 28% of negative instruction statements for Pro tasks to this task (Table IV). The paper itself states in Sec. V.B.1 that 'the too-difficult task Pro 1 likely played a significant role in this outcome,' but the aggregate Mann-Whitney tests in Sec. IV.A and the affective comparisons in Sec. IV.B do not examine whether the reported Exp-over-Pro differences survive after excluding Pro 1 or after modeling task as a random effect. Without such an analysis, the claim that the Exp format generally outperformed the Pro format cannot be distinguished from the effect of one poorly calibrated task.","section":"Sec. IV.A, Fig. 9; Sec. V.B.1; Table IV"},{"comment":"The decision to combine Exp-task scores from Com 1 and Com 2 despite a significant difference on disciplinary authenticity (Z = 3.14, p_B = 0.008) is not fully justified. The authors attribute the difference to context and timing, but for students who responded to both surveys the average of the two measurements conflates the two comparison conditions and the two measurement times. Reporting the Com 1 and Com 2 analyses separately, or including survey as a factor, would make the combined analysis easier to evaluate.","section":"Sec. III.E.2; Sec. IV.B"},{"comment":"The small and declining samples for the comparative analyses limit the strength of the conclusions. Participation fell from 188 (Exp 1) to 41 (Exp 9), and the three-way Friedman tests in Table X are based on N between 31 and 45. The non-significant '=' entries in Table VI should therefore be read as 'no significant difference detected' rather than equivalence; this is mostly handled in the text but the table itself invites the stronger reading. In addition, pairwise exclusion means the different rows in Table X are based on different subsets of students, which further complicates cross-scale comparisons.","section":"Sec. III.B; Table X"}],"minor_comments":[{"comment":"The sentence 'totaling in 313 codes' is unclear; it is not specified whether this refers to survey responses, matched participants, or something else.","section":"Sec. III.B"},{"comment":"The 'Mean Pro 1-3' and 'Mean Exp 1-9' bars average over tasks with different sample sizes per task; annotating the figures with the N range or using weighted means would improve interpretability.","section":"Fig. 9 and Fig. 19"},{"comment":"The comparison with Ref. [55] mixes matched-sample effect sizes from this study with independent-sample effect sizes from the earlier study; the figure caption should state that the confidence intervals for the reference data are not based on matched samples and may not be directly comparable.","section":"Sec. V.B, Fig. 15"},{"comment":"The statement that items are presented in German with English translations as supplementary material would benefit from an explicit reference to the supplementary material file in the main text.","section":"Sec. III.C"},{"comment":"The row for 'Use of technologies' indicates it was measured in weekly surveys and Com 1, but the text in Sec. IV.A reports only the weekly-survey comparison with N = 87; clarify where the N comes from.","section":"Table II; Sec. IV.A"}],"recommendation":"major_revision","confidential_remarks":"The main risk to the paper's central claim is the Exp-vs-Pro comparison. If the authors can add the sensitivity analysis (e.g., excluding Pro 1 or modeling task as a random effect) and adjust the abstract and conclusions accordingly, the paper would be suitable for publication. The phyphox developer conflict of interest is disclosed, but it should be given more prominence when interpreting self-report perception data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is worth your time for the feasibility data, not for the comparative claims. It reports one of the few semester-long implementations of smartphone-based experimental tasks in a large first-year mechanics course, with nine tasks, weekly surveys peaking at 188 respondents, and a transparent write-up. The descriptive finding is solid: students found the tasks manageable, the instructions mostly clear, and the open-text responses give a balanced picture of what worked and what did not. The instrument validation and the open data are real assets. The authors also disclose that four of them develop phyphox, which is the right call.\n\nThe soft spot is the Exp-over-Pro comparison. The programming arm is only three tasks, and one of them (Pro 1) was rated too difficult by 43% of students and generated most of the negative open-ended comments. The paper acknowledges this, but no sensitivity analysis is reported. With only three Pro tasks, removing Pro 1 would almost certainly erase the significant differences in goal clarity and experience of competence, which are the main basis for the abstract's 'tend to outperform' phrase. The Rec comparison is also fragile—measured once, days before the exam, and with students generalizing across all Rec tasks. The authors hedge this in the abstract, but the comparative claims in the discussion are stronger than the data can carry.\n\nThe predictor analysis (RQ3) is too underpowered to say much, and the many tests without clustering should be read as exploratory. These are limitations the authors mostly acknowledge, and they do not undermine the core feasibility contribution.\n\nMy take: this deserves a serious referee. The field needs more systematic, semester-long evaluations of smartphone experiments, and this paper provides a solid foundation. The authors should be asked to either soften the comparative conclusions into descriptive feasibility claims or add a sensitivity analysis that excludes Pro 1 and treats task as a random effect. With that revision, it would be a useful citable study, especially for PER colleagues working on technology-enhanced tasks.\n\nRecommendation: send it for review, with a request for revised comparative claims.","headline":"A serious semester-long feasibility study of smartphone experiments in a large intro course, but the comparative claims rest on a fragile three-task programming comparison.","tokens_in":32540,"tokens_out":2809,"would_cite":true,"duration_ms":29855,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that weekly smartphone-based experimental tasks can be integrated into a large first-year mechanics course, that students perceive them as well-suited homework, and that these tasks outperform newly introduced…","keywords":["smartphone-based experiments","introductory mechanics","recitation tasks","programming tasks","student perceptions","affective responses","phyphox","physics education"],"falsifier":"A matched comparison in which the three task types are newly designed with equal polish, comparable difficulty, and identical exam weight, and in which recitation tasks are rated mid-semester rather than days before the final exam, would settle whether the reported ordering persists; if experimental tasks then fall below programming tasks on the affective scales, or no longer beat recitation tasks on experience of competence, the central comparative claim would not generalize.","tokens_in":31549,"feed_emoji":"📱","tokens_out":3495,"duration_ms":42154,"temperature":0.7,"pith_summary":"The paper tries to establish that smartphone-based experimental tasks are a viable format for weekly homework in introductory university physics, not just a novelty for demonstration or lab courses. Nine such tasks were embedded in a first-year mechanics course alongside three Python programming tasks and the usual textbook recitation problems, and students' perceptions were surveyed throughout the semester. The central finding is that students rated the experimental tasks as equal to or better than the programming tasks on nearly all learning-perception and affective measures, and as comparable to, or only partly below, the long-established recitation tasks. If this holds, instructors can diversify homework with low-cost, at-home experiments that give students first-hand data collection and a genuine sense of experimental competence without sacrificing student acceptance.","feed_headline":"Phone experiments beat programming tasks in physics homework","feed_subtitle":"Nine weekly smartphone lab tasks earned student ratings equal to or better than Python tasks and close to classic problems.","key_machinery":"The carrying mechanism is an evaluation design rather than a physical apparatus: twelve short weekly surveys measuring perceptions of learning with each individual experimental and programming task, plus two comparative surveys in which students rated affective responses to experimental tasks against programming tasks and against standard recitation tasks. The survey items were consolidated through exploratory factor analysis into scales such as goal clarity, feasibility at home, curiosity, interest, authenticity (split into reference to reality and disciplinary authenticity), experience during the tasks (competence, curiosity/interest, autonomy/creativity), perceived educational effectiveness, autonomy, and linking to the lecture. Comparisons use Bonferroni-corrected nonparametric tests, with effect sizes and confidence intervals benchmarked against a prior study of video-based analysis tasks. The tasks themselves rely on the phyphox app—a smartphone app that turns built-in and external sensors into physics measurement tools—together with preset experiment configurations and lent equipment such as sensor boxes, wooden wheels, and balls.","core_discovery":"In an exploratory field study in a first-year mechanics course, nine smartphone-based experimental tasks using the phyphox app and external sensor boxes were implemented as weekly graded exercises, alongside three newly designed Python programming tasks and the course's traditional pen-and-paper recitation tasks. Students' responses show an overall ordering of recitation tasks at the top for curiosity, interest, perceived affective effectiveness, and linking to the lecture, with experimental tasks close behind and programming tasks generally lowest. The experimental tasks significantly outperformed the programming tasks in goal clarity, reference to reality, and experience of competence, and were rated significantly higher than the recitation tasks in experience of competence. On most other scales, including overall task rating, feasibility at home, time spent, use of technologies, autonomy, and disciplinary authenticity, the experimental tasks were statistically indistinguishable from the programming tasks or the recitation tasks. The authors interpret this as evidence that smartphone-based experimental tasks can be successfully integrated into undergraduate teaching and can enrich traditional recitation work, while noting that the programming tasks, one of which was rated too difficult by 43% of students, may not yet represent the format's potential.","pith_inferences":["An implication the paper leaves implicit is that the experimental-versus-recitation gap in curiosity, interest, and perceived affective effectiveness may be partly a timing artefact: the comparison survey ran days before an exam covering only recitation tasks, so students may have rated textbook problems as more relevant and engaging than they would have earlier in the semester.","A testable extension would be to compare matched task types with equivalent novelty: designing programming tasks as carefully calibrated as the experimental tasks, and measuring responses to recitation tasks at several points during the semester, would show whether the reported ordering Rec ≥ Exp ≥ Pro reflects the task formats themselves or the maturity of their implementation.","The study establishes perceived feasibility and positive affect, not learning gains; a natural next step is a performance-based assessment of experimental skills, measurement uncertainty, or conceptual understanding to see whether the favorable perceptions correspond to measurable outcomes.","The correlation between better high-school grades and a more favorable impression of the experimental tasks suggests that higher-achieving students may benefit most from this format, implying that differentiated support could be needed for students with weaker preparation."],"forward_implications":["Weekly smartphone-based experimental tasks can be run at scale in a large introductory course: logistics of lending equipment, grading submissions, and tying tasks to exam prerequisites were overcome for over a hundred student groups.","Instructors can expect first-iteration experimental tasks to be perceived as clearly more competence-building than first-iteration programming tasks, even when overall ratings are similar.","Newly introduced programming tasks need careful difficulty calibration: one task was rated too difficult by 43% of students and was the main drag on the programming-task format's affective scores.","Students are likely to judge new experimental tasks as less curiosity-provoking and less lecture-linked than polished, exam-aligned textbook problems, at least when the comparison is made near the final exam.","Affective perceptions of smartphone experiments are not automatically high simply because the technology is familiar; task design, difficulty, and perceived purpose appear to drive student responses."],"supporting_citations":[{"why":"Previous implementation and evaluation of smartphone-based experimental exercises in university physics courses, giving this study a direct predecessor and comparison point.","marker":"[33]"},{"why":"Quasi-experimental follow-up study on smartphone-based experimental exercises, whose null result on motivation and curiosity this study's findings both extend and contrast with.","marker":"[34]"},{"why":"Study of video-based analysis tasks using the same affective-response scales, providing the comparative dataset and effect-size benchmarks used in the discussion.","marker":"[55]"},{"why":"Pilot of smartphone experiments in introductory physics courses at two universities, supplying prior evidence on feasibility and on the instructional support students need.","marker":"[27]"},{"why":"Source of the task-quality, technology-use, and open-feedback items used in the weekly surveys of the experimental and programming tasks.","marker":"[53]"},{"why":"Evaluation of a lab course using smartphones and Arduinos, used as a reference for how smartphone-based activities affect student outcomes and beliefs in undergraduate labs.","marker":"[7]"}],"fun_headline_variants":["Phone experiments beat programming in physics homework","Smartphone labs outrank coding tasks in mechanics","Physics students prefer phone tasks to programming","Phone-based physics tasks rival classic recitation","Smartphone experiments trump Python in student ratings"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparative conclusions assume that the particular experimental, programming, and standard recitation tasks are fair representatives of their task types, meaning the observed differences are due to the task format itself rather than to differences in novelty, difficulty, instruction quality, or when the surveys happened to be administered.","fun_headline_variants_meta":{"raw":{"variants":["Phone experiments beat programming in physics homework","Smartphone labs outrank coding tasks in mechanics","Physics students prefer phone tasks to programming","Phone-based physics tasks rival classic recitation","Smartphone experiments trump Python in student ratings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000511,"raw_usage":{"total_tokens":2533,"prompt_tokens":1040,"completion_tokens":1493,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":656,"completion_tokens_details":{"reasoning_tokens":1428}},"tokens_in":656,"tokens_out":1493,"duration_ms":13572,"temperature":1.0,"reasoning_tokens":1428,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:28:50.421466+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A matched comparison in which the three task types are newly designed with equal polish, comparable difficulty, and identical exam weight, and in which recitation tasks are rated mid-semester rather than days before the final exam, would settle whether the reported ordering persists; if experimental tasks then fall below programming tasks on the affective scales, or no longer beat recitation tasks on experience of competence, the central comparative claim would not generalize.","supporting_citations":[{"cited_title":"Bernardini, M","cited_arxiv_id":null,"evidence_quote":"Previous implementation and evaluation of smartphone-based experimental exercises in university physics courses, giving this study a direct predecessor and comparison point."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Quasi-experimental follow-up study on smartphone-based experimental exercises, whose null result on motivation and curiosity this study's findings both extend and contrast with."},{"cited_title":"Hochberg, J","cited_arxiv_id":null,"evidence_quote":"Study of video-based analysis tasks using the same affective-response scales, providing the comparative dataset and effect-size benchmarks used in the discussion."},{"cited_title":"Troendle, Mapping physics students in Europe (2004)","cited_arxiv_id":null,"evidence_quote":"Pilot of smartphone experiments in introductory physics courses at two universities, supplying prior evidence on feasibility and on the instructional support students need."},{"cited_title":"Deslauriers, L","cited_arxiv_id":null,"evidence_quote":"Source of the task-quality, technology-use, and open-feedback items used in the weekly surveys of the experimental and programming tasks."},{"cited_title":"Chen, H.-C","cited_arxiv_id":null,"evidence_quote":"Evaluation of a lab course using smartphones and Arduinos, used as a reference for how smartphone-based activities affect student outcomes and beliefs in undergraduate labs."}],"review_version":1}