{"id":"0b3d53db-acfa-42de-b58d-d856216066a5","arxiv_id":"2411.14945","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A pilot study of the CTSkills app reports preliminary problem-decomposition scores across grades 4-9 from 75 students, but without validity evidence the scores only reflect the authors' answer key.","lead":"The paper introduces CTSkills, a web app that asks students in grades 4 to 9 to identify relevant objects and relations in a simple game, then scores their answers as a measure of problem decomposition. A pilot with 75 students shows scores generally rise with grade, though grade 9 dips below grade 8, and the authors argue this demonstrates a usable assessment tool.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Scoring formula in §3.3 is internally inconsistent: for |X|=2, |Y|=4 cells (Q3/Q4 Level 1 and 3) the rescaling denominator equals zero, and the optional-pair bonus from Table 2 cannot be awarded. All reported scores and grade comparisons rest on this formula.","rationale":"The reader’s weakest assumption was construct validity, but their rationale already noted that the scoring formula appears to omit the optional-pair bonus. I partially agree: the deeper issue is that the published scoring rule cannot reproduce the reported results at all. The min_score expression is not the true minimum, and for four of the twelve question-level cells the rescaling denominator is zero, so the standardised scores shown in Fig. 6 are undefined under the stated method. This is an internal mathematical inconsistency rather than a mere validity gap, and it is load-bearing because every claim about grade-related improvement is derived from these scores. However, the flaw is plausibly a typo and the paper is explicitly a pilot; if the formula is corrected and the analyses are rerun, the central trend may still survive. Therefore the reader’s CONDITIONAL verdict is appropriate, but the conditions must include a corrected scoring rule and a complete re-analysis rather than only additional external validity evidence.","tokens_in":12725,"tokens_out":6884,"duration_ms":63862,"concrete_test":"Take the raw JSON response logs (or regenerate them from the app for the 75 pilot students) and recompute all rescaled scores under three variants: (a) as published, (b) with min_score corrected to −|X|−|Y|, (c) with a +0.5 bonus per optional pair. Then rerun the grade-level ANOVA/LMM. If the mean-score trajectories or significance of grade effects change under (b) or (c), the reported age differences are an artifact of the scoring typo; if even variant (a) is undefined for Q4-L1, the published Figure 6 requires an explicit explanation of how those scores were obtained.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 defines score = |S_X| − (|X| − |S_X|) − |S_Y| and rescaling using min_score = −|X| + |Y|. Two problems. First, min_score is not the minimum achievable score: since S_X ⊆ X and S_Y ⊆ Y, the true minimum is −|X| − |Y| (select no targets, all non-targets). Second, for cells with |X|=2 and |Y|=4—Q3-L1, Q4-L1, Q3-L3, Q4-L3 in Table 2—the stated min_score equals 2, so the rescaling denominator |X| − min_score is 0, making rescaled_score undefined. Those cells nevertheless appear in Fig. 6 with nontrivial distributions, so the displayed values cannot be produced by the described method. Third, Table 2 marks optional pairs 'contributing a bonus of 0.5', but these pairs are not counted in |X| (e.g., Q4-L1 lists three pairs yet |X|=2), and the score formula contains no bonus term; selecting an optional pair yields no credit. Thus every quantitative result—means, Fig. 7, ANOVA and LMM tests—is computed from a scoring rule that is either mis-specified or undefined, directly undermining the central claim of a consistent grade-related improvement.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents CTSkills, a browser-based assessment app designed to measure students' problem decomposition skills as a component of computational thinking in K-12 education. Students play three levels of an interactive game and answer four questions per level about relevant objects, moving objects, state changes, and collisions. The authors define target and non-target sets for each question, compute per-question scores with a proposed rescaling, and average the scores across questions and levels. They report descriptive statistics, ANOVA, chi-square tests, Tukey post-hoc tests, and linear mixed-effect models on pilot data from 75 students in grades 4-9. The paper claims that the app effectively automates data collection (RQ1) and that the data show age-related improvement in decomposition skills, with a noted dip in Grade 9.","tokens_in":13001,"tokens_out":6077,"duration_ms":59488,"significance":"If the instrument were validated, this work could provide a scalable, classroom-ready tool for measuring a relatively under-explored CT skill, and the design is sensibly grounded in the decomposition framework of Rich et al. The usability testing with children and the attention to classroom data collection are strengths. However, the central measurement claim is not currently supported: the scoring formula in Section 3.3 is internally inconsistent and undefined for several question-level combinations, the author-defined target sets are not independently validated, and the grade-comparison conclusions rest on small and uneven samples. At this stage the contribution is best described as a feasibility prototype and data-collection pipeline, not a validated measure of problem decomposition.","major_comments":[{"comment":"The rescaling formula is undefined for several reported cells. With score = |S_X| - (|X|-|S_X|) - |S_Y| and min_score = -|X| + |Y|, the denominator |X| - min_score equals zero whenever |X|=2 and |Y|=4, which occurs for Q3-L1, Q4-L1, Q3-L3, and Q4-L3 according to Table 2. Yet Fig. 6 shows nontrivial score distributions for these cells. Either the paper does not report the scoring rule actually used, or the displayed scores cannot be produced by the described method. Because every aggregate and inferential result depends on these scores, the quantitative claims in Section 4 are not reproducible from the manuscript as written.","section":"Section 3.3"},{"comment":"The stated min_score is not the minimum achievable score. Since S_X is a subset of X and S_Y is a subset of Y, the lowest possible value of score is -|X|-|Y| (select no targets and all non-targets), not -|X|+|Y|. For some configurations the stated min_score even exceeds the maximum score |X|, so the rescaling can produce negative or inverted values. The claim that scores are standardised to a 0-5 range is therefore not correct for the stated formula, and all comparisons involving rescaled scores are affected.","section":"Section 3.3"},{"comment":"Table 2 is internally inconsistent with the scoring model. For Q4-L1 the table lists three target pairs (including the optional Apple spoiled red and Grass pair) but reports |X|=2; for Q4-L3 it lists five target pairs but reports |X|=2. The optional-pair bonus of 0.5 mentioned in the table is absent from the score formula in Section 3.3, so a student selecting an optional pair receives no credit under the stated formula, or the target counts used in the formula are wrong. Either way, the scoring rule does not match the described items.","section":"Appendix A.2, Table 2"},{"comment":"The central construct-validity question is unaddressed. The target and non-target sets are defined by the authors, and the paper itself acknowledges in Section 3.2 that the tree could reasonably be considered relevant in Q1 and admits in Section 5.1 that creative solutions such as grouping apples with baskets are difficult to quantify. No expert panel, inter-rater reliability check, item-level analysis, or comparison with an external decomposition measure is provided. Until such evidence is supplied, the score is best interpreted as agreement with the authors' answer key rather than as a validated measure of problem decomposition skill.","section":"Sections 3.2 and 5.1"},{"comment":"The conclusion of a 'consistent improvement in task performance as students progress through grades' is not consistent with the reported results. The post-hoc test shows Grade 9 scoring significantly lower than Grade 8 (MD = -0.6500, p = 0.0175), and the sample sizes are highly uneven: Grade 7 has only 4 students while Grade 9 has 24. The LMM also attributes only modest variance to grade (Variance = 0.061, SD = 0.247). The discussion should temper the developmental-claim language and address the small and unbalanced grade samples explicitly.","section":"Section 4 and Table 1"}],"minor_comments":[{"comment":"The text contains a duplicated phrase: 'potentially potentially indicating different difficulty levels.'","section":"Section 4, Fig. 6 paragraph"},{"comment":"The text refers to 'grade 8 to 11' when describing selection rates, but the study only covers grades 4-9; this appears to be a typo and should be corrected.","section":"Section 4, Fig. 5 discussion"},{"comment":"The value of max_scaled is never stated; the paper should specify that it is 5 (or whatever value was used) so the rescaling is reproducible.","section":"Section 3.3"},{"comment":"The caption and surrounding text describe the axes inconsistently: the text says difficulty levels are on the x-axis and score categories on the y-axis, but the figure layout appears to show score categories on the y-axis and levels on the x-axis. Please align the caption with the actual figure.","section":"Fig. 6"},{"comment":"The statement that 'the app proved effective in automating data collection' is presented as a finding, but the paper reports no usability or completion metrics (e.g., task completion rates, time-on-task, or error logs) to support this claim; 'feasible' would be a more accurate descriptor for the pilot evidence.","section":"Section 5, RQ1"}],"recommendation":"major_revision","confidential_remarks":"The scoring-formula problems are severe enough that the reported statistics cannot be taken at face value, but they are fixable by redefining the score, recalculating the rescaling, and re-running the analyses. The larger issue is that the paper currently frames a pilot feasibility study as a validated measurement contribution; the authors should either add validation evidence or reframe the claims. The paper may also be better suited to a venue or section that explicitly publishes instrument-development or pilot-study reports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know one thing up front: the scoring formula in §3.3 is broken, and it breaks the paper's central quantitative claims. The paper defines min_score = −|X| + |Y|, but the true minimum score is −|X| − |Y| (select no targets, all non-targets). For the Q3/Q4 cells with |X|=2, |Y|=4, the stated min_score equals 2, making the rescaling denominator |X| − min_score = 0. Those cells still show score distributions in Fig. 6, so the displayed values cannot come from the described method. The optional-pair bonus in Table 2 is also absent from the formula and from |X|, so selecting an optional pair gets no credit. Every mean, figure, ANOVA, and LMM is built on this misspecified rule.\n\nThat is a shame, because the underlying idea is worth pursuing. Decomposition is genuinely under-assessed in K-12 CS, and the app is a plausible new instrument that operationalizes Rich et al.'s framework in an interactive form. The authors deserve credit for building a working prototype, running a classroom pilot with 75 students across grades 4–9, and being transparent about the study's limitations. The task design—repeated scenarios with increasing complexity—is thoughtful.\n\nBut the soft spot is load-bearing. Beyond the formula, the target/non-target sets are author-defined with no external validity evidence, no reliability data, and no expert panel or inter-rater check. The paper acknowledges that the tree in Q1 is arguably relevant, yet the scoring key treats it as non-target. That is the circularity problem: the 'measurement' reduces to matching the authors' choices. The Grade 9 dip is based on one class of 20 students and is overinterpreted in the discussion. The sample is small and clustered by class, so the grade effects are fragile.\n\nWho is this for? Researchers working on CT assessment, especially those building automated instruments. The paper is a useful example of how to design a decomposition task, but it is not yet a validated assessment. My recommendation: send it to peer review, because the gap is real and the prototype is worth salvaging, but the authors must fix the scoring rule, re-analyze the data, and add at least some validity evidence (e.g., expert review, comparison with existing measures). As submitted, the quantitative conclusions should not be cited.","headline":"Promising prototype for assessing decomposition is undermined by a scoring formula that is internally inconsistent, so the reported grade trends cannot be trusted as they stand.","tokens_in":13573,"tokens_out":1785,"would_cite":false,"duration_ms":19275,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The CTSkills app measures problem decomposition in grades 4–9 by scoring students' selections of objects and relations, and its pilot data show performance improving with grade and dipping in grade 9.","keywords":["computational thinking","problem decomposition","assessment tool","K-12 education","pilot study","substantive decomposition","relational decomposition","scoring methodology"],"falsifier":"Re-analyze the Level 1, Q1 logs: count students who selected the tree while omitting an item the scoring rule counts as a target; if a substantial number of students who score high on Q2–Q4 did so, then the scoring rule penalizes a defensible decomposition rather than measuring the skill.","tokens_in":12511,"feed_emoji":"🧩","tokens_out":7297,"duration_ms":64743,"temperature":0.7,"pith_summary":"This paper tries to establish that problem decomposition—a component of computational thinking that is often named but rarely assessed—can be measured directly with a browser-based app in K-12 classrooms. The app presents three game-like scenarios and asks students to drag the relevant objects and relations into a solution field; the score rewards correct selections and penalizes missed or irrelevant ones. The paper reports on a pilot with 75 students in grades 4–9, finding a general improvement with grade, a dip for grade 9, and no gender difference. If the app measures what it claims, it would give schools and researchers an automated way to collect decomposition-skill baselines across a wide age range, a step toward the baseline the field currently lacks.","feed_headline":"App scores students' problem-splitting skill in grades 4–9","feed_subtitle":"In a 75-student pilot, scores climb with age, with a grade-9 dip; boys and girls score alike.","key_machinery":"The machinery is a formal target/non-target scoring scheme. For each question and level the paper defines a set $X$ of target items or item pairs and a set $Y$ of non-target items or pairs; from the student's selected sets $S_X$ and $S_Y$ the score is $$score = |S_X| - (|X| - |S_X|) - |S_Y|,$$ which rewards correct picks, penalizes missed targets, and penalizes false positives. The raw score is rescaled to a 0–5 range using the achievable minimum, then averaged over questions and levels. This scheme operationalizes decomposition as boundary-drawing accuracy, and the three sceneries are designed so the same objects reappear in different contexts, letting pattern recognition contribute to later levels.","core_discovery":"The central claim is that substantive and relational decomposition—picking out the objects that matter in a problem scenario and identifying how those objects change or collide—can be assessed automatically through a structured drag-and-drop interface. Each of the three levels is followed by four questions: which objects are relevant, which are moving, which objects change into which, and which collide. The paper's target/non-target scoring formula treats decomposition as accuracy in drawing the boundary between what belongs to the problem and what does not. In the pilot data, the paper reports a consistent improvement in task performance across grades, a statistically significant grade-9 drop relative to grade 8, and no significant effect of gender, and it interprets these as evidence that the app can automate data collection for a future decomposition proficiency baseline.","pith_inferences":["The current target sets may be too narrow; a validity study comparing app scores with open-ended decomposition explanations could reveal that some 'incorrect' selections are legitimate alternative decompositions.","Because the interface asks students to select pre-listed objects, the test measures recognition of relevant objects rather than the ability to generate the decomposition from scratch; a generative-response version might measure a different facet of the skill.","The cross-sectional age trend conflates age with cohort and curriculum; longitudinal retesting of the same students would be needed to claim individual development.","The format could transfer to non-programming subjects: any scenario with objects and relations could be scored the same way, making decomposition assessment domain-general."],"forward_implications":["Classrooms could run the same decomposition assessment across grades 4–9 and compare performance year to year without manual coding analysis.","The automated scoring formula gives researchers a compact outcome measure for studying how decomposition skill develops with age.","The pilot's null gender difference suggests decomposition skill is distributed evenly across genders in this sample, supporting inclusive computational thinking education.","The grade-9 dip points to a transitional period where secondary-school task complexity may outpace students' decomposition strategies, a target for intervention.","The same scoring scheme can be extended to functional decomposition once the app adds tasks that require grouping objects into reusable classes or procedures."],"supporting_citations":[{"why":"Supplies the substantive, relational, and functional decomposition categories that define the app's questions.","marker":"[23]"},{"why":"Provides the K-8 decomposition learning trajectories and consensus goals that motivate the level design and interpretation of progression.","marker":"[22]"},{"why":"Provides the foundational definition of computational thinking that frames decomposition as a key competence.","marker":"[32]"},{"why":"Reviews existing computational thinking assessments to establish the gap in direct decomposition measurement that the app addresses.","marker":"[28]"},{"why":"Represents the general computational thinking test used as a contrast; it lacks detailed decomposition sub-scores.","marker":"[24]"},{"why":"Shows an automatic artifact-based decomposition assessment that the app's process-oriented assessment contrasts with.","marker":"[19]"},{"why":"An item test that assesses decomposition, providing a prior example of decomposition assessment beyond artifacts.","marker":"[12]"}],"fun_headline_variants":["App scores problem-splitting skills in grades 4–9","Problem-decomposition app: skills peak in grade 8","New tool measures decomposition: no gender gap in grades 4–9","Drag-and-drop app automates decomposition scoring for students","75-student pilot: decomposition scores dip in grade 9"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The scoring depends on the author-defined lists of which objects and relations count as correct; if a student's answer reflects a legitimate alternative decomposition, such as counting the tree as relevant or grouping apples with the basket, the score will misclassify their ability.","fun_headline_variants_meta":{"raw":{"variants":["App scores problem-splitting skills in grades 4–9","Problem-decomposition app: skills peak in grade 8","New tool measures decomposition: no gender gap in grades 4–9","Drag-and-drop app automates decomposition scoring for students","75-student pilot: decomposition scores dip in grade 9"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000275,"raw_usage":{"total_tokens":1614,"prompt_tokens":887,"completion_tokens":727,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":641}},"tokens_in":503,"tokens_out":727,"duration_ms":7464,"temperature":1.0,"reasoning_tokens":641,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:40:29.844203+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze the Level 1, Q1 logs: count students who selected the tree while omitting an item the scoring rule counts as a target; if a substantial number of students who score high on Q2–Q4 did so, then the scoring rule penalizes a defensible decomposition rather than measuring the skill.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Represents the general computational thinking test used as a contrast; it lacks detailed decomposition sub-scores."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows an automatic artifact-based decomposition assessment that the app's process-oriented assessment contrasts with."}],"review_version":1}