REVIEW 4 major objections 5 minor 30 references
Measuring Computational Thinking Self-Efficacy (CT-SEI): Instrument development and preliminary evaluation
T0 review · 4 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Computational thinking self-efficacy can be measured with a 27-item, context-bound survey whose responses split into two factors: creating a solution and evaluating it.
desk verdict Useful instrument development with a genuinely good item pool, but the two-factor CFA claim rests on post-hoc model fitting with n=70 and is oversold in the abstract. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The instrument's three-part structure: a context-setting assignment, one overall task self-efficacy question, and 27 'I can' statements covering CT sub-skills on a 0-10 scale. The statistical machinery is principal component analysis with oblimin rotation for item reduction, followed by confirmatory factor analysis on a separate 70-response sample. The two-factor CFA model (create versus evaluate) is the load-bearing result that translates the theoretical skill list into a measurable two-dimensional belief structure.
What would settle it
Administer the final 27-item instrument to a new, larger sample of higher-education students in a preregistered confirmatory factor analysis. If the create/evaluate two-factor model shows RMSEA above 0.06 or CFI/TLI below 0.95, or if the two factors correlate near 1, the claimed structure fails to reproduce. A simpler check: low test-retest correlations on the same students would indicate the instrument does not measure a stable self-efficacy belief.
Extended reading notes
Core claim
The central claim is that CT self-efficacy in higher education can be measured as a context-dependent belief, and that when measured with 'I can' statements tied to a Python assignment, the 27 retained items do not reproduce the five theoretical CT skills. Instead, the best-fitting model groups the items into two categories: creating the solution (abstraction, algorithmic thinking, decomposition, and generalization items) and evaluating the solution (evaluation items), with reported fit statistics of RMSEA = 0.050, CFI = 0.972, TLI = 0.969. The paper also reports that women in the sample rated their CT self-efficacy lower than men on most items, a pattern consistent with broader findings in
Load-bearing premise
The load-bearing premise is that a two-factor structure fitted with only 70 respondents and with models chosen after inspecting the data describes a stable structure in the broader student population rather than a small-sample artifact.
Editorial extensions
If this is right
- Educators can use the 27 items to profile a student's CT self-efficacy separately for creating and evaluating solutions, rather than relying on a single overall score.
- The scale can serve as an outcome measure in intervention studies, including the authors' planned comparison of worked examples versus practice problems.
- Because the instrument is context-dependent, the same items can be adapted to other CT tasks by changing the assignment text, though validity in new contexts still needs testing.
- The reported gender differences suggest that CT instruction may need to address self-efficacy directly, especially for women in computing courses.
- The CFA results imply that separate self-efficacy beliefs for algorithmic thinking, decomposition, abstraction, and generalization may not be distinguishable in this population.
Reading between the lines
- If the two-factor create/evaluate structure replicates, it would suggest that students' self-efficacy does not track the five-skill taxonomy that guided item construction; the taxonomy may describe competence better than perceived competence.
- The gender differences here could reflect measurement non-invariance across groups or particularities of the sample; a formal invariance test would show whether the scale measures the same construct for men and women.
- Because the CFA sample is small and models were selected after inspecting the data, a preregistered replication on a new sample, including test-retest reliability, is the natural next test of whether the two-factor structure is stable.
- The instrument's link to actual behavior is untested; correlating CT-SEI scores with performance on the programming task or course grades would show whether self-reported confidence predicts the outcomes that matter.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents the development and preliminary validation of the Computational Thinking Self-Efficacy Instrument (CT-SEI), a 27-item scale measuring self-efficacy for abstraction, algorithmic thinking, decomposition, evaluation, and generalization in the context of a Python programming assignment. Ninety-one candidate items were reduced by expert review to 54, then by principal component analysis on 200 respondents to 27. A confirmatory factor analysis on the remaining 70 respondents is used to claim the items fall into two categories: creating the solution and evaluating the solution. The paper also reports gender differences favoring men on assignment-level self-efficacy and some items.
Significance. If the factor structure were supported, the CT-SEI would be a useful, theory-aligned instrument for computing education research. The paper has genuine strengths: item wording follows Bandura's self-efficacy principles, content validity is addressed via independent expert review, the assignment context is clearly specified, and the final item set is reported in Table 4. However, the central confirmatory claim is not supported by the reported analyses: the two-category structure was selected post hoc and estimated on a very small sample. The contribution at this stage is a carefully constructed item pool and a set of hypotheses about its structure, not a validated two-factor scale.
major comments (4)
- [§3.2.3 and Table 3] The headline two-category result is not a confirmatory finding. Model 1 was inadmissible, Model 2 was formed by collapsing algorithmic thinking and decomposition, and Model 3 was added 'based on an inspection of the items included in the final item set' after seeing the earlier results. All models are fit to the same n=70 sample. At this sample size, χ² has low power, and a model selected after inspecting the data can reproduce plausible-looking RMSEA/CFI/TLI values even if it is overfit. The abstract's 'showed the items can be divided into two categories' overstates what the analysis establishes. Please re-estimate the structure on independent data, or recast the analysis as exploratory and temper the conclusion accordingly. The existing limitation statement in §3.4 acknowledges the sample size but does not address the post-hoc selection problem.
- [§3.2.3 and Table 4] The model described as a 'two-factor' model is not a two-factor CFA at the item level. In Model 3, the 'Create solution' factor contains three sub-factors (Abstraction, Generalization, and Algorithmic Thinking & Decomposition), while 'Evaluate solution' contains only the Evaluation items. This is a higher-order structure, not a direct division of 27 items into two factors. The abstract and Section 3.2.3 use 'two categories' loosely. In addition, the text says the first factor consists of 'abstraction, generalization, and a combination of algorithmic thinking and abstraction', but Table 4 labels the combined factor 'Algorithmic Thinking & Decomposition'; this discrepancy should be fixed. Please specify the exact model (first-order factors, second-order loadings, constraints) that produced the fit indices in Table 3. A second-order factor with only one first-order factor (Evaluation) is no
- [§3.2.2] The item-reduction PCA was run separately for each of the five CT skills on its own item subset, rather than on the full 54-item set. This procedure cannot reveal cross-loadings or assess whether the five theoretically defined components are empirically distinct. It therefore partly entrenches the five-factor theory instead of testing it, and it makes the later CFA results harder to interpret. Given the strong interfactor correlations implied by the inadmissible Model 1 and the eventual collapsing of AT and DC, an exploratory factor analysis on all 54 items would be a more appropriate item-selection step. Please report the full PCA results (eigenvalues, loadings, cross-loadings) or justify the per-skill approach more strongly.
- [§3.2.2, Cronbach's alpha] The combined alpha of .976 for 27 items is very high, which raises the possibility that a single general factor accounts for most of the variance. To support the two-category interpretation, the paper should report factor correlations from Model 3 and compare Model 3 against a unidimensional model and/or a two-factor model without the sub-factors. Without discriminant validity evidence, the high internal consistency is ambiguous and does not by itself validate the 'creating vs. evaluating' distinction.
minor comments (5)
- [Throughout] The phrase 'principle component analysis' should be 'principal component analysis' (abstract, §3.2, §3.2.2).
- [§3.2.2] 'Kaiser maximization' is presumably 'Kaiser normalization'; please correct.
- [§3.2.3] The typo 'combination of algorithmic thinking and abstraction' should read 'algorithmic thinking and decomposition' to match Table 4.
- [Footnote 1 and Table 4] The full 91-item pool is said to be 'available from the first author upon request'; for reproducibility, the candidate items and expert review outcomes should be included in a supplement or online repository.
- [§3.3] The gender comparison makes many Mann-Whitney tests without a multiple-comparison correction. Report effect sizes, and consider a conservative correction for the 27/54 item-level tests.
Circularity Check
No significant circularity: the factor structure is empirically tested on a hold-out sample, and the post-hoc model selection is a validity limitation, not a circular reduction.
full rationale
The paper's derivation chain is: (1) items generated from the literature and the theoretical framework in Table 1; (2) independent experts filter the items; (3) PCA on 200 respondents reduces the item set to 27; (4) CFA on the remaining 70 respondents tests a series of models. The abstract's claim that the items 'can be divided into two categories' is a summary of Model 3, which was specified 'based on an inspection of the items included in the final item set' (Section 3.2.3) and then fitted to the CFA sample. This is a specification search after Model 1 was inadmissible, but it is not circular: the two-category structure is a hypothesis imposed by the researchers, not an output that is used to redefine the inputs; the CFA loadings and fit statistics could have failed to support the model, and indeed Model 1 did fail. The use of a separate hold-out sample (the 49 Prolific responses plus 21 Costa Rican responses, distinct from the PCA's 200) provides some empirical separation between item reduction and factor-structure testing. The small CFA sample (n=70) and the post-hoc nature of Model 3 are legitimate methodological weaknesses, and the authors acknowledge the sample-size limitation in Section 3.4, but they do not constitute definitional or fit-forced circularity. No load-bearing argument reduces to a self-citation: the authors' own prior works are cited only for background context about CT interventions and programming skills, not as the basis for the factor structure. Therefore no circular step meeting the required evidentiary standard can be identified.
Assumptions & free parameters
free parameters (2)
- Model 3 factor structure =
Two broad factors: 'create solution' (abstraction, algorithmic thinking+decomposition, generalization) and 'evaluate sol
- Abstraction item exclusions =
Two abstraction items removed after loading on a second PCA component
assumptions (5)
- domain assumption Bandura's self-efficacy theory is the correct framework for measuring confidence in CT skills.
- domain assumption Selby/Woollard's five CT skills (abstraction, algorithmic thinking, decomposition, evaluation, generalization) with Dagiene's sub-skill refinement are a valid taxonomy of CT.
- domain assumption The CFA sample of 70 respondents is sufficient to test a 27-item multi-factor model.
- domain assumption Data from Prolific and the Costa Rica university setting can be pooled despite language/context differences.
- domain assumption Self-report on a 0-10 scale accurately operationalizes self-efficacy for CT tasks.
invented entities (1)
-
Two-category factor structure ('create solution' vs 'evaluate solution')
Cite this review
Pith. "Pith review of Measuring Computational Thinking Self-Efficacy (CT-SEI): Instrument development and preliminary evaluation." pith.science (2026). https://pith.science/paper/MLRNHYMW
@misc{pith2026260717704,
author = {Pith},
title = {Pith review of: Measuring Computational Thinking Self-Efficacy (CT-SEI): Instrument development and preliminary evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/MLRNHYMW}},
note = {Machine review of arXiv:2607.17704}
}
read the original abstract
Much effort is put into helping students at different educational levels develop Computational Thinking (CT) skills. Self-efficacy is important for skill development. It can predict perseverance, engagement and success on educational tasks. We created an instrument to measure self-efficacy of students in higher education for the CT skills abstraction, algorithmic thinking, decomposition, evaluation and generalization. First, 91 candidate items were created by including, adapting and extending items found in the literature. These items were evaluated by experts in the field of CT and education. 54 items remained and to reduce the number of items further, data was collected from 270 students in higher education recruited both through Prolific and a university setting in Costa Rica. Through principle component analysis (PCA) using a subset of 200 responses, the number of items was reduced to 27. Confirmatory factor analysis (CFA) using the remaining responses in the dataset showed the items can be divided into two categories: (1) creating the solution and (2) evaluating the solution. The created instrument can be valuable when assessing CT self-efficacy of students in higher education. With additional validation (e.g. examination of test-retest validity), we believe the scale could be used to evaluate the effectiveness of interventions, or decide what interventions should be provided to foster CT skill development.
Reference graph
Works this paper leans on
-
[1]
1997.Self-efficacy: the exercise of control
Albert Bandura. 1997.Self-efficacy: the exercise of control. W.H. Freeman and Company, New York, NY
1997
-
[2]
Albert Bandura. 2006. Guide for constructing self-efficacy scales. InSelf-efficacy beliefs of adolescents, F. Pajares and T.Editors Urdan (Eds.). Vol. 5. Information Age Publishing, 307–337. Measuring Computational Thinking Self-Efficacy (CT-SEI) UKICER 2026, September 03–04, 2026, Cambridge, United Kingdom
2006
-
[3]
2017.Computational Thinking : A Beginner’s Guide to Problem- solving and Programming
Karl Beecher. 2017.Computational Thinking : A Beginner’s Guide to Problem- solving and Programming. BCS, The Chartered Institute for IT, Swindon, UK
2017
-
[4]
Godfred O. Boateng, Torsten B. Neilands, Edward A. Frongillo, Hugo R. Melgar- Quiñonez, and Sera L. Young. 2018. Best Practices for Developing and Validating Scales for Health, Social, and Behavioral Research: A Primer.Frontiers in Public Health6 (2018). doi:10.3389/fpubh.2018.00149
arXiv 2018
-
[5]
Danielle Cadieux Boulden, Arif Rachmatullah, Kevin M Oliver, and Eric Wiebe
-
[6]
2006.Confirmatory factor analysis for applied research
Timothy A Brown. 2006.Confirmatory factor analysis for applied research. Guil- ford publications
2006
-
[7]
Computer Science Teachers Association (CSTA), and International Society for Technology in Education (ISTE). 2011. Computational Thinking Leadership Toolkit. http://www.iste.org/docs/ct-documents/ct-leadershipt-toolkit.pdf
2011
-
[8]
Valentina Dagien˙e, Sue Sentance, and Gabriel ˙e Stupurien˙e. 2017. Developing a two-dimensional categorization system for educational tasks in informatics. Informatica28, 1 (2017), 23–44. doi:10.15388/Informatica.2017.119
Show all 30 references
-
[9]
Imke de Jong and Johan Jeuring. 2020. Computational thinking interventions in higher education: A scoping literature review of interventions used to teach com- putational thinking. InProceedings of the 20th koli calling international conference on computing education research. 1–10
2020
-
[10]
2019.Computational thinking
Peter J Denning and Matti Tedre. 2019.Computational thinking. Mit Press
2019
-
[11]
AlMothana Gasaymeh and Reham AlMohtadi. 2024. College of education stu- dents’ perceptions of their computational thinking proficiency.Frontiers in Education9 (2024). doi:10.3389/feduc.2024.1478666
2024
-
[12]
Yasemin Gülbahar, Serhat Bahadır Kert, and Filiz Kalelioğlu. 2019. The self- efficacy perception scale for computational thinking skill: Validity and reliability study.Turkish Journal of Computer and Mathematics Education (TURCOMAT)10, 1 (2019), 01–29. doi:10.17762/turcomat.v10i1.194
2019 doi
-
[13]
Xinying Hou, Barbara Jane Ericson, and Xu Wang. 2024. Understanding the Ef- fects of Using Parsons Problems to Scaffold Code Writing for Students with Vary- ing CS Self-Efficacy Levels. InProceedings of the 23rd Koli Calling International Conference on Computing Education Rese...
2024
-
[14]
Ting-Chia Hsu, Shao-Chen Chang, and Yu-Ting Hung. 2018. How to learn and how to teach computational thinking: Suggestions based on a review of the literature.Computers & Education126 (2018), 296–310. doi:10.1016/j.compedu. 2018.07.004
2018 doi
-
[15]
Chiungjung Huang. 2013. Gender differences in academic self-efficacy: A meta- analysis.European journal of psychology of education28 (2013), 1–35
2013
-
[16]
Cynthia Hunt, Spencer Yoder, Taylor Comment, Thomas Price, Bita Akram, Lina Battestilli, Tiffany Barnes, and Susan Fisk. 2022. Gender, Self-Assessment, and Persistence in Computing: How gender differences in self-assessed ability reduce women’s persistence in computer science....
2022
-
[17]
Johan Jeuring, Roel Groot, and Hieke Keuning. 2023. What Skills Do You Need When Developing Software Using ChatGPT? (Discussion Paper). InProceedings of the 23rd Koli Calling International Conference on Computing Education Research. arXiv:2310.05998 [cs.SE]
2023 arXiv
-
[18]
Yaşar Özden
Özgen Korkmaz, Recep Çakir, and M. Yaşar Özden. 2017. A validity and reliability study of the computational thinking scales (CTS).Computers in Human Behavior 72 (2017), 558–569. doi:10.1016/j.chb.2017.01.005
2017 doi
-
[19]
Volkan Kukul and Serçin Karatas. 2019. Computational thinking self-efficacy scale: Development, validity and reliability.Informatics in Education18, 1 (2019), 151–164. doi:10.15388/infedu.2019.07
2019 doi
-
[20]
Robert C MacCallum, Keith F Widaman, Shaobo Zhang, and Sehee Hong. 1999. Sample size in factor analysis.Psychological methods4, 1 (1999), 84
1999
-
[21]
Masaki Matsunaga. 2010. How to Factor-Analyze Your Data Right: Do’s, Don’ts, and How-To’s.International journal of psychological research3, 1 (2010), 97–110
2010
-
[22]
Vidushi Ojha, Leah West, and Colleen M. Lewis. 2024. Computing Self-Efficacy in Undergraduate Students: A Multi-Institutional and Intersectional Analysis. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1(Portland, OR, USA)(SIGCSE 2024). Ass...
2024
-
[23]
Dale H. Schunk. 2012.Learning Theories: An Educational Perspective. Pearson, Boston, MA
2012
-
[24]
Cynthia Selby and John Woollard. 2013. Computational thinking: the developing definition. (2013). https://eprints.soton.ac.uk/356481/
2013
-
[25]
Cynthia C. Selby. 2015. Relationships: Computational Thinking, Pedagogy of Programming, and Bloom’s Taxonomy. InProceedings of the Workshop in Primary and Secondary Computing Education(London, United Kingdom)(WiPSCE ’15). Association for Computing Machinery, New York, NY, USA,...
2015
-
[26]
Shute, Chen Sun, and Jodi Asbell-Clarke
Valerie J. Shute, Chen Sun, and Jodi Asbell-Clarke. 2017. Demystifying computa- tional thinking.Educational Research Review22 (2017), 142–158. doi:10.1016/j. edurev.2017.09.003
2017 doi
-
[27]
Sónia Rolland Sobral. 2021. The Old Question: Which Programming Language Should We Choose to Teach to Program?. InAdvances in Digital Science, Tatiana Antipova (Ed.). Springer International Publishing, Cham, 351–364
2021
-
[28]
Xiaodan Tang, Yue Yin, Qiao Lin, Roxana Hadad, and Xiaoming Zhai. 2020. Assessing computational thinking: A systematic review of empirical studies. Computers & Education148 (2020). doi:10.1016/j.compedu.2019.103798
2020
-
[29]
Jeannette Wing. 2010. Computational thinking: What and Why. https://www.cs. cmu.edu/~CompThink/resources/TheLinkWing.pdf
2010
-
[2021]
Measuring in-service teacher self-efficacy for teaching computational think- ing: development and validation of the T-STEM CT.Education and Information technologies26, 4 (2021), 4663–4689
2021
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.