Pith. sign in

REVIEW 4 major objections 5 minor 30 references

Measuring Computational Thinking Self-Efficacy (CT-SEI): Instrument development and preliminary evaluation

T0 review · 4 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read Computational thinking self-efficacy can be measured with a 27-item, context-bound survey whose responses split into two factors: creating a solution and evaluating it.

desk verdict Useful instrument development with a genuinely good item pool, but the two-factor CFA claim rests on post-hoc model fitting with n=70 and is oversold in the abstract. read the letter →

arxiv 2607.17704 v1 pith:MLRNHYMW submitted 2026-07-20 cs.CY

classification cs.CY
keywords ComputationalThinkingself-efficacyinstrumentdevelopmentCT-SEIconfirmatoryfactoranalysisprincipalcomponenthighereducationgenderdifferences
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to measure a psychological belief: whether higher-education students think they can perform computational thinking (CT) skills such as abstraction, algorithmic thinking, decomposition, evaluation, and generalization. Because self-efficacy depends on context, the survey anchors every item in a concrete Python programming assignment. Starting from 91 candidate statements, expert review and principal component analysis reduced the set to 27, and confirmatory factor analysis on 70 additional responses suggests the items form two high-level factors: creating the solution and evaluating it. If this factor structure holds, educators and researchers gain a practical instrument for profiling student confidence and tailoring CT instruction.

What carries the argument

The instrument's three-part structure: a context-setting assignment, one overall task self-efficacy question, and 27 'I can' statements covering CT sub-skills on a 0-10 scale. The statistical machinery is principal component analysis with oblimin rotation for item reduction, followed by confirmatory factor analysis on a separate 70-response sample. The two-factor CFA model (create versus evaluate) is the load-bearing result that translates the theoretical skill list into a measurable two-dimensional belief structure.

What would settle it

Administer the final 27-item instrument to a new, larger sample of higher-education students in a preregistered confirmatory factor analysis. If the create/evaluate two-factor model shows RMSEA above 0.06 or CFI/TLI below 0.95, or if the two factors correlate near 1, the claimed structure fails to reproduce. A simpler check: low test-retest correlations on the same students would indicate the instrument does not measure a stable self-efficacy belief.

Watch

Extended reading notes

Core claim

The central claim is that CT self-efficacy in higher education can be measured as a context-dependent belief, and that when measured with 'I can' statements tied to a Python assignment, the 27 retained items do not reproduce the five theoretical CT skills. Instead, the best-fitting model groups the items into two categories: creating the solution (abstraction, algorithmic thinking, decomposition, and generalization items) and evaluating the solution (evaluation items), with reported fit statistics of RMSEA = 0.050, CFI = 0.972, TLI = 0.969. The paper also reports that women in the sample rated their CT self-efficacy lower than men on most items, a pattern consistent with broader findings in

Load-bearing premise

The load-bearing premise is that a two-factor structure fitted with only 70 respondents and with models chosen after inspecting the data describes a stable structure in the broader student population rather than a small-sample artifact.

Editorial extensions

If this is right

  • Educators can use the 27 items to profile a student's CT self-efficacy separately for creating and evaluating solutions, rather than relying on a single overall score.
  • The scale can serve as an outcome measure in intervention studies, including the authors' planned comparison of worked examples versus practice problems.
  • Because the instrument is context-dependent, the same items can be adapted to other CT tasks by changing the assignment text, though validity in new contexts still needs testing.
  • The reported gender differences suggest that CT instruction may need to address self-efficacy directly, especially for women in computing courses.
  • The CFA results imply that separate self-efficacy beliefs for algorithmic thinking, decomposition, abstraction, and generalization may not be distinguishable in this population.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the two-factor create/evaluate structure replicates, it would suggest that students' self-efficacy does not track the five-skill taxonomy that guided item construction; the taxonomy may describe competence better than perceived competence.
  • The gender differences here could reflect measurement non-invariance across groups or particularities of the sample; a formal invariance test would show whether the scale measures the same construct for men and women.
  • Because the CFA sample is small and models were selected after inspecting the data, a preregistered replication on a new sample, including test-retest reliability, is the natural next test of whether the two-factor structure is stable.
  • The instrument's link to actual behavior is untested; correlating CT-SEI scores with performance on the programming task or course grades would show whether self-reported confidence predicts the outcomes that matter.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents the development and preliminary validation of the Computational Thinking Self-Efficacy Instrument (CT-SEI), a 27-item scale measuring self-efficacy for abstraction, algorithmic thinking, decomposition, evaluation, and generalization in the context of a Python programming assignment. Ninety-one candidate items were reduced by expert review to 54, then by principal component analysis on 200 respondents to 27. A confirmatory factor analysis on the remaining 70 respondents is used to claim the items fall into two categories: creating the solution and evaluating the solution. The paper also reports gender differences favoring men on assignment-level self-efficacy and some items.

Significance. If the factor structure were supported, the CT-SEI would be a useful, theory-aligned instrument for computing education research. The paper has genuine strengths: item wording follows Bandura's self-efficacy principles, content validity is addressed via independent expert review, the assignment context is clearly specified, and the final item set is reported in Table 4. However, the central confirmatory claim is not supported by the reported analyses: the two-category structure was selected post hoc and estimated on a very small sample. The contribution at this stage is a carefully constructed item pool and a set of hypotheses about its structure, not a validated two-factor scale.

major comments (4)
  1. [§3.2.3 and Table 3] The headline two-category result is not a confirmatory finding. Model 1 was inadmissible, Model 2 was formed by collapsing algorithmic thinking and decomposition, and Model 3 was added 'based on an inspection of the items included in the final item set' after seeing the earlier results. All models are fit to the same n=70 sample. At this sample size, χ² has low power, and a model selected after inspecting the data can reproduce plausible-looking RMSEA/CFI/TLI values even if it is overfit. The abstract's 'showed the items can be divided into two categories' overstates what the analysis establishes. Please re-estimate the structure on independent data, or recast the analysis as exploratory and temper the conclusion accordingly. The existing limitation statement in §3.4 acknowledges the sample size but does not address the post-hoc selection problem.
  2. [§3.2.3 and Table 4] The model described as a 'two-factor' model is not a two-factor CFA at the item level. In Model 3, the 'Create solution' factor contains three sub-factors (Abstraction, Generalization, and Algorithmic Thinking & Decomposition), while 'Evaluate solution' contains only the Evaluation items. This is a higher-order structure, not a direct division of 27 items into two factors. The abstract and Section 3.2.3 use 'two categories' loosely. In addition, the text says the first factor consists of 'abstraction, generalization, and a combination of algorithmic thinking and abstraction', but Table 4 labels the combined factor 'Algorithmic Thinking & Decomposition'; this discrepancy should be fixed. Please specify the exact model (first-order factors, second-order loadings, constraints) that produced the fit indices in Table 3. A second-order factor with only one first-order factor (Evaluation) is no
  3. [§3.2.2] The item-reduction PCA was run separately for each of the five CT skills on its own item subset, rather than on the full 54-item set. This procedure cannot reveal cross-loadings or assess whether the five theoretically defined components are empirically distinct. It therefore partly entrenches the five-factor theory instead of testing it, and it makes the later CFA results harder to interpret. Given the strong interfactor correlations implied by the inadmissible Model 1 and the eventual collapsing of AT and DC, an exploratory factor analysis on all 54 items would be a more appropriate item-selection step. Please report the full PCA results (eigenvalues, loadings, cross-loadings) or justify the per-skill approach more strongly.
  4. [§3.2.2, Cronbach's alpha] The combined alpha of .976 for 27 items is very high, which raises the possibility that a single general factor accounts for most of the variance. To support the two-category interpretation, the paper should report factor correlations from Model 3 and compare Model 3 against a unidimensional model and/or a two-factor model without the sub-factors. Without discriminant validity evidence, the high internal consistency is ambiguous and does not by itself validate the 'creating vs. evaluating' distinction.
minor comments (5)
  1. [Throughout] The phrase 'principle component analysis' should be 'principal component analysis' (abstract, §3.2, §3.2.2).
  2. [§3.2.2] 'Kaiser maximization' is presumably 'Kaiser normalization'; please correct.
  3. [§3.2.3] The typo 'combination of algorithmic thinking and abstraction' should read 'algorithmic thinking and decomposition' to match Table 4.
  4. [Footnote 1 and Table 4] The full 91-item pool is said to be 'available from the first author upon request'; for reproducibility, the candidate items and expert review outcomes should be included in a supplement or online repository.
  5. [§3.3] The gender comparison makes many Mann-Whitney tests without a multiple-comparison correction. Report effect sizes, and consider a conservative correction for the 27/54 item-level tests.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the factor structure is empirically tested on a hold-out sample, and the post-hoc model selection is a validity limitation, not a circular reduction.

full rationale

The paper's derivation chain is: (1) items generated from the literature and the theoretical framework in Table 1; (2) independent experts filter the items; (3) PCA on 200 respondents reduces the item set to 27; (4) CFA on the remaining 70 respondents tests a series of models. The abstract's claim that the items 'can be divided into two categories' is a summary of Model 3, which was specified 'based on an inspection of the items included in the final item set' (Section 3.2.3) and then fitted to the CFA sample. This is a specification search after Model 1 was inadmissible, but it is not circular: the two-category structure is a hypothesis imposed by the researchers, not an output that is used to redefine the inputs; the CFA loadings and fit statistics could have failed to support the model, and indeed Model 1 did fail. The use of a separate hold-out sample (the 49 Prolific responses plus 21 Costa Rican responses, distinct from the PCA's 200) provides some empirical separation between item reduction and factor-structure testing. The small CFA sample (n=70) and the post-hoc nature of Model 3 are legitimate methodological weaknesses, and the authors acknowledge the sample-size limitation in Section 3.4, but they do not constitute definitional or fit-forced circularity. No load-bearing argument reduces to a self-citation: the authors' own prior works are cited only for background context about CT interventions and programming skills, not as the basis for the factor structure. Therefore no circular step meeting the required evidentiary standard can be identified.

Assumptions & free parameters 2 free parameters · 5 assumptions · 1 invented entities

This is an empirical scale-development paper rather than a parametric derivation, so there are no fitted physical constants. The burden rests instead on untested psychometric assumptions and on model/selection choices made after seeing the data.

free parameters (2)
  • Model 3 factor structure = Two broad factors: 'create solution' (abstraction, algorithmic thinking+decomposition, generalization) and 'evaluate sol
    Chosen by hand after Model 1 failed to produce an admissible covariance (covariance > 1) and after inspecting item content; the number and composition of factors were not fixed a priori.
  • Abstraction item exclusions = Two abstraction items removed after loading on a second PCA component
    Data-driven deletion based on factor loadings and the observation that the items used double negation; affects the final 27-item set.
assumptions (5)
  • domain assumption Bandura's self-efficacy theory is the correct framework for measuring confidence in CT skills.
    The entire instrument design, including 'I can' phrasing and context dependency, rests on Bandura's definition (Section 3).
  • domain assumption Selby/Woollard's five CT skills (abstraction, algorithmic thinking, decomposition, evaluation, generalization) with Dagiene's sub-skill refinement are a valid taxonomy of CT.
    The item pool and Table 1 are built around this taxonomy; if the taxonomy is wrong, the instrument lacks construct coverage.
  • domain assumption The CFA sample of 70 respondents is sufficient to test a 27-item multi-factor model.
    The paper cites MacCallum et al. for stability with smaller samples, but 70 respondents for 27 items is far below conventional guidance and the authors themselves flag it as a limitation (Section 3.4).
  • domain assumption Data from Prolific and the Costa Rica university setting can be pooled despite language/context differences.
    The questionnaire was translated to Spanish for 21 respondents, but no measurement-invariance analysis is reported (Section 3.2.1).
  • domain assumption Self-report on a 0-10 scale accurately operationalizes self-efficacy for CT tasks.
    The instrument assumes participants can reliably introspect and report their confidence; no validation against external performance is provided.
invented entities (1)
  • Two-category factor structure ('create solution' vs 'evaluate solution')
    purpose: Organizes the 27 CT self-efficacy items into a parsimonious higher-order structure.
    The structure was derived post hoc from the same dataset and has not been replicated or validated on an external sample.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Measuring Computational Thinking Self-Efficacy (CT-SEI): Instrument development and preliminary evaluation." pith.science (2026). https://pith.science/paper/MLRNHYMW

@misc{pith2026260717704,
  author       = {Pith},
  title        = {Pith review of: Measuring Computational Thinking Self-Efficacy (CT-SEI): Instrument development and preliminary evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MLRNHYMW}},
  note         = {Machine review of arXiv:2607.17704}
}
read the original abstract

Much effort is put into helping students at different educational levels develop Computational Thinking (CT) skills. Self-efficacy is important for skill development. It can predict perseverance, engagement and success on educational tasks. We created an instrument to measure self-efficacy of students in higher education for the CT skills abstraction, algorithmic thinking, decomposition, evaluation and generalization. First, 91 candidate items were created by including, adapting and extending items found in the literature. These items were evaluated by experts in the field of CT and education. 54 items remained and to reduce the number of items further, data was collected from 270 students in higher education recruited both through Prolific and a university setting in Costa Rica. Through principle component analysis (PCA) using a subset of 200 responses, the number of items was reduced to 27. Confirmatory factor analysis (CFA) using the remaining responses in the dataset showed the items can be divided into two categories: (1) creating the solution and (2) evaluating the solution. The created instrument can be valuable when assessing CT self-efficacy of students in higher education. With additional validation (e.g. examination of test-retest validity), we believe the scale could be used to evaluate the effectiveness of interventions, or decide what interventions should be provided to foster CT skill development.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

30 extracted references · 3 canonical work pages

  1. [1]

    1997.Self-efficacy: the exercise of control

    Albert Bandura. 1997.Self-efficacy: the exercise of control. W.H. Freeman and Company, New York, NY

  2. [2]

    Albert Bandura. 2006. Guide for constructing self-efficacy scales. InSelf-efficacy beliefs of adolescents, F. Pajares and T.Editors Urdan (Eds.). Vol. 5. Information Age Publishing, 307–337. Measuring Computational Thinking Self-Efficacy (CT-SEI) UKICER 2026, September 03–04, 2026, Cambridge, United Kingdom

  3. [3]

    2017.Computational Thinking : A Beginner’s Guide to Problem- solving and Programming

    Karl Beecher. 2017.Computational Thinking : A Beginner’s Guide to Problem- solving and Programming. BCS, The Chartered Institute for IT, Swindon, UK

  4. [4]

    Boateng, Torsten B

    Godfred O. Boateng, Torsten B. Neilands, Edward A. Frongillo, Hugo R. Melgar- Quiñonez, and Sera L. Young. 2018. Best Practices for Developing and Validating Scales for Health, Social, and Behavioral Research: A Primer.Frontiers in Public Health6 (2018). doi:10.3389/fpubh.2018.00149

  5. [5]

    Danielle Cadieux Boulden, Arif Rachmatullah, Kevin M Oliver, and Eric Wiebe

  6. [6]

    2006.Confirmatory factor analysis for applied research

    Timothy A Brown. 2006.Confirmatory factor analysis for applied research. Guil- ford publications

  7. [7]

    Computer Science Teachers Association (CSTA), and International Society for Technology in Education (ISTE). 2011. Computational Thinking Leadership Toolkit. http://www.iste.org/docs/ct-documents/ct-leadershipt-toolkit.pdf

  8. [8]

    Valentina Dagien˙e, Sue Sentance, and Gabriel ˙e Stupurien˙e. 2017. Developing a two-dimensional categorization system for educational tasks in informatics. Informatica28, 1 (2017), 23–44. doi:10.15388/Informatica.2017.119

Show all 30 references
  1. [9]

    Imke de Jong and Johan Jeuring. 2020. Computational thinking interventions in higher education: A scoping literature review of interventions used to teach com- putational thinking. InProceedings of the 20th koli calling international conference on computing education research. 1–10

  2. [10]

    2019.Computational thinking

    Peter J Denning and Matti Tedre. 2019.Computational thinking. Mit Press

  3. [11]

    AlMothana Gasaymeh and Reham AlMohtadi. 2024. College of education stu- dents’ perceptions of their computational thinking proficiency.Frontiers in Education9 (2024). doi:10.3389/feduc.2024.1478666

  4. [12]

    Yasemin Gülbahar, Serhat Bahadır Kert, and Filiz Kalelioğlu. 2019. The self- efficacy perception scale for computational thinking skill: Validity and reliability study.Turkish Journal of Computer and Mathematics Education (TURCOMAT)10, 1 (2019), 01–29. doi:10.17762/turcomat.v10i1.194

  5. [13]

    Xinying Hou, Barbara Jane Ericson, and Xu Wang. 2024. Understanding the Ef- fects of Using Parsons Problems to Scaffold Code Writing for Students with Vary- ing CS Self-Efficacy Levels. InProceedings of the 23rd Koli Calling International Conference on Computing Education Rese...

  6. [14]

    Ting-Chia Hsu, Shao-Chen Chang, and Yu-Ting Hung. 2018. How to learn and how to teach computational thinking: Suggestions based on a review of the literature.Computers & Education126 (2018), 296–310. doi:10.1016/j.compedu. 2018.07.004

  7. [15]

    Chiungjung Huang. 2013. Gender differences in academic self-efficacy: A meta- analysis.European journal of psychology of education28 (2013), 1–35

  8. [16]

    Cynthia Hunt, Spencer Yoder, Taylor Comment, Thomas Price, Bita Akram, Lina Battestilli, Tiffany Barnes, and Susan Fisk. 2022. Gender, Self-Assessment, and Persistence in Computing: How gender differences in self-assessed ability reduce women’s persistence in computer science....

  9. [17]

    Johan Jeuring, Roel Groot, and Hieke Keuning. 2023. What Skills Do You Need When Developing Software Using ChatGPT? (Discussion Paper). InProceedings of the 23rd Koli Calling International Conference on Computing Education Research. arXiv:2310.05998 [cs.SE]

  10. [18]

    Yaşar Özden

    Özgen Korkmaz, Recep Çakir, and M. Yaşar Özden. 2017. A validity and reliability study of the computational thinking scales (CTS).Computers in Human Behavior 72 (2017), 558–569. doi:10.1016/j.chb.2017.01.005

  11. [19]

    Volkan Kukul and Serçin Karatas. 2019. Computational thinking self-efficacy scale: Development, validity and reliability.Informatics in Education18, 1 (2019), 151–164. doi:10.15388/infedu.2019.07

  12. [20]

    Robert C MacCallum, Keith F Widaman, Shaobo Zhang, and Sehee Hong. 1999. Sample size in factor analysis.Psychological methods4, 1 (1999), 84

  13. [21]

    Masaki Matsunaga. 2010. How to Factor-Analyze Your Data Right: Do’s, Don’ts, and How-To’s.International journal of psychological research3, 1 (2010), 97–110

  14. [22]

    Vidushi Ojha, Leah West, and Colleen M. Lewis. 2024. Computing Self-Efficacy in Undergraduate Students: A Multi-Institutional and Intersectional Analysis. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1(Portland, OR, USA)(SIGCSE 2024). Ass...

  15. [23]

    Dale H. Schunk. 2012.Learning Theories: An Educational Perspective. Pearson, Boston, MA

  16. [24]

    Cynthia Selby and John Woollard. 2013. Computational thinking: the developing definition. (2013). https://eprints.soton.ac.uk/356481/

  17. [25]

    Cynthia C. Selby. 2015. Relationships: Computational Thinking, Pedagogy of Programming, and Bloom’s Taxonomy. InProceedings of the Workshop in Primary and Secondary Computing Education(London, United Kingdom)(WiPSCE ’15). Association for Computing Machinery, New York, NY, USA,...

  18. [26]

    Shute, Chen Sun, and Jodi Asbell-Clarke

    Valerie J. Shute, Chen Sun, and Jodi Asbell-Clarke. 2017. Demystifying computa- tional thinking.Educational Research Review22 (2017), 142–158. doi:10.1016/j. edurev.2017.09.003

  19. [27]

    Sónia Rolland Sobral. 2021. The Old Question: Which Programming Language Should We Choose to Teach to Program?. InAdvances in Digital Science, Tatiana Antipova (Ed.). Springer International Publishing, Cham, 351–364

  20. [28]

    Xiaodan Tang, Yue Yin, Qiao Lin, Roxana Hadad, and Xiaoming Zhai. 2020. Assessing computational thinking: A systematic review of empirical studies. Computers & Education148 (2020). doi:10.1016/j.compedu.2019.103798

  21. [29]

    Jeannette Wing. 2010. Computational thinking: What and Why. https://www.cs. cmu.edu/~CompThink/resources/TheLinkWing.pdf

  22. [2021]

    Measuring in-service teacher self-efficacy for teaching computational think- ing: development and validation of the T-STEM CT.Education and Information technologies26, 4 (2021), 4663–4689

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.