{"id":"32fc3d20-eb7f-4b91-8497-2dc4f8430fad","arxiv_id":"2506.06193","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The Critical Reflection and Agency in Computing Index shows evidence of validity, and computing ethics training is associated with higher scores on some ethical dimensions but also with stronger techno-solutionism.","lead":"The paper validates a survey index for measuring critical reflection and agency in computing, using two Prolific samples of computing students and professionals. If the index holds up, ethics educators gain a standard, theory-based tool for comparing teaching approaches and tracking student development.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Techno-Solutionism factor is built entirely from reverse-coded items, so the flagship ethics-course finding may reflect response style rather than a measured belief.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the all-reverse-coded Techno-Solutionism factor may reflect a method artifact rather than a substantive construct. This concern is concrete, located in the reported EFA and CFA results, and directly threatens the novelty claim that ethics courses are associated with stronger techno-solutionist beliefs. The paper is otherwise transparent and follows standard psychometric procedures, so the appropriate verdict remains CONDITIONAL: the index and its validation are promising, but the key techno-solutionism finding should be treated as exploratory until the item-valence confound is addressed. No change to the reader's verdict is needed; the conditional status already captures this uncertainty.","tokens_in":18268,"tokens_out":4356,"duration_ms":40048,"concrete_test":"Collect new data (or re-administer to a fresh sample) with a balanced set of six techno-solutionism items: three positively keyed (e.g., \"Technology can solve most social problems\") and three negatively keyed (e.g., \"Some problems cannot be fixed by technology\"). Run the same EFA/CFA and the ethics-training regression, additionally controlling for a response-style covariate such as each participant's intra-individual standard deviation across items. If a single techno-solutionism factor still emerges and the ethics-training coefficient remains significant after the response-style control, the concern is resolved; if the factor splits by item valence or the coefficient vanishes, the current Techno-Solutionism scale and the ethics-course association are artifacts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is the Techno-Solutionism factor. In Section 4.1.2, the authors report that \"several reverse-coded items\" merged into a single factor, which they then labeled Techno-Solutionism. Appendix A and the confirmatory factor analysis (Table 11) confirm that all four retained items are reverse-coded relative to the original \"recognizing computing/data embeds values/power\" construct. A factor defined entirely by items sharing the same response direction can emerge from method effects such as acquiescence, careless responding, or differential interpretation of negatively keyed statements, rather than from a substantive latent belief. If that is what happened, then the construct-validity correlations involving techno-solutionism (e.g., r = .44 with System Responsiveness in Table 1) and the paper's headline result that ethics training increases techno-solutionism (Table 6, β = .17, p < .001) do not measure what they claim. The paper's own account—that hypothesized \"data has limits\" items failed to load reliably and that the remaining reverse-coded items clustered together—makes the method-effect explanation plausible and, critically, untested. This is the single assumption whose failure would most change the paper's conclusions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a two-sample validation study (N=474 and N=464) of the Critical Reflection and Agency in Computing Index, a scale intended to measure critical reflection and critical agency in computing students and professionals. Using exploratory factor analysis, confirmatory factor analysis, reliability analysis, construct-validity correlations, and regressions against workplace ethical behavior and prior ethics training, the authors argue that the index has a stable factor structure, adequate reliability for most subscales, construct validity, and predictive criterion validity. They also report the headline finding that completing computing ethics courses is associated with higher scores on some reflection and agency dimensions but also with higher techno-solutionism scores. The paper concludes that the index is a validated tool for assessment and that ethics pedagogy may have both beneficial and unintended effects.","tokens_in":18485,"tokens_out":6373,"duration_ms":64356,"significance":"If the validity evidence holds, this paper would give computing education research a standardized, theory-grounded instrument for measuring ethical development, filling a gap identified in prior systematic reviews. The use of two independent samples, attention checks, and external behavioral outcome questions are genuine strengths, as is the transparent reporting of the factor analytic process and limitations. However, the validity claim is currently undermined by several load-bearing issues: the techno-solutionism factor is composed entirely of reverse-coded items and may reflect response style; the CFA fit is reported only by threshold, not by numeric values; and the Bonferroni-adjusted criterion-validity results are not consistently reported. These issues do not necessarily invalidate the work, but they require additional analysis and reporting before the instrument's validity can be considered established.","major_comments":[{"comment":"The Techno-Solutionism factor is defined entirely by reverse-coded items. In the EFA in Table 7, every item loading above .4 on Factor 1 is reverse-coded, and the four items retained in the final CFA in Table 11 are all marked (R) in Table 3. A factor composed exclusively of negatively keyed items can emerge from method effects such as acquiescence, careless responding, or differential interpretation of reverse wording, rather than from a substantive latent belief. This is not merely a psychometric nuance: the construct-validity interpretation of the moderate correlations involving techno-solutionism (e.g., r = .44 with System Responsiveness in Table 1) and the headline RQ3 result that ethics training is associated with higher techno-solutionism (Table 6, beta = .17, p < .001) both depend on this factor measuring what it is claimed to measure. The authors themselves note that the hypothesized 'data has limits' items failed to load reliably and that reverse-coded items merged, which makes the method-effect explanation plausible. To support the interpretation, the paper should present evidence that the factor is not an artifact of item keying, for example by comparing the current model with one that includes a method factor for reverse-coded items, or by validating the factor against a separately collected set of non-reverse-coded techno-solutionism items. Without such evidence, the techno-solutionism findings should be reported with the caveat that they may partly reflect response style.","section":"Section 4.1.2, Tables 7 and 11; Section 4.2.3; Table 6"},{"comment":"The confirmatory factor analysis is claimed to support the final model, but the paper gives no numeric fit statistics. The text states that RMSEA, TLI, CFI, and SRMR scored 'adequately' relative to thresholds, yet no actual values are reported anywhere in the paper or appendices. Without the numeric RMSEA, TLI, CFI, and SRMR values (and ideally chi-square and degrees of freedom), the dimensionality claim that underlies the entire validation cannot be independently checked. This is a central piece of evidence for RQ1 and RQ2, so the fit indices must be reported explicitly.","section":"Section 4.2.1"},{"comment":"The criterion-validity results are reported in a way that is inconsistent with the stated Bonferroni correction. The text says that for the 'noticing ethical issues' regressions the significance level was adjusted to alpha = .017, but Table 2 reports stars using conventional thresholds (* p < .05, ** p < .01) and the text states that 'valuing ethics training' and 'valuing marginalized perspectives' showed a significant association without indicating whether these associations survive the .017 threshold. A coefficient with one star does not necessarily meet a .017 threshold. The authors should report exact p-values or clearly mark which coefficients meet the adjusted threshold. Because the criterion-validity claim is a key contribution, this ambiguity is load-bearing and should be resolved.","section":"Section 4.2.3, Table 2"},{"comment":"The final index includes items from the 'computing has limits' factor after that factor was dropped for very low reliability (alpha = .38). The paper retains two such items as standalone 'reference points,' and they appear in Table 3 under the construct label 'Recognizing computing/data embeds values/power.' However, these two items are not included in the CFA model shown in Tables 11-13 and therefore have no reliability or validity evidence in this study. This creates ambiguity about whether the final 29-item index is fully validated as claimed. The paper should clearly separate validated subscales from the two standalone reference items, for example by placing them in a distinct table or clearly labeling them as not part of any validated subscale, and should avoid implying that the CFA supports the full 29-item index.","section":"Sections 4.2.1 and 4.3, Tables 3 and 11"}],"minor_comments":[{"comment":"Reliability coefficients are inconsistent between the text and the CFA tables: the text reports Valuing Ethics Training alpha = .89 and System Responsiveness alpha = .84, while Table 12 reports .87 for the ethics-training factor and Table 13 reports .77 for System Responsiveness; Personal Effectiveness is .89 in the text and .87 in Table 13. These discrepancies should be reconciled.","section":"Section 4.2.2 and Appendix C"},{"comment":"The second sample is described as N=464 in Section 3.1.1 and Section 4.1.1, but Table 10, described as the CFA sample, reports N=474, and Table 6 reports N=472. The sample size should be consistent across the text and tables.","section":"Section 3.1.1 and Table 10"},{"comment":"The regression tables report standardized coefficients with standard errors in parentheses, but standard errors of standardized coefficients are not usually reported this way and the reader cannot tell whether the values in parentheses are standard errors of the standardized or unstandardized coefficient. Clarify the reporting convention in the table notes or text.","section":"Table 2 and Table 6"},{"comment":"The Likert scale is described as '(Strongly Disagree (1/6) to Strongly Agree (6/6))'; this should read '1 to 6' rather than '1/6' and '6/6'.","section":"Section 3.1.1"},{"comment":"The item 'Ethics discussions in computing should only involve computer scientists. (R)' loads above .4 on the techno-solutionism factor in the EFA but does not appear in the final index or in the CFA table. The paper should explain why this item was removed despite meeting the loading threshold.","section":"Appendix A, Table 7"},{"comment":"The first row of Table 3 labels the two retained 'computing has limits' items as 'Recognizing computing/data embeds values/power,' but this is not a validated factor in the final CFA. Add a footnote to indicate that these items are standalone reference items and are not part of any validated subscale.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of ICER and addresses an important measurement gap. The main concern is the reverse-coded techno-solutionism factor, which is central to the headline finding and may reflect method effects; this needs to be addressed with additional analysis or substantially weakened claims. The missing numeric CFA fit statistics and the Bonferroni reporting inconsistency are also fixable but must be corrected before the paper can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent, transparent psychometric validation of an index the authors built earlier, and the validation itself is worth having. The most interesting result—ethics training associated with stronger techno-solutionism—is probably not a real finding, because every Techno-Solutionism item is reverse-keyed. The stress-test note is correct, and the paper's own EFA description makes the method-effect explanation the natural one.\n\nWhat's good: two independent samples (N=474, N=464) with attention checks; standard EFA/CFA; model fit adequate; construct validity correlations mostly follow theory; criterion validity work with workplace behavior is a nice step that most scale papers skip; and the authors are unusually honest about the 'computing has limits' factor failing (alpha .38) and about keeping two of its items anyway. They also cite and follow Boateng et al.'s best practices. For the field, a reliable index for critical reflection and agency in computing is a real gap, so this is a useful contribution.\n\nWhere it frays: the Techno-Solutionism factor is four reverse-coded items, all sharing the same response direction. The EFA in Appendix A shows those items cluster; the CFA confirms. A factor composed entirely of reverse-keyed items can easily be an acquiescence/careless responding artifact. The paper never tests for this, and it then builds two of its most provocative claims on it: the r = .44 with System Responsiveness and the beta = .17 for ethics training. I would not put weight on either until keying-balanced items are tested. The 'valuing marginalized perspectives' alpha is .63, below the usual .7, which the authors admit. The ethics training comparison is self-selected, so the direction of causality is not established (the paper says 'associated' most of the time, but the abstract phrases it as 'showed higher scores').\n\nThe bottom line: the validation evidence for the core agency subscales is decent, and the index is likely usable after revision. The techno-solutionism result should be treated as exploratory at best, and the paper should be asked to do a method-effects check or revise the claims.","headline":"Honest validation study with a likely artifact in its most interesting result: the Techno-Solutionism factor is all reverse-coded items, so the ethics-course association may be method noise.","tokens_in":18987,"tokens_out":1938,"would_cite":true,"duration_ms":19700,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper provides psychometric evidence that the Critical Reflection and Agency in Computing Index is valid, and finds that computing ethics courses are associated with higher scores on some reflection and agency dimensions but also…","keywords":["computing ethics education","critical consciousness","psychometric validation","factor analysis","techno-solutionism","critical reflection","critical agency","computing education research"],"falsifier":"Administer the four techno-solutionism items in both their original reverse-coded form and a positively worded version to the same sample, and compare factor structure and correlations with ethics training; if the positive-wording version no longer forms a factor or no longer correlates with course completion, the construct is a wording artifact.","tokens_in":18074,"feed_emoji":"📋","tokens_out":8605,"duration_ms":79221,"temperature":0.7,"pith_summary":"The paper argues that the Critical Reflection and Agency in Computing Index is a psychometrically valid instrument for measuring how computing students and professionals reflect on ethical issues and feel able to act on them. The case rests on two survey samples and shows that the index has a distinct factor structure, acceptable reliability for most subscales, construct validity, and predictive criterion validity. The paper then uses the validated index to ask whether computing ethics courses matter, finding that course completion is associated with higher scores on some dimensions of ethical reflection and agency but also with stronger techno-solutionist beliefs. If correct, the field gains a common measurement tool and evidence that current ethics pedagogy has both intended and unintended associations.","feed_headline":"Computing ethics courses raise agency and solutionism, index finds","feed_subtitle":"A two-sample validation links ethics training to more critical agency but also to stronger techno-solutionist beliefs.","key_machinery":"The central object is the Critical Reflection and Agency in Computing Index, a Likert-scale survey instrument grounded in Critically Conscious Computing theory, which distinguishes critical reflection from critical agency. The argument is carried by a two-phase psychometric procedure: exploratory factor analysis with promax rotation in the first sample to discover how items cluster, then confirmatory factor analysis in the second sample to test whether that cluster structure fits, followed by reliability estimates and regression-based checks of construct and criterion validity. The same regression setup is reused to compare respondents who had completed a computing ethics or professionalism course with those who had not, making the instrument the bridge between the validation result and the education-effect finding.","core_discovery":"On its own terms, the paper establishes that the index measures six related but distinct constructs: valuing marginalized perspectives, techno-solutionism, valuing ethics training, valuing technical training, personal effectiveness, and system responsiveness. A first sample (N=474) supplied exploratory factor analysis; a second (N=464) supplied confirmatory factor analysis, with model-fit indices meeting standard thresholds and all retained loadings above 0.55. Cronbach's alpha was .63-.89, with most subscales above .7; the \"computing has limits\" subscale failed and was dropped from the confirmatory analysis while two of its items were kept as standalone reference points. Construct validity was supported by predicted correlations (for example, personal effectiveness and system responsiveness at r=.62), and predictive criterion validity by regressions showing that valuing ethics training and valuing marginalized perspectives relate to noticing ethical issues at work while personal effectiveness and system responsiveness relate to acting on them. Controlling for demographics, completing a computing ethics course was associated with higher personal effectiveness (beta=.30), higher system responsiveness (beta=.27), higher valuing of ethics training (beta=.11), and higher techno-solutionism (beta=.17), with no significant association for valuing marginalized perspectives.","pith_inferences":["All four retained techno-solutionism items are reverse-coded; a plausible alternative explanation is that the factor reflects agreement with negatively worded statements, so the ethics-course association with techno-solutionism may be partly a response-style artifact.","For the same reason, the positive correlations between techno-solutionism and the agency measures could reflect shared method variance rather than a genuine link between faith in technical fixes and confidence in taking ethical action.","A direct test would be to rewrite the techno-solutionism items in positive wording and re-run the factor analysis; if the factor dissolves, the construct as currently measured is not a substantive belief.","A longitudinal pre/post design with balanced item wording would convert the course-completion associations into a stronger case for causality and would reveal whether the unintended solutionism bump is temporary or persistent."],"forward_implications":["The index gives computing ethics researchers a shared, theory-grounded measure, replacing intervention-specific surveys that lack validity evidence.","Educators can use the subscales to identify which dimensions, such as valuing marginalized perspectives, are not shifting and target instruction accordingly.","The positive association between ethics training and techno-solutionism implies that standalone ethics courses may unintentionally reinforce the belief that ethical problems are technical puzzles, suggesting course designs should explicitly address technology's limits.","Because personal effectiveness and system responsiveness predict real-world ethical action, changes in these subscale scores are meaningful signals for whether pedagogy will affect workplace behavior.","The unusable \"computing has limits\" subscale shows the instrument does not yet capture students' understanding of technology's limitations, and that dimension needs new item development."],"supporting_citations":[{"why":"Supplies the index items and the Critically Conscious Computing theory the validation is designed to test.","marker":"[47]"},{"why":"Provides the scale-development and validation procedure the study follows, from item extraction to reliability and validity tests.","marker":"[5]"},{"why":"Documents the lack of validated psychometric surveys in computing ethics education, the gap the index fills.","marker":"[7]"},{"why":"Provides the critical consciousness scale whose theoretical structure and expected correlations anchor the construct validity argument.","marker":"[17]"},{"why":"Grounds the criterion measures of noticing and acting on ethical issues in real workplace practices.","marker":"[60]"},{"why":"Names and defines techno-solutionism, the belief used to interpret the emerged reverse-coded factor.","marker":"[41]"},{"why":"Frames the pedagogical concern about balancing technology's power and limitations that motivates the techno-solutionism finding.","marker":"[35]"}],"fun_headline_variants":["Ethics courses boost agency but also solutionist beliefs, study finds","Validation links ethics courses to agency and techno-solutionism","Ethics education tied to higher agency, but also more solutionism","Courses raise critical agency and techno-solutionism, index shows","New index reveals ethics courses aid agency yet reinforce solutionism"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that ethics courses strengthen techno-solutionism assumes that the reverse-worded survey items measure a real belief; if they only capture how respondents answer negatively phrased statements, that finding is an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Ethics courses boost agency but also solutionist beliefs, study finds","Validation links ethics courses to agency and techno-solutionism","Ethics education tied to higher agency, but also more solutionism","Courses raise critical agency and techno-solutionism, index shows","New index reveals ethics courses aid agency yet reinforce solutionism"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000455,"raw_usage":{"total_tokens":2270,"prompt_tokens":914,"completion_tokens":1356,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":1268}},"tokens_in":530,"tokens_out":1356,"duration_ms":9590,"temperature":1.0,"reasoning_tokens":1268,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:58:22.690827+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Administer the four techno-solutionism items in both their original reverse-coded form and a positively worded version to the same sample, and compare factor structure and correlations with ethics training; if the positive-wording version no longer forms a factor or no longer correlates with course completion, the construct is a wording artifact.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the index items and the Critically Conscious Computing theory the validation is designed to test."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the scale-development and validation procedure the study follows, from item extraction to reliability and validity tests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the lack of validated psychometric surveys in computing ethics education, the gap the index fills."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the critical consciousness scale whose theoretical structure and expected correlations anchor the construct validity argument."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the criterion measures of noticing and acting on ethical issues in real workplace practices."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Names and defines techno-solutionism, the belief used to interpret the emerged reverse-coded factor."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames the pedagogical concern about balancing technology's power and limitations that motivates the techno-solutionism finding."}],"review_version":1}