{"id":"ebb2d34c-c251-455b-9dd5-460b85b84b1e","arxiv_id":"2606.27398","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The study implements and evaluates a Bloom-aligned GenAI framework in SE/CS courses, finding students value it most for higher-order tasks and that explicit guidance promotes reflective use.","lead":"This paper reports on embedding a Bloom's taxonomy-aligned framework for generative AI use into software engineering and computer science courses at two universities, then surveying student and instructor perceptions. A smart generalist might read it to see one concrete way educators are trying to make AI tools support deeper learning rather than replace it.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Evidence is limited to thematic analysis of self-reported perceptions without controls or objective learning metrics.","rationale":"The reader's weakest assumption directly names the same vulnerability in the qualitative evidence base. The abstract already signals this limitation; the method description reinforces that it is load-bearing for any causal claim about the framework's value.","tokens_in":1782,"tokens_out":234,"duration_ms":21540,"concrete_test":"Add a matched control cohort in one course (identical content, no Bloom guidance) and have independent raters score final artifacts for cognitive level and independent problem-solving; if no reliable difference appears versus the intervention cohort, attribution to the framework is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that Bloom-aligned guidance produces genuine shifts toward reflective GenAI use and educational value. The study embeds the framework then interprets anonymous questionnaires and artifacts via thematic analysis at two universities. No control sections, pre-intervention baselines, or blinded outcome scoring are described, so reported influences on behavior could reflect social-desirability bias, instructor expectations, or course-specific factors rather than the taxonomy alignment itself.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper describes the design and multi-course deployment of a Bloom's taxonomy-aligned framework for guiding GenAI use in SE/CS education at two universities. It reports thematic analysis of anonymous student/instructor questionnaires and learning artifacts, finding that students perceive GenAI as more valuable for higher-order tasks (analysis, evaluation, reflection), that explicit guidance promotes reflective use and selective non-use, and that both groups note pedagogical benefits alongside increased cognitive and design effort. The central claim is that Bloom's taxonomy supplies a scalable, pedagogy-first alternative to enforcement-focused approaches for responsible GenAI integration.","tokens_in":1842,"tokens_out":531,"duration_ms":21507,"significance":"If the perception data can be shown to reflect genuine behavioral shifts rather than reporting bias, the work supplies a concrete, replicable instructional framework that links cognitive-level goals to GenAI roles. This is a practical contribution to the growing literature on AI in computing education and could inform curriculum design at other institutions.","major_comments":[{"comment":"Methods: The description of data collection provides no sample sizes, response rates, number of courses or sections involved, number of learning artifacts examined, or inter-rater reliability statistics for the thematic analysis. These omissions make it impossible to evaluate the robustness or generalizability of the reported themes.","section":"Methods"},{"comment":"Results/Discussion: Claims that the Bloom-aligned guidance produced changes in student behavior (e.g., \"delayed or intentional non-use\") rest entirely on post-intervention self-reports without pre-intervention baselines, control sections, or objective outcome measures. This leaves open the possibility that observed perceptions reflect social-desirability bias, instructor expectations, or course-specific factors rather than the framework itself.","section":"Results"},{"comment":"Analysis: The thematic analysis uses Bloom's taxonomy as an analytic lens, yet the manuscript does not describe how codes were mapped to taxonomy levels, whether coding was performed blind to the framework, or how disagreements between coders were resolved.","section":"Analysis"}],"minor_comments":[{"comment":"The abstract and introduction would benefit from a brief statement of the number of courses and approximate participant numbers to give readers an immediate sense of scale.","section":"Abstract"},{"comment":"Figure or table summarizing the Bloom-level GenAI roles would improve clarity and allow readers to assess the framework without reconstructing it from prose.","section":"Framework"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed feedback. We address each major comment below and indicate the revisions planned for the next version of the manuscript.","responses":[{"response":"We agree that these details are necessary to assess robustness. The revised manuscript will add the number of courses and sections at each institution, total enrolled students, questionnaire response rates, number of learning artifacts examined, and inter-rater reliability statistics (including Cohen's kappa) for the thematic analysis.","revision_made":"yes","referee_comment":"[Methods] The description of data collection provides no sample sizes, response rates, number of courses or sections involved, number of learning artifacts examined, or inter-rater reliability statistics for the thematic analysis. These omissions make it impossible to evaluate the robustness or generalizability of the reported themes."},{"response":"We acknowledge the limitation. The study was exploratory and focused on post-implementation perceptions; no pre-intervention baselines or control sections were collected. The revision will add an explicit Limitations section discussing reliance on self-reports, potential social-desirability bias, and the absence of objective behavioral measures. Language will be adjusted to frame findings as student-reported perceptions supported by convergent evidence from artifacts, rather than demonstrated behavioral change.","revision_made":"partial","referee_comment":"[Results] Claims that the Bloom-aligned guidance produced changes in student behavior (e.g., \"delayed or intentional non-use\") rest entirely on post-intervention self-reports without pre-intervention baselines, control sections, or objective outcome measures. This leaves open the possibility that observed perceptions reflect social-desirability bias, instructor expectations, or course-specific factors rather than the framework itself."},{"response":"We will expand the Methods section to describe the analysis process in detail. This will cover how codes were iteratively mapped to Bloom's levels, that coding was conducted with awareness of the framework (full blinding was not feasible due to the framework's embedding in course materials) while prioritizing data-driven coding, and that coder disagreements were resolved through discussion until consensus was reached.","revision_made":"yes","referee_comment":"[Analysis] The thematic analysis uses Bloom's taxonomy as an analytic lens, yet the manuscript does not describe how codes were mapped to taxonomy levels, whether coding was performed blind to the framework, or how disagreements between coders were resolved."}],"tokens_in":1436,"tokens_out":508,"duration_ms":21373,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core thing here is a concrete framework that spells out GenAI roles at each Bloom level and then gets embedded into course materials, labs, and assessments across several SE/CS classes at Queen's University Belfast and Azerbaijan Technical University. Students reportedly saw more value at higher cognitive levels and shifted toward more deliberate use or non-use when the guidance pointed that way. Instructors noted both benefits and added workload.\n\nWhat the work actually does is take an existing taxonomy and apply it systematically to a current tool in one domain. The multi-course, two-institution rollout plus the mix of questionnaires and artifacts gives it a bit more reach than a single-class case study. That part is useful for anyone who needs a ready template rather than starting from scratch.\n\nThe soft spots sit in the evidence. Thematic analysis of anonymous responses and artifacts is described, yet there are no sample sizes, response rates, inter-rater reliability figures, or pre-intervention baselines. No control sections or blinded scoring appear, so the claimed shifts in reflective use could easily trace to social-desirability effects, instructor expectations, or course-specific factors instead of the taxonomy alignment itself. The central claim about genuine behavioral change therefore rests on perception data alone.\n\nThis is for computing educators who are already dealing with GenAI in their classes and want structured guidance rather than blanket rules. A reader running similar courses could borrow the framework and adapt it quickly. It is not aimed at learning theorists or researchers needing rigorous causal evidence.\n\nI would send it to peer review. The idea is timely and the execution is straightforward; referees can ask for the missing methodological details and objective checks without starting over.","headline":"The paper supplies a practical Bloom-aligned framework for GenAI in SE/CS courses at two universities, but the supporting data are thin self-reported perceptions without controls or objective measures.","tokens_in":2405,"tokens_out":410,"would_cite":false,"duration_ms":16210,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Bloom's taxonomy aligns GenAI use with cognitive levels in software engineering and computer science education.","keywords":["GenAI","Bloom's taxonomy","software engineering education","computer science education","responsible AI use","cognitive levels","student perceptions","instructional guidance"],"falsifier":"A controlled comparison showing no increase in reflective or delayed GenAI use when Bloom-level guidance is added versus standard instructions.","tokens_in":2669,"feed_emoji":"🎓","tokens_out":544,"duration_ms":27157,"temperature":0.7,"pith_summary":"The paper tests whether mapping generative AI roles to Bloom's taxonomy levels helps students and instructors use the tools responsibly in SE and CS courses. Explicit guidance led students to favor AI for analysis, evaluation, and reflection while sometimes skipping it for foundational work to preserve independent thinking. Data came from questionnaires and artifacts across courses at two universities, with thematic analysis showing both benefits and added design effort. A sympathetic reader would care because the approach offers a concrete way to integrate AI without defaulting to bans or unchecked adoption.","feed_headline":"Bloom's taxonomy guides responsible GenAI use in SE and CS courses","feed_subtitle":"Students apply AI more reflectively to analysis and evaluation when instructions match cognitive goals.","key_machinery":"Bloom-aligned GenAI framework that articulates appropriate GenAI roles at different cognitive levels.","core_discovery":"GenAI's educational value lies in intentional alignment between cognitive learning goals, instructional guidance, and learner self-regulation. Bloom's taxonomy provides a scalable, pedagogy-driven framework for responsible GenAI use in SE/CS education, offering a practical alternative to enforcement-focused responses.","pith_inferences":["Other fields could map similar cognitive taxonomies to define AI assistance boundaries.","Longer-term tracking could check whether reflective habits continue after the guided courses end.","Replicating the framework in additional universities would test how much the two-site results depend on local context."],"forward_implications":["Students perceive GenAI as most valuable for higher-order cognitive activities such as analysis, evaluation, and reflection.","Explicit Bloom-level guidance influences students to use GenAI reflectively, with delayed or intentional non-use when independent thinking is prioritized.","Both students and instructors report pedagogical benefits alongside challenges in cognitive effort and instructional design workload."],"fun_headline_variants":["Bloom taxonomy links GenAI use to cognitive levels in SE courses","Students see GenAI value in analysis evaluation with Bloom guidance","Bloom alignment promotes reflective GenAI use in CS education","GenAI value stems from alignment with Bloom cognitive goals"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Thematic analysis of anonymous questionnaires and learning artifacts accurately reflects genuine changes in student behavior and perceptions rather than social-desirability bias or course-specific effects.","fun_headline_variants_meta":{"raw":{"variants":["Bloom taxonomy links GenAI use to cognitive levels in SE courses","Students see GenAI value in analysis evaluation with Bloom guidance","Bloom alignment promotes reflective GenAI use in CS education","GenAI value stems from alignment with Bloom cognitive goals"]},"model":"grok-4.3","cost_usd":0.008295,"raw_usage":{"total_tokens":3753,"prompt_tokens":655,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":82949500,"prompt_tokens_details":{"text_tokens":655,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3033,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":655,"tokens_out":65,"duration_ms":29608,"temperature":1.0,"reasoning_tokens":3033,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T01:50:43.795380+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled comparison showing no increase in reflective or delayed GenAI use when Bloom-level guidance is added versus standard instructions.","supporting_citations":[],"review_version":1}