{"id":"6c090f22-2103-4f78-9e94-fb33275848fe","arxiv_id":"2507.17230","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Two semester-long student case studies show trade-offs between GenAI proficiency, ethics, and career confidence, suggesting that GenAI-integrated teaching should be evaluated and designed as a multi-goal 'wicked problem.'","lead":"A two-student longitudinal study of a generative AI-integrated creative media course found that students who gained GenAI skills sometimes lost ethical ground, while students who became more ethically aware used GenAI less and grew fewer skills. The paper argues this makes teaching with generative AI a 'wicked problem' that needs balancing, not optimizing, of competing goals.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed trade-offs are not established by the reported data: Pat's ethics were already absent at Interview 1, and Jay's 'impeded skill development' is inferred from reduced use rather than measured skill.","rationale":"I read the paper as an intentionally small qualitative illustration; the limitations section is honest about n=2 and non-generalizability. The pen-portrait method and longitudinal design are appropriate for generating hypotheses. The concern is not that the study exists but that the central inference is stronger than the evidence. The reader's conditional verdict already captures much of this, so my recommendation is UNCHANGED. My specific addition is that even within the two selected cases, the claimed inverse relationships are not demonstrated: Pat's ethics appear constant-low rather than declining, and Jay's 'impeded skill development' is inferred from reduced use, not measured skill. This is a reasoning gap rather than a contradiction, and it can be addressed by re-analyzing the full sample and by coding the focal dimensions independently. The paper's contribution as an illustrative case series remains valuable; the title and conclusion should be scaled back unless the requested test supports the trade-offs.","tokens_in":11806,"tokens_out":4514,"duration_ms":48453,"concrete_test":"Have two independent coders, blind to the wicked-problem hypothesis, code all 14 participants' first and fourth interviews using a pre-registered codebook with separate ordinal scales for GenAI proficiency/use, ethical stance/behavior, career confidence, and explicitly articulated conflicts between goals. Then report (a) how many of the 14 exhibit each of the two claimed trade-off patterns, (b) how many exhibit counterexamples (e.g., increased ethical awareness without reduced GenAI skill use), and (c) whether Pat's ethics code declines from T1 to T4. If the two patterns appear in only a minority of participants, or Pat's ethics do not decline, the current title and abstract overstate the evidence and should be revised to describe two illustrative cases rather than a general wicked-problem property.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that progress on one learning goal impedes another—requires both (a) measured change on each dimension from Interview 1 to Interview 4 and (b) evidence that one dimension's gain caused another's loss. Neither is supplied. Section 3 says Pat and Jay were chosen because their accounts were 'the most detailed and thought-provoking,' which is selection on the outcome of interest; the other 12 participants are not analyzed, so we cannot tell whether these trade-offs are systemic or idiosyncratic. More seriously, the two case narratives do not actually show the claimed trade-offs. For Pat, §4.1.1 reports he described himself as a 'notorious cheater' at Interview 1 and in Interview 4 said 'I really didn't have any ethical views before, and I still don't really'—that is stable low ethical engagement, not 'increasing GenAI use skills can lower ethics.' For Jay, the paper infers 'impeded skill development' solely from a self-imposed 10-minute daily usage limit (§4.2.1); no GenAI skill measure is reported, and less use is not itself evidence of less skill acquisition. Likewise, 'career confidence' is inferred from narrative statements rather than a validated measure. The abstract and title generalize to 'designing for learning with GenAI is a wicked problem,' but the data are two purposively selected, self-reported trajectories with no evidence of causal trade-offs. The appropriate claim would be that these cases illustrate potential tensions worth studying, not that the problem is wicked.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues, based on a longitudinal qualitative study of two students in a GenAI-integrated creative media course, that designing for learning with generative AI is a 'wicked problem' in which progress on one educational goal (e.g., GenAI use skills) can impede progress on another (e.g., ethics or career confidence). The two case studies—Pat, who increased his use of GenAI while reporting unchanged low ethical engagement, and Jay, who developed ethical concerns and reduced his GenAI use—are presented as illustrations of such trade-offs. The paper uses Social Cognitive Career Theory (SCCT) as an interpretive lens, analyzes interviews from the beginning and end of the semester, and calls for multi-dimensional evaluation of GenAI-integrated curricula rather than optimizing any single outcome.","tokens_in":12101,"tokens_out":6558,"duration_ms":67667,"significance":"If the central claim were backed by the presented evidence, the paper would make an important contribution to computing education by cautioning against single-outcome evaluations of GenAI integration and highlighting potential feedback loops among learning, ethics, and career outcomes. The paper's strength is its rich, detailed narrative data from a course that explicitly integrated ethics instruction and its attention to understudied career-outcome dimensions. However, as presented, the evidence does not establish the claimed causal trade-offs: the two cases are selected on the outcome of interest, there are no direct measures of skill or ethical change, and the analysis collapses to two timepoints. The paper therefore functions best as a hypothesis-generating case study rather than as a demonstration of the wicked-problem claim; the title and abstract overstate what the data can support.","major_comments":[{"comment":"The selection of Pat and Jay because their accounts were 'the most detailed and thought-provoking' is a selection on the outcome of interest; without any analysis of the other 12 participants, the paper cannot support the general claim in the title that designing for learning with GenAI is a wicked problem, nor the Introduction's assertion that 'our findings demonstrate that in this setting, it exhibited clear characteristics of one.' The authors should either analyze and report on the full cohort (even concisely) or explicitly limit the claims to 'these two cases illustrate potential tensions' and adjust the title, abstract, and discussion accordingly.","section":"Section 3 (Method)"},{"comment":"Pat's narrative does not show 'increasing GenAI use skills can lower ethics' (Abstract). At Interview 1 he already describes himself as a 'notorious cheater' who avoided GenAI because if he started, 'I'm never gonna not use it'; at Interview 4 he states 'I really didn't have any ethical views before, and I still don't really.' The data indicate stable low ethical engagement, not a decline triggered by skill gains. The paper should either present evidence of temporal change in ethical stance or revise the claim to one about stability of low ethics under skill acquisition, which is a weaker and different finding.","section":"Section 4.1.1 (Pat)"},{"comment":"The claim that Jay's ethical awakening 'impeded skill development' is unsupported: the only evidence is his self-imposed 10-minute daily usage limit and reduced use. No GenAI skill measure (e.g., output quality, task performance, self-efficacy) is reported, and reduced use does not logically imply reduced skill. The causal chain from ethical concern to usage limit to skill deficit is assumed rather than demonstrated; the authors should provide direct or indirect skill evidence or rephrase the finding as a potential risk that warrants further study.","section":"Section 4.2.1 (Jay)"},{"comment":"Despite the longitudinal design with four monthly interviews, the analysis uses only interviews 1 and 4, as stated in Section 3: 'the analysis for this paper focuses specifically on the first and final interviews.' Two timepoints cannot reveal the feedback loops, 'evolving dilemmas,' and 'cumulative effects' that the wicked-problem framing requires (Section 1, Discussion). The authors should either analyze and report at least one intermediate interview or discuss why the intermediate data were excluded, and temper claims about change over time accordingly.","section":"Section 3 (Method) and Section 4 (Results)"},{"comment":"The a priori development of the codebook from SCCT and the introduction of the wicked-problem framing before data collection create a risk of circular interpretation: the analysis may only confirm the pre-existing framework. The paper should explain how the analysis allowed for disconfirming evidence and what alternative explanations (e.g., individual differences, pre-existing attitudes, course context) were considered. This concern does not invalidate the descriptive narratives but weakens the theoretical contribution as stated.","section":"Section 3.1 and Section 1"}],"minor_comments":[{"comment":"The phrase 'an GenAI-integrated' should be 'a GenAI-integrated' because 'GenAI' begins with a consonant sound.","section":"Abstract"},{"comment":"Reference [17] is cited awkwardly as 'Studies [17] like Yasar and Karagücük's' inside a sentence; add the author names or rephrase to 'Studies by Yasar and Karagücük [17]...'.","section":"Section 2"},{"comment":"The text uses both 'GenAI' and 'generative AI' inconsistently; choose one convention at first mention and use it consistently thereafter.","section":"Throughout"},{"comment":"The phrase 'contradicted with my intended meaning' should be 'contradicted my intended meaning'.","section":"Section 4.2.1"},{"comment":"The claim that this is 'the first longitudinal, in-depth qualitative study' of this kind requires a more precise comparison set and a defense that no prior longitudinal qualitative studies exist; as written, the claim is too strong.","section":"Section 1"},{"comment":"The limitations section does not mention the purposive selection of two extreme cases from a 14-participant cohort; this is a central limitation that should be acknowledged explicitly alongside the existing caveats about generalizability.","section":"Section 5 (Limitations)"}],"recommendation":"major_revision","confidential_remarks":"The paper has a potentially valuable case-study contribution, but the current evidence is not sufficient to support the broad wicked-problem claim. I would ask the editor to require a revision that either substantially expands the empirical base (e.g., analyzing the full 14-participant cohort to show the range of trajectories) or explicitly reframes the paper as an illustrative, hypothesis-generating study. The case-selection and missing-outcome-measure issues are load-bearing, not cosmetic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—quick take: this paper has two genuinely new longitudinal portraits of students in a GenAI-integrated course, and that part is worth reading. But the headline claim that designing for learning with GenAI is a wicked problem isn't established by the reported evidence. The stress-test note is right: Pat doesn't show progress on GenAI skills causing lower ethics. At interview 1 he already calls himself a \"notorious cheater\" and avoids GenAI; at interview 4 he still says he has no ethical views. That is stable low ethical engagement, not a trade-off generated by skill gains. Jay's \"impeded skill development\" is inferred from a self-imposed 10-minute daily limit; there's no measure of GenAI skill. Less use is not the same as less skill acquisition. Career confidence is similarly narrative.\n\nWhat is new: the longitudinal tracing itself. Pat's arc from avoidance to dependency, Jay's ethics-driven limits and growing career anxiety, are the kind of within-person change that cross-sectional surveys miss. The authors also position the work honestly—the limitations section notes self-report and the two-case design. The writing is clear and the related-work coverage is fair.\n\nSoft spots, in proportion: selection on the outcome. They picked Pat and Jay because their accounts were \"most detailed and thought-provoking,\" which biases toward extreme cases. We don't see the other 12 participants, so we can't know whether these patterns are systemic. No inter-rater reliability is reported, and the codebook was built from SCCT plus the wicked-problem lens, so part of the interpretation is theory-driven. These are real but not fatal for an illustrative case series. What is more serious is the mismatch between the abstract/title and what the data can support. \"Progress on one goal impedes another\" requires measured change on both dimensions and some evidence of relation. The paper has neither. The defensible claim is: these cases illustrate potential tensions worth studying, not that the problem is wicked.\n\nWho this is for: computing educators and researchers working on GenAI integration, especially people designing courses. The pedagogical discussion at the end is sensible.\n\nRecommendation: deserving of a serious referee. I'd send it out with the expectation of revision—re-scope the title and abstract to \"two illustrative cases,\" show the full cohort analysis or justify selection more carefully, and avoid causal trade-off language where only sequence or correlation exists. The case data are valuable; the framing needs to shrink.","headline":"Two rich case trajectories, but the wicked-problem trade-off claim overreaches the data: Pat's ethics were already low, Jay's skill loss is inferred from reduced use.","tokens_in":12589,"tokens_out":2047,"would_cite":false,"duration_ms":22329,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two student paths show why GenAI teaching is a wicked problem","keywords":["generative AI","wicked problem","student development","longitudinal qualitative","case series","ethics","career confidence","computing education"],"falsifier":"Re-analyze all fourteen participants' first and final interviews and count how many show a clear gain on one dimension matched by a decline on another; if most students improve on several dimensions at once, or if Pat and Jay are clear outliers, the claim that progress on one goal impedes another in this setting would not hold.","tokens_in":11613,"feed_emoji":"🤖","tokens_out":7230,"duration_ms":65137,"temperature":0.7,"pith_summary":"Designing courses around generative AI is a \"wicked problem,\" this paper argues: making progress on one educational goal can undermine another. Over a semester of interviews with students in a GenAI-integrated creative media course, the authors trace two trajectories showing this trade-off. One student's growing GenAI fluency coincided with a loss of ethical restraint and a self-described turn to cheating; another's deepening ethical awareness led to self-imposed usage limits that stalled skill growth and intensified career anxiety. The paper concludes that GenAI-integrated learning must be evaluated and designed multi-dimensionally, not by optimizing any single outcome.","feed_headline":"GenAI teaching trades skills for ethics, two cases show","feed_subtitle":"Interviews show one student's GenAI skills rising as ethics fell, another's ethics rising as skills stalled.","key_machinery":"The central object is the concept of a \"wicked problem\" from planning theory, applied to GenAI-integrated education. The argument is carried by a longitudinal qualitative case series: four semi-structured interviews with each of fourteen students, with the analysis focused on the first and final interviews and rendered as \"pen portraits\" that narrate each student's trajectory over the semester. Social Cognitive Career Theory supplies the interpretive lens, positing that self-efficacy, outcome expectations, and career confidence reinforce one another; the two cases show GenAI short-circuiting that loop. The specific mechanism that makes the case is the observed inverse relationship—skill gains accompanying ethical decline in one student, ethical gains accompanying skill stagnation in another.","core_discovery":"On the authors' own terms, the central claim is that a course deliberately built to teach GenAI skills, ethical reasoning, and career awareness together produced students in which progress on one of these goals impeded another. Pat began the semester avoiding GenAI and calling it \"trash,\" but by the final interview he was using it to \"get all the right answers,\" described himself as a \"notorious cheater,\" and stated he had no ethical views about the technology. Jay began confident that human writing could beat GenAI, but ethical concerns raised in the course—about environmental cost and artists' consent—led to a self-imposed ten-minute daily usage limit that curtailed skill development and made them \"nervous to do any type of writing professionally.\" The authors use these two cases to argue that GenAI-integrated education in this setting exhibited the defining features of a wicked problem: competing values, unpredictable outcomes, and evolving dilemmas without straightforward resolution.","pith_inferences":["A testable extension: courses that frame GenAI use through personal values and agency, rather than efficiency, may reduce the observed skills–ethics trade-off.","If the wicked-problem framing generalizes, institutional mandates to use GenAI in classrooms could produce unanticipated ethical and motivational harms for students who respond like Jay.","The cases suggest that Social Cognitive Career Theory may need enrichment: self-efficacy can be decoupled from actual learning when students attribute success to the tool rather than to themselves.","A larger-sample replication could quantify the trade-off and check whether it is robust or an artifact of selecting the two most dramatic cases."],"forward_implications":["Curricula should be evaluated on multiple dimensions—learning, ethics, motivation, and career confidence—rather than on any single outcome like GenAI proficiency.","Ethics instruction can suppress skill building when it drives students toward avoidance rather than reflective engagement.","Career confidence does not automatically follow from tool proficiency; students need explicit help connecting GenAI skills to their own career narratives.","Assessments of GenAI-integrated courses should include longitudinal measures of student development over time, not just end-of-task performance.","The \"illusion of competence\" can persist even when students explicitly acknowledge they are learning less, as Pat's case shows."],"supporting_citations":[{"why":"Defines \"wicked problems\" as ill-structured, value-laden dilemmas; this definition is the paper's central frame.","marker":"[25]"},{"why":"Presents Social Cognitive Career Theory, the model of self-efficacy and career confidence that the two cases are interpreted through.","marker":"[19]"},{"why":"Documents the \"illusion of competence\" in GenAI-assisted learning, which explains Pat's confidence despite acknowledging he learns less.","marker":"[24]"},{"why":"Reports complex relationships between GenAI use and novice programmers' self-efficacy and fear of failure, supporting the career-confidence findings.","marker":"[20]"},{"why":"Provides the \"pen portrait\" analytic technique used to construct the two longitudinal case narratives.","marker":"[29]"},{"why":"Describes the ethics activity that led Jay to impose a ten-minute daily usage limit, the mechanism behind Jay's stalled skill growth.","marker":"[18]"},{"why":"Documents how GenAI can erode social interactions and learning communities, background for the destructive cycle the paper describes.","marker":"[12]"}],"fun_headline_variants":["Two students reveal GenAI's wicked trade-offs","GenAI education's catch-22: skills vs ethics","Skills and ethics trade off in GenAI course","Ethics and skills at odds in GenAI education"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the two selected students, Pat and Jay, are representative enough of the fourteen participants to illustrate the claimed trade-offs; if they are atypical, the observed conflicts could be artifacts of case selection rather than systemic features of GenAI-integrated learning.","fun_headline_variants_meta":{"raw":{"variants":["Two students reveal GenAI's wicked trade-offs","GenAI education's catch-22: skills vs ethics","Skills and ethics trade off in GenAI course","Ethics and skills at odds in GenAI education"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001285,"raw_usage":{"total_tokens":5285,"prompt_tokens":1013,"completion_tokens":4272,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":4210}},"tokens_in":629,"tokens_out":4272,"duration_ms":33544,"temperature":1.0,"reasoning_tokens":4210,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:53:10.245619+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze all fourteen participants' first and final interviews and count how many show a clear gain on one dimension matched by a decline on another; if most students improve on several dimensions at once, or if Pat and Jay are clear outliers, the claim that progress on one goal impedes another in this setting would not hold.","supporting_citations":[{"cited_title":"W., and Webber, M","cited_arxiv_id":null,"evidence_quote":"Defines \"wicked problems\" as ill-structured, value-laden dilemmas; this definition is the paper's central frame."},{"cited_title":"W., and Brown, S","cited_arxiv_id":null,"evidence_quote":"Presents Social Cognitive Career Theory, the model of self-efficacy and career confidence that the two cases are interpreted through."},{"cited_title":"N., Leinonen, J., MacNeil, S., Randrianasolo, A","cited_arxiv_id":null,"evidence_quote":"Documents the \"illusion of competence\" in GenAI-assisted learning, which explains Pat's confidence despite acknowledging he learns less."},{"cited_title":"E., Prather, J., Reeves, B","cited_arxiv_id":null,"evidence_quote":"Reports complex relationships between GenAI use and novice programmers' self-efficacy and fear of failure, supporting the career-confidence findings."},{"cited_title":"How to analyse longitudinal data from multiple sources in qualitative health research: the pen portrait analytic technique","cited_arxiv_id":null,"evidence_quote":"Provides the \"pen portrait\" analytic technique used to construct the two longitudinal case narratives."},{"cited_title":"O., and Ko, A","cited_arxiv_id":null,"evidence_quote":"Describes the ethics activity that led Jay to impose a ten-minute daily usage limit, the mechanism behind Jay's stalled skill growth."},{"cited_title":"all roads lead to ChatGPT","cited_arxiv_id":null,"evidence_quote":"Documents how GenAI can erode social interactions and learning communities, background for the destructive cycle the paper describes."}],"review_version":1}