{"id":"8642f971-71fb-41c6-ac37-0268b0a2bb23","arxiv_id":"2509.00852","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The author argues that AI learning tools should be designed to require deliberate effort at certain stages, using the effort paradox and the IKEA effect to deepen learning and metacognition.","lead":"This position paper argues that students should sometimes be required to make more effort when using generative AI tools, rather than only exploiting them for convenience. It proposes designing GenAI interactions that deliberately create cognitive effort at certain stages to deepen learning and metacognition.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Effort→learning claim rests on extrapolation from decision-making tasks; no educational outcome measures connect tool-imposed effort to durable learning.","rationale":"The reader's weakest assumption correctly identifies the same load-bearing concern: the extrapolation from small decision-making studies and surveys to educational learning gains. I concur. The paper is a position paper, so UNVERDICTED is appropriate; the central claim is plausible but unverified. My concern doesn't change that verdict, but it sharpens the specificity: the empirical gap is not just sample size but outcome type—no study measures learning or metacognition as a function of effort imposed by tool design. A concrete test would resolve whether the claim is a useful design principle or a hopeful analogy.","tokens_in":6769,"tokens_out":3479,"duration_ms":42137,"concrete_test":"Run a pre-registered randomized experiment with, say, 120 students. Condition A: students use a GenAI tutoring tool that requires typing a rationale before receiving feedback (effort-imposing). Condition B: same tool with immediate feedback (no rationale). Measure: pre/post content knowledge, delayed retention (1 week), self-reported frustration, engagement, and metacognitive calibration. If Condition A does not significantly outperform B on retention/transfer, the central claim fails; if A only increases frustration, the design premise is wrong. This directly tests the extrapolation from the 20-participant decision study.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that requiring more effort when using GenAI 'could result in them learning more'—is unsupported because every cited empirical anchor measures something other than learning. The ExtendAI vs RecommendAI study (20 participants, investment decisions) measured decision outcomes, not knowledge retention, transfer, or metacognitive skill. The SelVReflect and VoiceViz studies measured reflection during a reasoning task, not learning gains. The Kreijkes et al. RCT measured comprehension and memory, but compared note-taking vs LLM use; it did not test a GenAI tool that imposes effort, and its result actually favored traditional note-taking over LLM use. The knowledge-worker survey (Lui et al., 2025) is self-reported workflow changes, not learning. The paper's own sentence 'Extrapolating from these findings into the domain of learning suggests that it can be beneficial' admits the leap. Moreover, the effort manipulation is confounded: in ExtendAI, writing a rationale is a self-explanation prompt, a known learning technique, so the benefit might come from the content of the reasoning rather than the 'effort cost.' The effort paradox/IKEA effect literature applies to tangible goods and successful completion; learning outcomes are intangible and delayed, so the analogy is untested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that students' over-reliance on generative AI (GenAI) for homework may undermine critical thinking and writing skills. The author proposes the 'effort paradox' and the IKEA effect as conceptual lenses for understanding why requiring more effort when using GenAI might deepen learning and produce a sense of achievement. The paper sketches five design directions for 'tools for thought' that deliberately impose effort, describes the author's prior empirical work with VoiceViz, SelVReflect, and ExtendAI vs. RecommendAI, and draws on literature by Sharples, Tang et al., Kreijkes et al., and Lui et al. The central claim is that shifting cognitive effort to later stages of a task (e.g., critiquing AI output) or combining GenAI with traditional note-taking can foster critical thinking and metacognition. The paper concludes that the effort paradox provides a useful framework for rethinking learning with GenAI, while acknowledging that more research is needed.","tokens_in":1910,"tokens_out":1719,"duration_ms":46582,"significance":"The paper addresses an important and timely question: how to reconcile the convenience of GenAI with the need to develop students' critical thinking and writing skills. Its strength is in framing the problem through the effort paradox and offering concrete design ideas (e.g., requiring rationale before AI feedback, constraining interactions, combining GenAI with note-taking). The author's own studies, though small, provide proof-of-concept demonstrations that tool design can shift users' reflective behavior. If the central claim were supported by educational outcome data, the framework would meaningfully contribute to the HCI and AI-in-education literature. However, as presented, the claim rests on extrapolation from non-educational decision-making studies and self-report surveys; no study in the manuscript directly tests whether tool-imposed cognitive effort improves learning or metacognitive skills in educational contexts. The paper is therefore best read as a design manifesto or hypothesis-generating essay rather than an evidence-backed finding.","major_comments":[{"comment":"The central claim that 'the additional effort involved could result in them learning more' (Abstract) is not directly supported by the cited studies. The ExtendAI vs. RecommendAI study involved 20 participants making simulated investment decisions and measured decision outcomes and self-reported reflection, not knowledge retention, transfer, or metacognitive skill. The paper itself acknowledges the leap: 'Extrapolating from these findings into the domain of learning suggests that it can be beneficial.' This extrapolation is load-bearing, so the strength of language in the Abstract and Summary ('have shown') should be tempered or supplemented with educational outcome measures.","section":"Designing new GenAI tools (ExtendAI paragraph)"},{"comment":"The Kreijkes et al. RCT is cited as supporting the idea of combining GenAI with note-taking, but that study compared traditional note-taking with LLM use, not a GenAI tool designed to impose effort. Its result actually favored note-taking for comprehension and memory, which says nothing about whether a deliberately effortful GenAI tool—rather than effortful traditional study strategies—yields learning gains. The distinction between effort imposed by tool design and effort from doing the task oneself is confounded, and the manuscript does not address this.","section":"Other opportunities for learning (Kreijkes et al. paragraph)"},{"comment":"The mechanism attributed to 'extra effort' is confounded with a known learning technique: writing one's own rationale before receiving AI feedback is a form of self-explanation/elaboration. Any benefit of ExtendAI could stem from the content of the reasoning produced, not from the perceived cost of effort. To support the effort-paradox mechanism, an effort-matched control (e.g., requiring the same amount of typing but without self-explanatory content) would be needed. Without this, the paper's central mechanism is not isolated.","section":"Designing new GenAI tools (ExtendAI description)"},{"comment":"The analogy to the IKEA effect and effort paradox relies on studies of tangible goods and successful completion (Norton et al. 2012; Inzlicht et al. 2018). Learning outcomes are intangible, delayed, and uncertain; there is a real risk that imposed effort leads to frustration or disengagement rather than a sense of achievement. The paper acknowledges the cost side only in passing ('which they found burdensome') and does not present any evidence on boundary conditions. This is a significant gap for the central claim and should be explicitly discussed as an open empirical question, not only as a design opportunity.","section":"The effort paradox"}],"minor_comments":[{"comment":"Typo: 'sometime using AI' should be 'sometimes using AI'.","section":"Designing new GenAI tools (item ii)"},{"comment":"Typo: 'critiquing and checking the validity of the this' should be 'of this' or 'of the output'.","section":"Designing new GenAI tools (ExtendAI paragraph)"},{"comment":"The phrase 'especially those whose English is not their second language' appears to be an error; it should likely be 'not their first language.'","section":"Introduction"},{"comment":"The typography 'Vo i c e Vi z' and 'S e l V R e f l e c t' is distracting; use regular formatting for tool names.","section":"Designing new GenAI tools (VoiceViz paragraph)"},{"comment":"The Freeman (2025) reference is incomplete: 'Student Generative AI Survey 202' should include the issue number and full title; the HEPI policy note number is listed but not the URL's accessible title. Minor citation formatting issues also appear in the Kreijkes et al. reference, which should include the SSRN preprint status.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a perspective/manifesto that would benefit from reframing as such in the title and abstract. The central claim is plausible but currently goes beyond the evidence; the author is encouraged to either gather pilot educational data or explicitly reposition the paper as a design agenda with testable hypotheses. I would not recommend rejection, as the topic is timely and the design directions are useful, but the overstatement in the Abstract and Summary should be corrected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my take. This is a position paper, not a research report, and it should be read as a design provocation. The fresh idea is to apply the effort paradox to GenAI learning: effort is both costly and valued, so we might design GenAI tools to deliberately shift students between low-effort and high-effort modes. The five design directions—tools for thought, slowing thinking down, externalizing half-baked ideas, constraining tasks, and building metacognition—are concrete enough to spark empirical work. The paper draws on a reasonable literature base, including Kreijkes et al., Sharples, Tang et al., and the author's own earlier tools (Proberbots, SelVReflect, ExtendAI).\n\nThe main soft spot is the evidence bridge. ExtendAI had 20 participants and measured investment decisions, not learning. SelVReflect measured reflection during a VR experience, not learning gains. Kreijkes et al. compared note-taking with LLM use; the effortful condition was traditional note-taking, not a GenAI tool that imposes effort. The knowledge-worker survey is self-report. None of these directly support the conclusion that imposed effort in GenAI interactions improves learning outcomes. To her credit, Rogers mostly hedges with 'suggests' and 'could', and even says the key question remains open. But the summary's 'I have shown' overstates the case; that sentence needs softening. A related gap is the lack of attention to failure modes: forced effort can frustrate students, push them to use other AI tools covertly, or disadvantage those with weaker prior knowledge. These deserve a paragraph in any revision.\n\nThe stress-test note is right about the thinness of the empirical anchors, but it slightly overreaches if it expects a position paper to supply outcome data. The value here is framing and agenda-setting. As such, it is a clear, readable piece that would be useful for people designing GenAI learning tools. I would send it to peer review—not to verify any new result, but to sharpen the claims and get the limitations and failure modes on the page.","headline":"A worthwhile design provocation that applies the effort paradox to GenAI learning, with the empirical support a little thinner than the summary lets on.","tokens_in":7494,"tokens_out":4220,"would_cite":true,"duration_ms":50433,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Requiring students to use generative AI more effortfully—writing rationales before feedback, critiquing output, taking handwritten notes—may deepen learning and give a greater sense of achievement.","keywords":["generative AI in education","effort paradox","IKEA effect","critical thinking","metacognition","human-AI interaction","learning design","cognitive effort"],"falsifier":"A randomized classroom experiment: one group writes a rationale before receiving AI feedback, another gets AI recommendations directly; if the rationale group does not score higher on delayed tests of understanding or reports more frustration, the claim that required effort deepens learning is weakened.","tokens_in":6717,"feed_emoji":"🧠","tokens_out":10380,"duration_ms":112316,"temperature":0.7,"pith_summary":"This paper argues that the widespread student habit of letting ChatGPT produce essays and homework at minimal effort is quietly eroding the thinking skills that writing itself builds. To counter that, it introduces the effort paradox—people avoid effort yet also value it, as in the IKEA effect—and applies it to learning: AI tools should sometimes make students work harder, not always easier. The paper proposes designing GenAI tools that stage effort across a task, letting AI help with getting started and then requiring students to evaluate, critique, verify, and build on what it generated. It supports this with early evidence from studies of AI assistants that scaffold reflection, a comparison of an advisory AI with one that demands users write their own rationale first, and a survey showing knowledge workers already shift their effort to verifying AI output. The payoff, if the argument holds, is a way to keep the convenience of GenAI without sacrificing critical thinking, metacognition, or the sense of accomplishment that comes from hard-won understanding.","feed_headline":"Design AI to demand more effort and students learn more","feed_subtitle":"Staging low-effort AI help early and high-effort critique later could build critical thinking.","key_machinery":"The paper's central mechanism is the \"effort paradox\"—the finding that effort is simultaneously costly and valued—deployed through a design pattern of staging cognitive effort across task phases: low effort when starting with GenAI (getting ideas, plans, drafts) and higher effort afterward (verifying, critiquing, iterating, note-taking). Named vehicles include \"Proberbots\" (chatbots that nudge reflection), \"ExtendAI\" (an AI that requires users to write their rationale before receiving feedback), and \"SelVReflect\" (a voice/VR tool for guided reflection), plus the \"IKEA effect\" as the motivational analogy explaining why invested labor increases felt value.","core_discovery":"The central claim is that the effort students avoid when they outsource homework to ChatGPT is the same effort that makes learning feel worthwhile, so the aim should not be to ban generative AI but to design it to deliberately require more effort at the right moments. The paper names this the effort paradox—effort is costly and also valued—and uses the IKEA effect, where labor invested in an object increases its value, to explain why effortful interaction could make learning more rewarding. Concretely, it proposes staging cognitive effort across a task: use GenAI for low-effort starts, plans, and explanations, then require students to invest more effort when evaluating, critiquing, iterating","pith_inferences":["One extension the author leaves implicit: the effort-staging design implies an inverted-U curve between imposed effort and learning, with too much required effort pushing students toward disengagement; experiments could identify the point at which effortful AI use stops being rewarding.","The argument suggests a product-design direction only gestured at in the paper: GenAI tools could offer an explicit \"deep mode\" that withholds recommendations until the user has written a rationale or made a prediction, turning effort into a feature rather than a bug.","If the effort paradox generalizes, it predicts that students who use AI to produce a draft will later value the final work more when they have revised it substantially themselves—an education-specific instance of the IKEA effect that could be measured with ownership and achievement surveys.","The paper's evidence comes from decision-making and short-term studies; an unstated corollary is that the metacognitive benefits depend on students noticing the value of the effort, so simply forcing effort without explanation or reflection may fail."],"forward_implications":["Students can learn to use GenAI with less effort at the start of a task and more effort later—checking, critiquing, and iterating—without losing the time-saving benefits.","Combining GenAI use with traditional methods such as handwritten note-taking can make complex material accessible while preserving the memory benefits of effortful note-making.","Educators can assess the process as well as the product by asking students to document their prompts, disagreements, revisions, and acceptance decisions.","AI assistants that ask users to articulate their own rationale before giving feedback can make decisions more reflective, at the cost of being burdensome—a cost students may nonetheless value.","Dialogic, persona-based uses of GenAI can provoke critical questioning and follow-up research in classrooms, rather than passive acceptance of output."],"supporting_citations":[{"why":"Establishes the effort paradox—effort is both costly and valued—which the paper uses as its core lens for learning with GenAI.","marker":"Inzlicht et al. (2018)"},{"why":"Supplies the IKEA effect, the analogy that labor invested in a task increases its felt value, used to explain why effortful AI use could feel rewarding.","marker":"Norton et al. (2012)"},{"why":"Provides the ExtendAI-versus-RecommendAI comparison, the paper's main empirical evidence that requiring users to state their rationale first yields more reflective decisions.","marker":"Reicherts et al. (2025)"},{"why":"Reports the VoiceViz proberbot study showing that proactive questioning prompts deeper collaborative reasoning, supporting the scaffolding design.","marker":"Reicherts et al. (2022a)"},{"why":"Describes SelVReflect, the voice/VR guided-reflection tool whose 20-participant study indicates that scaffolded prompts foster self-understanding.","marker":"Wagener et al. (2023)"},{"why":"Provides evidence that handwritten note-taking with LLM support improves comprehension and retention, backing the mixed-tools recommendation.","marker":"Kreijkes et al. (2024)"},{"why":"Documents knowledge workers shifting effort from information gathering to verification, the real-world pattern the paper generalizes to student learning.","marker":"Lui et al. (2025)"},{"why":"Reports the secondary-school dialogic GenAI study showing students engage in critical questioning, supporting conversational tool designs.","marker":"Tang et al. (2024)"}],"fun_headline_variants":["Effort paradox: why harder AI use teaches more","Make students work harder with AI, not less","Design AI to cost effort, not save it","The IKEA effect for learning with GenAI","More effort with AI, more learning payoff"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise—which the paper states explicitly as \"Extrapolating from these findings into the domain of learning suggests that it can be beneficial\"—is that extra effort imposed by AI tool design converts into learning and achievement rather than frustration, a link drawn from small decision-making studies and a survey rather than from direct learning-outcome experiments.","fun_headline_variants_meta":{"raw":{"variants":["Effort paradox: why harder AI use teaches more","Make students work harder with AI, not less","Design AI to cost effort, not save it","The IKEA effect for learning with GenAI","More effort with AI, more learning payoff"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000149,"raw_usage":{"total_tokens":1044,"prompt_tokens":772,"completion_tokens":272,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":516,"completion_tokens_details":{"reasoning_tokens":200}},"tokens_in":516,"tokens_out":272,"duration_ms":3935,"temperature":1.0,"reasoning_tokens":200,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:07:52.587243+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A randomized classroom experiment: one group writes a rationale before receiving AI feedback, another gets AI recommendations directly; if the rationale group does not score higher on delayed tests of understanding or reports more frustration, the claim that required effort deepens learning is weakened.","supporting_citations":[],"review_version":1}