{"id":"2b1c67b5-5b87-438b-babe-4e589c81148a","arxiv_id":"2102.09692","paper_version":1,"verdict":"ACCEPT","confidence":"LOW","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Cognitive forcing interventions reduce overreliance on AI recommendations more than simple explanations, with effects moderated by individual need for cognition.","lead":"The study tested three cognitive forcing designs meant to make users think harder about AI suggestions instead of accepting them automatically. These designs cut overreliance on incorrect AI advice more than standard explanations, though users gave them lower satisfaction scores and the benefit was larger for people who enjoy effortful thinking.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No direct measures confirm that forcing interventions increased analytical engagement vs. friction-induced caution","rationale":"The reader's weakest assumption matches the mechanism gap exactly. This does not falsify the empirical reduction in overreliance but requires the interpretation to be qualified as 'effective at reducing overreliance' rather than 'effective via deeper analytical engagement.' The low reader confidence is warranted given the abstract-only review; full-text verification of process measures would resolve it.","tokens_in":1734,"tokens_out":304,"duration_ms":21798,"concrete_test":"In the methods/results sections, locate any secondary engagement metrics (time reading explanations, number of explanation views, or post-task effort ratings). If absent, recompute overreliance reduction while controlling for subjective difficulty ratings; if the effect size drops below significance after this control, the mechanism claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim attributes reduced overreliance to compelled deeper processing (dual-process theory). The design compares outcome metrics across conditions but reports no direct process measures of engagement depth—such as response latencies on explanations, eye-tracking dwell time, think-aloud protocols, or comprehension probes tied to the AI rationale. Without these, the observed drop in overreliance could arise from non-analytic mechanisms (e.g., added steps raising perceived cost of following AI, or general slowing). The Need-for-Cognition moderation is consistent with differential susceptibility but does not isolate the intended cognitive pathway.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript reports results from a controlled experiment (N=199) comparing three cognitive forcing interventions—designed to promote analytical engagement with AI explanations per dual-process theory—to two simple explainable-AI baselines and a no-AI condition. It claims that the forcing designs significantly reduce overreliance on incorrect AI recommendations, albeit with lower subjective ratings, and that the benefit is moderated by Need for Cognition such that higher-NFC participants gain more from the interventions.","tokens_in":1816,"tokens_out":453,"duration_ms":48715,"significance":"If the core empirical result holds, the work supplies concrete design evidence that cognitive forcing can mitigate overreliance in AI-assisted decisions, documents a satisfaction trade-off, and identifies cognitive motivation as a moderator relevant to equitable XAI deployment.","major_comments":[{"comment":"The central interpretation—that reduced overreliance results from compelled deeper analytical processing rather than non-analytic mechanisms such as added friction—is load-bearing for the dual-process framing, yet the design reports only outcome metrics (overreliance rates) without direct process measures (response latencies on explanations, eye-tracking dwell times, or comprehension probes of the AI rationale). This leaves the mechanism unverified.","section":"Methods and Results sections"},{"comment":"The moderation analysis by Need for Cognition is presented as evidence of differential benefit, but the manuscript does not report the full regression model (including interaction term, covariates, and effect-size details) or power calculations for the subgroup comparisons, making it difficult to assess whether the reported average benefit for higher-NFC participants is robust.","section":"Results section"}],"minor_comments":[{"comment":"The abstract states the sample size and key comparisons but omits the decision task domain and the precise operationalization of overreliance; adding one sentence would improve standalone readability.","section":"Abstract"},{"comment":"Subjective rating scales are mentioned but the exact items, anchors, and reliability statistics are not tabulated; a supplementary table would clarify the reported trade-off.","section":"Results"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed comments. We address each major point below, indicating planned revisions to improve clarity, transparency, and completeness of the manuscript.","responses":[{"response":"We agree that the absence of direct process measures (e.g., response latencies, eye-tracking, or comprehension probes) leaves the precise cognitive mechanism somewhat inferential rather than directly verified. Our study was designed to evaluate behavioral outcomes of the forcing interventions, which were constructed to interrupt heuristic reliance and require explicit engagement with the AI rationale, consistent with dual-process theory and prior medical decision-making research. We cannot rule out non-analytic factors such as added friction with the current data. In revision we will (1) expand the Discussion to explicitly address alternative mechanisms, (2) add a dedicated Limitations subsection noting the lack of process-tracing data, and (3) propose future studies that incorporate such measures. This is a partial revision focused on improved interpretation and transparency.","revision_made":"partial","referee_comment":"[Methods and Results sections] The central interpretation—that reduced overreliance results from compelled deeper analytical processing rather than non-analytic mechanisms such as added friction—is load-bearing for the dual-process framing, yet the design reports only outcome metrics (overreliance rates) without direct process measures (response latencies on explanations, eye-tracking dwell times, or comprehension probes of the AI rationale). This leaves the mechanism unverified."},{"response":"We thank the referee for highlighting this reporting gap. The original manuscript summarized the NFC moderation but omitted the full model specification. In the revised version we will include the complete regression results: the full model equation with the condition × NFC interaction term, all covariates, coefficient estimates with confidence intervals, effect-size metrics, and any available power or robustness checks for the subgroup analyses. This will allow readers to evaluate the strength and stability of the moderation finding directly.","revision_made":"yes","referee_comment":"[Results section] The moderation analysis by Need for Cognition is presented as evidence of differential benefit, but the manuscript does not report the full regression model (including interaction term, covariates, and effect-size details) or power calculations for the subgroup comparisons, making it difficult to assess whether the reported average benefit for higher-NFC participants is robust."}],"tokens_in":1377,"tokens_out":489,"duration_ms":49857,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that the three cognitive forcing interventions reduced overreliance on wrong AI suggestions more than the simple XAI baselines, though participants rated those forcing designs lowest and the gains went mostly to people high in Need for Cognition. The paper adapts forcing functions from medical decision-making to general AI-assisted tasks and tests three concrete designs in a controlled experiment with 199 participants against two explanation conditions and a no-AI baseline. The outcome data show lower overreliance in the forcing arms, which is a straightforward empirical result worth having on record. They also checked for differential effects by NFC and reported that higher-NFC participants benefited more on average, which addresses one fairness angle directly. The experiment setup itself looks clean with clear condition comparisons and a reasonable sample size for the comparisons they ran. The main limitation is the missing process evidence. The design measures final accuracy and overreliance rates but does not include response times on the explanations, comprehension checks, or any other direct indicator that people actually engaged more analytically with the AI rationale rather than simply slowing down because of added steps. That leaves open the possibility that the effect comes from friction or caution rather than the intended dual-process shift. The lower subjective ratings for the conditions that performed best on overreliance is another practical issue that would matter in deployment. This paper is for researchers working on human-AI decision support and XAI evaluation. It supplies specific intervention examples and documents both the benefit and the trade-off in one study. I would send it to peer review. The empirical comparison is solid enough to deserve referee time even if the mechanism story needs tightening in revision.","headline":"Cognitive forcing cuts overreliance more than plain explanations but without process measures to confirm deeper thinking and with clear usability costs.","tokens_in":2284,"tokens_out":391,"would_cite":true,"duration_ms":36305,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"Foundation.LawOfExistence","rs_theorem":null,"paper_passage":"Informed by the dual-process theory of cognition, we posit that people rarely engage analytically with each individual AI recommendation and explanation, and instead develop general heuristics about whether and when to follow the AI suggestions."},{"relation":"unclear","rs_module":"Foundation.DiscretenessForcing","rs_theorem":null,"paper_passage":"The results demonstrate that cognitive forcing significantly reduced overreliance compared to the simple explainable AI approaches."}],"headline":"HCI study on cognitive forcing for AI overreliance is orthogonal to RS cost-minimization framework","alignment":"orthogonal","rationale":"The paper empirically tests dual-process interventions to reduce AI overreliance via explanations and forcing functions, focusing on behavioral outcomes and Need-for-Cognition moderation. It does not engage RS concepts such as J-cost uniqueness, defect collapse, ledger forcing, or phi-ladder geometry. No passages reference recognition cost, 8-tick periodicity, or the Law of Existence; the work remains in standard HCI/psychology territory without tapping RS-shaped structures like cosh-cost reasoning or ratio symmetry.","tokens_in":280598,"confidence":"high","tokens_out":289,"duration_ms":41656,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"lean_confirmation":{"model":"grok-4.3","status":"out_of_scope","citations":[],"rationale":"The paper is an empirical HCI/cs.HC study with no load-bearing mathematical claim. Shape-of-logic is a Lean corpus on foundational physics, logic forcing, and structural theorems across domains; it has no theorems relevant to cognitive forcing functions or AI overreliance. Status is therefore out_of_scope.","tokens_in":280369,"confidence":"moderate","tokens_out":197,"duration_ms":38165,"inferential_bridge":"The paper's central result is an empirical finding from an MTurk experiment (N=199) comparing cognitive forcing designs to simple XAI baselines. No mathematical or structural premise is present that could be machine-checked in Lean; the result rests on dual-process theory of cognition and measured participant behavior. Shape-of-logic contains no theorems about cognitive forcing, overreliance, or HCI decision-making.","load_bearing_premise":"Cognitive forcing functions compel deeper analytical engagement with AI explanations (reducing overreliance on incorrect AI suggestions in decision-making tasks).","cache_read_input_tokens":64,"cache_creation_input_tokens":0},"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Cognitive forcing interventions reduce overreliance on incorrect AI suggestions by prompting deeper analysis.","keywords":["cognitive forcing","overreliance","explainable AI","AI-assisted decision-making","dual-process theory","human-AI interaction","need for cognition","trust in AI"],"falsifier":"An experiment that measures actual analytical processing (for example via eye-tracking or think-aloud protocols) and finds no increase in depth of engagement despite the forcing designs.","tokens_in":2635,"feed_emoji":"🤔","tokens_out":584,"duration_ms":34190,"temperature":0.7,"pith_summary":"People often accept AI recommendations even when wrong because they apply general heuristics instead of analyzing each case and its explanation. The paper tests three cognitive forcing designs that require users to engage analytically with the AI output before deciding. In an experiment with 199 participants, these designs cut overreliance compared with standard explainable AI approaches. The designs that worked best received the lowest user satisfaction ratings, and their benefits were larger for participants who score high on Need for Cognition. The work therefore shows that explainable AI success depends on whether users are motivated to think through the provided information.","feed_headline":"Cognitive forcing reduces overreliance on AI suggestions","feed_subtitle":"Forcing users to analyze explanations lowers acceptance of wrong AI advice, though users rate those interfaces lower.","key_machinery":"Cognitive forcing interventions: interface designs that require users to perform additional analytical steps with the AI explanation before accepting or rejecting the suggestion.","core_discovery":"Cognitive forcing functions compel people to engage more thoughtfully with AI-generated explanations rather than relying on heuristics, and this engagement significantly reduces overreliance on wrong AI suggestions relative to simple explainable AI baselines, although it lowers subjective satisfaction and benefits people higher in Need for Cognition more.","pith_inferences":["Forcing mechanisms might transfer to high-stakes domains such as medical or financial decisions where overreliance carries larger costs.","Designers could explore milder versions of forcing that preserve user satisfaction while still increasing analysis.","Personalizing the level of forcing based on a user's measured Need for Cognition could improve both effectiveness and acceptance."],"forward_implications":["Cognitive forcing can be used to lower acceptance of erroneous AI advice in decision tasks.","Any reduction in overreliance comes with lower subjective ratings of the system.","The benefit of forcing is moderated by individual differences in motivation to think effortfully.","Explainable AI solutions will not work equally well for all users without accounting for cognitive motivation."],"fun_headline_variants":["Cognitive forcing reduces overreliance on faulty AI","Thoughtful AI analysis cuts overreliance via forcing","Cognitive forcing functions lower wrong AI acceptance","Overreliance falls with cognitive forcing interventions"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The three interventions actually compel deeper analytical thinking rather than simply adding friction or prompting other behavioral changes.","fun_headline_variants_meta":{"raw":{"variants":["Cognitive forcing reduces overreliance on faulty AI","Thoughtful AI analysis cuts overreliance via forcing","Cognitive forcing functions lower wrong AI acceptance","Overreliance falls with cognitive forcing interventions"]},"model":"grok-4.3","cost_usd":0.006447,"raw_usage":{"total_tokens":2947,"prompt_tokens":683,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":64465500,"prompt_tokens_details":{"text_tokens":683,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2209,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":683,"tokens_out":55,"duration_ms":32000,"temperature":1.0,"reasoning_tokens":2209,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-15T17:58:59.505519+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An experiment that measures actual analytical processing (for example via eye-tracking or think-aloud protocols) and finds no increase in depth of engagement despite the forcing designs.","supporting_citations":[],"review_version":1}