{"id":"326bc116-078f-40a7-950b-a45bc18e992f","arxiv_id":"2608.08882","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper introduces a framework and experimental protocol for measuring whether AI-assisted verification leaves users with improved, unchanged, or weakened independent judgment on later claims.","lead":"This paper argues that AI fact-checking tools are usually tested only while they are switched on. It proposes two measures, the Epistemic Transfer Effect and Tool-Removal Cost, to capture what people can still do on their own after the tool is gone.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The protocol's matched removal probe violates the no-practice control: Phase 3 gives the no-practice arm an unassisted verification block, so ETE versus no-practice does not estimate the advertised contrast.","rationale":"The reader's weakest assumption focused on the validity of the delayed test as a measure of retained skill, citing item memorability, test-retest practice, and domain familiarity. That is a legitimate construct-validity concern, but it does not identify the more concrete protocol-level contradiction: the removal probe, which the paper itself flags as a learning event, is applied to the no-practice control in a way that turns it into a practice condition. This is a direct inconsistency between Table 1, Section 3.2, Section 4.6, and the Phase 3 instructions in Section 4.4. Because the paper explicitly lists 'each AI condition versus no practice' as one of the two main ETE contrasts, this contamination is load-bearing for the advertised measurement protocol. The central conceptual claim that point-of-use performance should not be equated with learning survives, and the active-practice-based diagnostic space in Figure 2 remains usable, so rejection is not warranted. However, acceptance should be conditional on either demonstrating empirically that the matched probe has no differential learning effect for the no-practice arm, or revising the protocol to include a genuine no-practice comparator that does not receive unassisted verification practice. This is a concrete, fixable flaw rather than a fundamental objection to the framework, and it deserves a conditional verdict rather than an unconditional accept.","tokens_in":11107,"tokens_out":7765,"duration_ms":83349,"concrete_test":"Run a small three-arm pilot in which no-practice participants are randomized to (a) the matched unassisted probe block exactly as specified in Section 4.4, Phase 3, or (b) an equal-length non-verification filler activity, followed by the same delayed unassisted test. If delayed accuracy or discernment differs between (a) and (b) by more than the preregistered smallest effect of interest, the Section 4.4 protocol's ETE-versus-no-practice estimate is contaminated. The fix would be to relabel the comparator as 'minimal unassisted practice' or to add a separate no-probe no-practice arm in a two-wave design.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Table 1 defines the no-practice control as completing an unrelated matched-duration activity, and Section 3.2/4.6 treat 'versus no practice' as one of the two main ETE contrasts. Yet Section 4.4, Phase 3, instructs that 'the active-practice and no-practice groups should complete a matched unassisted probe block of equal length.' Unassisted verification is practice; it is the exact behavior the no-practice arm is supposed to avoid. The no-practice condition therefore becomes a minimal one-session unassisted practice condition, and the ETE(c, no-practice) contrast in Eq. (2) is estimated against this contaminated comparator, not against doing nothing. The paper acknowledges that the probe can act as an additional learning event and proposes matched exposure to keep conditions comparable, but equal exposure does not imply equal learning: giving the no-practice group unassisted retrieval practice may differentially raise its delayed performance, shrinking or even reversing the estimated ETE. The limitation section discusses designs that omit the probe to avoid this influence, but that sacrifices TRC; the joint protocol as written cannot simultaneously deliver a clean no-practice ETE and a TRC estimate in a single study. This does not invalidate the active-practice contrast or the conceptual framework, but it undermines one of the two primary ETE comparisons and the paper's RQ1 framing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a conceptual framework and an evaluation protocol for studying what users retain from AI-assisted verification tools. It defines epistemic transfer as the effect of prior interaction with an AI verification system on later unassisted performance with novel claims, and introduces two complementary quantitities: the Epistemic Transfer Effect (ETE), comparing delayed unassisted performance across conditions, and the Tool-Removal Cost (TRC), measuring the immediate within-person drop when the tool is removed. The protocol specifies four practice conditions (answer-first AI, evidence-first AI, active practice, no-practice control), a Phase 3 immediate removal probe, and a delayed unassisted test 7–14 days later, together with a diagnostic space of four descriptive profiles. The paper is explicitly a working paper and a proposal; it provides no empirical validation, but it does provide a detailed design, analysis recommendations (mixed-effects models, equivalence tests), and a discussion of boundary conditions and limitations.","tokens_in":11376,"tokens_out":4510,"duration_ms":45732,"significance":"If adopted, the framework would give the field a common vocabulary and a concrete protocol for moving beyond point-of-use evaluations of AI assistance. The separation of ETE and TRC is conceptually useful, and the diagnostic space makes visible a policy-relevant distinction between 'verification on loan' and genuine capability building. The paper is well grounded in cognitive-offloading, learning, and automation research, and it is transparent about the proposal status, the lack of empirical validation, and several measurement limitations. Its practical value is as a design template for future preregistered studies; that value is considerable, provided the protocol's internal inconsistency regarding the no-practice control is resolved.","major_comments":[{"comment":"The no-practice control is described in Table 1 as completing an unrelated matched-duration activity, but Phase 3 instructs that 'the active-practice and no-practice groups should complete a matched unassisted probe block of equal length.' This turns the no-practice arm into a minimal unassisted-practice condition, so the ETE(c, no-practice) contrast defined in Eq. (2) does not estimate the advertised comparison against no practice. Because unassisted verification is retrieval practice, this contamination may differentially raise the no-practice group's delayed performance and shrink or even reverse the estimated ETE, undermining the second of the two primary ETE comparisons and the RQ1 framing. The protocol should either restrict the unassisted probe to the AI conditions and collect TRC on a separate sample, or explicitly redefine the comparator as a 'no-AI, minimal-practice' condition and adjust the interpretation of the ETE(c, no-practice) contrast accordingly.","section":"Section 4.4 Phase 3; Section 5.1 Eq. (2); Table 1"},{"comment":"The framework treats a single delayed unassisted test, administered 7–14 days after practice, as the measure of retained capability. No evidence is provided for the reliability or stability of this measurement, and the delayed test itself—like the Phase 3 probe—is an additional learning event whose effect can interact with condition. The paper acknowledges that the probe adds retrieval practice, but it does not address the analogous issue for the delayed test. To support the claim that ETE is a valid and useful measurement, the protocol should include a test-retest component, multiple delayed-test waves, or a psychometric analysis of delayed-test scores, or at least specify how the interpretational risk would be handled in the analysis plan.","section":"Section 3.1 and Section 4.4 Phase 4"}],"minor_comments":[{"comment":"The feasibility paragraph states that detecting a three-to-five percentage-point difference 'will typically require several hundred participants per condition' and recommends 1,200–2,000 total participants, but no simulation details, code, or parameter assumptions are provided; because the numbers are labeled 'illustrative rather than prescriptive,' the wording 'will typically require' gives them unwarranted concreteness and should be softened or backed by a supplementary simulation.","section":"Section 4.7"},{"comment":"The TRC estimand is defined as a difference in expectations across tool-available and tool-removed states, but the equation does not specify the within-person pairing of items; while Section 4.4 mentions randomization and counterbalancing, making the within-person matched-item structure explicit in Eq. (3) would prevent ambiguity.","section":"Section 5.2 Eq. (3)"},{"comment":"The paragraph beginning 'Participants should match the intended users' appears twice, once before and once after Figure 1, and the figure itself is dense; consider removing the duplicate and moving the detailed sampling text to a subsection so the figure can be read more easily.","section":"Figure 1 and surrounding text"},{"comment":"Kothe and Ling (2019) is a PsyArXiv preprint; the claim about high retention of panel participants over 7–14 days would be better supported by a peer-reviewed source or by a more cautious statement acknowledging that retention rates vary widely across panels and populations.","section":"References"},{"comment":"The phrase 'the treated claim' is standard in the correction literature but may be opaque to HCI readers; a brief gloss such as 'the claim that was the target of the correction' would improve accessibility.","section":"Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written methods proposal that is honest about being a working paper. The conceptual distinction between ETE and TRC is likely to be useful to the community, but the protocol as written cannot simultaneously deliver a clean no-practice ETE and a TRC estimate in one study; this is a fixable design issue rather than a fatal flaw. The paper would also benefit from addressing the measurement-reliability question for the delayed test. The preprint citations (e.g., Kothe & Ling, Liu et al., Shen & Tamkin) should be checked for later peer-reviewed versions before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core contribution is real: the paper names a neglected outcome, epistemic transfer, and gives it two concrete, measurable quantities (ETE and TRC) plus a diagnostic space that separates capability building from verification on loan. That synthesis is overdue, and the paper does it cleanly. The protocol is actionable, the distinctions from trust, reliance, and team performance are sharp, and the literature coverage is current and well chosen, including the striking clinical deskilling result from colonoscopy. The paper knows it is a proposal and says so; that honesty counts. The soft spot is the no-practice condition. Table 1 defines no-practice as an unrelated matched-duration filler activity, and RQ1 explicitly contrasts against no practice. But Phase 3 instructs the no-practice group to complete a matched unassisted probe block, equal in length to the other arms. Unassisted verification is practice; it is exactly the behavior the no-practice manipulation is supposed to keep out. So ETE(c, no-practice) in Eq. (2) is not system use vs. doing nothing; it is system use vs. one short session of unassisted verification. Equal exposure does not mean equal learning. The paper acknowledges in the limitations that the probe adds retrieval practice for everyone and may compress differences, but that underestimates the problem: the comparator itself changes meaning. This undermines one of the two primary ETE contrasts and the RQ1 framing, though not the active-practice contrast or the overall framework. It is fixable: omit the probe for the no-practice arm, or explicitly reframe the contrast as vs. minimal unassisted session practice. Other concerns are minor for a working paper. There is no empirical validation, and the power figures in Section 4.7 are illustrative without simulation details; that is fine as long as readers treat this as a proposal. The delayed-test measurement is plausible but untested, as the reader noted. Who benefits: HCI, CSCW, misinformation, and human-AI interaction researchers who need a shared metric for evaluating assistant tools. This deserves a serious referee; I would also use it as a starting point for discussion. The no-practice contamination is the thing I would push the authors to fix before publication.","headline":"A genuinely useful framework for measuring what AI verification tools leave behind, but the protocol's own no-practice control is contaminated by the immediate removal probe, which gives that arm unassisted verification practice.","tokens_in":737,"tokens_out":647,"would_cite":true,"duration_ms":21564,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI tools should be judged by what users can still do without them.","keywords":["epistemic transfer","AI-assisted verification","fact-checking evaluation","cognitive offloading","tool-removal cost","human-AI interaction","evaluation methodology","de-skilling"],"falsifier":"Run the protocol twice, once with and once without the immediate removal probe, holding the practice phase identical; if the ETE contrast between an AI condition and active practice appears only in the version with the probe, then the probe's extra unassisted retrieval practice, not the AI assistance itself, is driving the measured transfer effect.","tokens_in":1683,"feed_emoji":"🧠","tokens_out":3214,"duration_ms":52614,"temperature":0.7,"pith_summary":"The paper argues that evaluating AI verification tools only while the tool is present can mislead: strong assisted performance does not show what users take away. It defines epistemic transfer as the effect of prior AI-assisted verification on later unassisted performance with new claims, and proposes two quantities—the Epistemic Transfer Effect (ETE) for delayed unassisted performance and the Tool-Removal Cost (TRC) for the immediate drop when help disappears. These two measures together sort systems into four profiles: capability building, capability plus tool advantage, verification on loan, and epistemic inertness or de-skilling. The paper also lays out a five-phase evaluation protocol with answer-first and evidence-first AI conditions, active-practice and no-practice controls, and a delayed 7 to 14 day test on held-out claims. A sympathetic reader would take the core point to be: point-of-use accuracy is not a proxy for what users learn.","feed_headline":"Measure AI tools by what users retain, not just live accuracy","feed_subtitle":"A proposed protocol separates real capability building from performance rented from the tool.","key_machinery":"The load-bearing object is the definition of epistemic transfer together with the two estimands built from it. Epistemic transfer is defined in Section 3.1 as the effect of prior interaction with an AI verification system on a person's later unassisted performance when evaluating new claims, with three defining features: no tool at test, novel claims, and a retention interval. ETE(c,k,d,b) is the difference in delayed unassisted performance between an AI condition and a comparator (active practice or no practice); TRC(c) is the within-person difference between matched probe items with and without the tool. These feed a five-phase protocol—baseline, practice, immediate removal probe, delayed test, debriefing—and a mixed-effects model with participant and item random intercepts. The mechanism doing the work is the joint reading of ETE and TRC, which distinguishes learning from rented performance.","core_discovery":"The central claim is that a system's value for independent judgment is not captured by how well human and AI perform together at the moment of use. The paper introduces epistemic transfer—the effect of prior interaction with an AI verification system on later unassisted performance on new claims—as a distinct outcome, and defines ETE and TRC as complementary estimands: ETE compares delayed unassisted performance across conditions, while TRC measures the immediate within-person drop when the tool is removed. Crossing ETE against active practice with TRC yields a diagnostic space separating capability building, capability plus tool advantage, verification on loan, and epistemic inertness or de-skilling. The paper's own conclusion states the claim directly: AI assistance does not have to teach in every setting, but we should stop assuming that strong point-of-use performance tells us what users learn.","pith_inferences":["If ETE and TRC are adopted, the same protocol could be adapted to other AI-assisted tasks such as search, medical diagnosis, or programming by swapping the verification task for the target skill, and the four profiles would likely map onto existing findings about tutor-versus-answer systems.","The removal probe's dual role as measurement and learning event suggests a testable refinement: vary probe length across conditions to estimate and subtract its contribution to ETE.","The framework implies a design target: interface features that maximize ETE relative to active practice, even at some cost to TRC, may be worth optimizing for when independent judgment matters.","When a system repeatedly scores as verification on loan, designers could add periodic unassisted retrieval practice to convert rented performance into retained capability."],"forward_implications":["Evaluations of AI verification tools should report the comparator, delay, access regime, item novelty, transfer distance, and outcome family, or two studies can appear to disagree while actually measuring different things.","The choice between answer-first and evidence-first support becomes an empirical trade-off: a fluent interface may maximize short-term performance while producing little transfer, and a more demanding interface may do the opposite.","A system can show a large TRC while still building capability, so tool advantage alone is not harm; classification requires uncertainty intervals and preregistered smallest effects of practical interest.","Procurement and policy settings can ask whether repeated use improves or weakens later independent performance, surfacing problems such as the reported decline in unassisted adenoma detection before deployment.","Equivalence tests, not null-hypothesis significance tests alone, are needed to claim that a system has no transfer effect."],"supporting_citations":[{"why":"Supplies the learning-versus-performance distinction that motivates separating assisted performance from retention.","marker":"Soderstrom and Bjork (2015)"},{"why":"Supplies retrieval-practice evidence that supports the delayed-test and active-practice logic.","marker":"Roediger III and Karpicke (2006)"},{"why":"Supplies the cognitive-offloading account that explains why tool availability may shift effort and memory.","marker":"Risko and Gilbert (2016)"},{"why":"Supplies the automation deskilling premise that high support can reduce opportunity to maintain skill.","marker":"Bainbridge (1983)"},{"why":"Provides an empirical example of generative AI improving practice performance but reducing post-removal performance in mathematics.","marker":"Bastani et al. (2025)"},{"why":"Provides real-world clinical evidence of unassisted performance decline after routine AI use in colonoscopy.","marker":"Budzyń et al. (2025)"},{"why":"Provides evidence that imperfect LLM guidance can reduce headline discernment.","marker":"DeVerna et al. (2024)"},{"why":"Shows a dialogue intervention that reduced false beliefs about discussed claims without building lasting discernment for unseen claims.","marker":"Rani et al. (2026)"},{"why":"Supplies the taxonomy of transfer distance used for near, intermediate, and far delayed-test claims.","marker":"Barnett and Ceci (2002)"},{"why":"Provides the equivalence-testing method for claims of practically null transfer effects.","marker":"Lakens (2017)"}],"fun_headline_variants":["Test AI tools by what they leave behind","Measure AI by what users retain after removal","AI verification: building skills or borrowing them?","Epistemic transfer: a new way to test AI tools","What users can do without AI: the real metric"],"cache_read_input_tokens":13952,"weakest_assumption_plain":"The framework assumes that one unassisted test on held-out claims, given 7 to 14 days after practice, is a stable and valid measure of what a user has actually retained; the paper offers no reliability evidence for this measure and notes that the removal probe itself can act as an additional learning event.","fun_headline_variants_meta":{"raw":{"variants":["Test AI tools by what they leave behind","Measure AI by what users retain after removal","AI verification: building skills or borrowing them?","Epistemic transfer: a new way to test AI tools","What users can do without AI: the real metric"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000283,"raw_usage":{"total_tokens":1680,"prompt_tokens":962,"completion_tokens":718,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":645}},"tokens_in":578,"tokens_out":718,"duration_ms":7403,"temperature":1.0,"reasoning_tokens":645,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:20:59.869641+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the protocol twice, once with and once without the immediate removal probe, holding the practice phase identical; if the ETE contrast between an AI condition and active practice appears only in the version with the probe, then the probe's extra unassisted retrieval practice, not the AI assistance itself, is driving the measured transfer effect.","supporting_citations":[],"review_version":1}