{"id":"8c028364-9d31-4a02-a38c-bfd6c5df1a21","arxiv_id":"2507.21090","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Interactive simulations improved knowledge transfer to new AI scenarios more than no training, but were not significantly better than static PDFs on most measures.","lead":"This paper tested whether interactive online simulations, called Explorables, teach critical AI literacy better than static text or no training. In a controlled study of 605 adults, the interactive format improved knowledge transfer to new scenarios more than no training, but was not significantly better than static reading on most measures.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pre/post and transfer items are not calibrated; Basic Control's large Fairness gain (0.39 to 0.75, Table 4) shows the reported learning effects are confounded with item difficulty, undermining the central effectiveness claim.","rationale":"The reader's weakest assumption identifies the uncalibrated AI scenario questions, and my stress-test converges on the same point. The load-bearing issue is that the outcome measure itself may not measure learning: because pre- and post-test item sets differ per topic and are not equated, a pre-to-post improvement could simply reflect that the post items are easier. The Basic Control condition provides a direct existence proof of this confound: on Fairness, participants who received no instruction improved by 0.36 (p<.001), so the post-test set is substantially easier than the pre-test set. This means the Explorable condition's Fairness gain of 0.47 to 0.77 cannot be confidently attributed to the interactive tutorial. The same concern applies to the transfer analysis, where 'no significant difference' between target and non-target scores is interpreted as evidence of generalization without calibration or power analysis. The paper otherwise has merits: it is preregistered, uses random assignment, has a reasonable sample size, and includes attention checks and engagement logging. But the abstract overstates the evidence: the main between-condition regression found no significant differences, and the LLM topic declined across all conditions. These issues are addressable, so the conditional verdict is appropriate. My recommendation is unchanged: the paper should be accepted only if the authors can demonstrate item equivalence or re-analyze the data with calibrated measures, and the abstract should be softened to reflect the absence of significant between-condition differences and the possibility of item effects.","tokens_in":14548,"tokens_out":4308,"duration_ms":51375,"concrete_test":"Fit a two-parameter IRT model to all scenario-item responses pooled across conditions and topics; estimate item difficulty for each pre-test and post-test item. If post-test items are significantly easier than pre-test items on Fairness (or any topic), the reported learning gains are confounded with item difficulty. This test uses the existing response data and requires no new data collection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that interactive simulations 'effectively enhance AI literacy across topics' rests on within-subject pre-to-post gains and target-vs-non-target transfer comparisons. However, pre-test and post-test use different, non-overlapping question sets per topic (Table 3, Appendix D): Fairness pre uses Q2, post uses Q1; Diversity pre uses Q1, post uses Q2; LLM pre uses Q2, post uses Q1; Worldview pre uses Q1, post uses Q2. No item calibration, equating, or IRT analysis is reported, and no items appear in both pre and post, so a pre-to-post gain cannot be separated from item difficulty. The Basic Control group, which receives no instruction, improved substantially on Fairness overall (0.39 to 0.75, p<.001, Table 4), directly demonstrating that the post-test Fairness set is easier than the pre-test set. Since the Explorable Fairness gain is 0.47 to 0.77, a large part of the effect could be item difficulty rather than learning. Similarly, the generalization analysis treats 'no significant difference between target and non-target' as evidence of transfer (p>0.05), but with n≈50 this is also consistent with low power, and non-target items are uncalibrated across topics. The Limitations section (5.4) does not mention item equivalence, so the missing support is unacknowledged.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This preregistered online experiment (n≈605) compares interactive 'Explorable' tutorials, static PDF tutorials, and a no-tutorial Basic Control across four AI topics (Worldview, Diversity, Fairness, LLM). Learning is measured with scenario-based multiple-choice questions (pre/post, target/non-target) plus self-reported AI literacy items from the MAILS scale. The paper claims that interactive simulations effectively enhance AI literacy across topics, support knowledge transfer to new scenarios, and increase self-reported confidence, while engagement quantity alone does not predict learning.","tokens_in":14770,"tokens_out":4428,"duration_ms":49284,"significance":"If substantiated, the paper would provide rare controlled evidence on whether interactive, inquiry-driven tutorials improve critical AI literacy beyond static materials, which would be a useful contribution to AI education. Strengths include the preregistered design, random assignment, multiple AI topics, use of real-world scenarios from the AI Incident Database, logging of user interactions, and transparent reporting of null between-condition regressions in the results. However, the central effectiveness and transfer claims are not currently supported by the primary analyses: the within-subject gains rest on non-equivalent pre/post item sets, the Basic Control also shows large gains, and the transfer analysis treats non-significance as evidence of transfer. The abstract and conclusions therefore overstate the findings.","major_comments":[{"comment":"The within-subject pre-to-post gains are not interpretable as learning because no scenario item appears in both pre-test and post-test within a topic; for example, in the Fairness group the pre-test uses Question Set 2 and the post-test uses Question Set 1 (Table 3). The Basic Control group, which received no instruction, improved on Fairness overall from 0.39 to 0.75 (p<.001) (Table 4), demonstrating that the post-test items are easier or that substantial practice effects are present. Since the between-condition OLS regression found no significant condition effects (Static β=-0.01, p=.642; Basic β=-0.04, p=.127), the abstract's claim that 'interactive simulations effectively enhance AI literacy across topics' is not supported by the primary analysis; item calibration or equating is needed before these gains can be attributed to the intervention.","section":"§4.1, Tables 3 and 4"},{"comment":"The generalization analysis uses non-significance (p>0.05) as evidence of transfer, stating that 'effective generalization is indicated by statistically similar performance across target and non-target questions.' With roughly 50 participants per condition, these comparisons have low power, and no equivalence bounds, effect sizes, or item-difficulty controls are reported. The non-target questions are drawn from different topics and are uncalibrated, so equal mean scores may reflect item difficulty rather than conceptual transfer; therefore the claim that Explorables 'support greater knowledge transfer' is not established.","section":"§4.2, Table 1"},{"comment":"The LLM topic showed significant declines in the Explorable condition (overall 0.76 to 0.66, p=.01) and the Static condition (0.79 to 0.68, p=.01), with a non-significant decline in Basic Control (0.69 to 0.63, p=.11). This contradicts the abstract's 'across topics' generalization and should be acknowledged as a topic-dependent boundary condition, with the abstract and conclusion claims adjusted accordingly.","section":"§4.1, Table 4"},{"comment":"The Limitations section does not mention the non-equivalence of pre- and post-test items or the Basic Control's large gains on Fairness, even though these directly threaten the central effectiveness claim. The omitted limitation should be added, and the conclusions should be revised to reflect the null between-condition regression and the measurement confound.","section":"§5.4 Limitations"}],"minor_comments":[{"comment":"The abstract states '605 participants' while §3 reports n=612 recruited, and the final valid samples per condition range from 47 to 51; please reconcile these numbers and provide a participant flow diagram.","section":"Abstract and §3"},{"comment":"Table 1 presents significance with asterisks but no numeric p-values or test statistics for most cells, and the caption does not fully define all abbreviations; consider adding a supplementary table with means, standard deviations, and effect sizes.","section":"Table 1"},{"comment":"The phrase 'statistically similar performance (p>0.05)' should be replaced by proper equivalence-testing language or an explicit power analysis, since a non-significant difference is not evidence of similarity.","section":"§4.2"},{"comment":"The text refers to 'Incident 375' for the Amazon recruiting scenario while footnote 5 links to incident 37; please verify and correct the incident identifiers.","section":"§3, Topic Selection"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable empirical study with a preregistered protocol and a rich dataset, but the abstract and conclusions materially overstate the findings. The primary between-condition regression is null, and the within-subject gains are confounded with item difficulty and practice effects, as shown by the Basic Control's large Fairness gain. The authors should be asked to either reanalyze with item-level calibrations or reframe the claims to align with the evidence; the current presentation is not acceptable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read it. First, the headline claim is not supported by the paper's own between-condition analysis: the interactive Explorables did not significantly outperform static text or even no instruction on the objective post-test, once pre-test scores were in the regression. Second, the within-subject gains the abstract leans on are confounded by uncalibrated pre/post items: the Basic Control group, which got no instruction, improved on Fairness from 0.39 to 0.75 (p<.001), strongly suggesting the post items were easier. The stress-test note is right about that.\n\nWhat's genuinely new: this is one of the first preregistered, randomized controlled comparisons of interactive 'Explorable' tutorials against static text and a no-training control for AI ethics literacy, using real Google PAIR materials across four topics, with transfer and engagement measures. That is a sensible design and a real contribution to the AI-education literature, even if the mechanism (scientific discovery learning) is not new.\n\nWhat it does well: preregistration, random assignment, reasonable sample (600+), multiple topics, transfer questions, and engagement logging. The paper also reports the null between-condition regression honestly in the body, and the discussion concedes 'we did not observe statistically significant differences between conditions.' The tables are detailed enough for a careful reader to see what happened.\n\nSoft spots, in order of size. The biggest is the outcome measure. Pre-test and post-test use different, non-overlapping scenario items per topic; no equating, calibration, or IRT is reported. The Basic Control's Fairness gain is a direct red flag: the post set is easier. That alone undercuts any claim that the gains are learning rather than item difficulty. Second, the transfer analysis treats 'no significant target/non-target difference' as evidence of generalization, which is too generous with n≈50 per cell; the regression shows Explorable better than Basic but not better than Static, so the interactive advantage is not clean. Third, the LLM topic declined in every condition, including a significant decline in the Explorable arm, which contradicts the 'across topics' phrasing in the abstract. Minor: no multiple-comparison correction, no effect sizes or CIs, and data/code are not shared. The limitations section acknowledges short-term gains and sample issues but never mentions item equivalence.\n\nWho's this for? AI-literacy and learning-science researchers, and people building tutorial tools. The value is in the design and the comparative data, not in the current conclusions.\n\nMy take: the paper deserves a serious referee — the question is timely and the design is mostly careful — but it needs substantial revision: calibrate or validate the items, re-analyze with proper metrics, and revise the abstract and conclusions to the defensible claim: Explorables are at least as effective as static text and better than no training on some transfer measures, with no robust evidence of a general interactive advantage. As is, the abstract overstates.","headline":"A mostly well-designed preregistered study whose central claim doesn't survive the uncalibrated pre/post items; the real finding is that interactive Explorables are at least as good as static, not that they're better.","tokens_in":15310,"tokens_out":2792,"would_cite":false,"duration_ms":31462,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Interactive simulations can teach critical AI literacy and transfer to new AI scenarios.","keywords":["critical AI literacy","interactive simulation","Explorable explanations","scientific discovery learning","knowledge transfer","algorithmic fairness","AI education","human-computer interaction"],"falsifier":"A replication that counterbalances the scenario items across pre- and post-tests and equates item difficulty with item response theory would falsify the transfer claim if the interactive condition no longer outperformed the no-instruction control on non-target questions.","tokens_in":14318,"feed_emoji":"🧪","tokens_out":6465,"duration_ms":61493,"temperature":0.7,"pith_summary":"This paper asks whether interactive simulations can teach people to think critically about AI systems. In a controlled study with over 600 participants, it compares interactive 'Explorable' tutorials against static PDFs and a no-instruction control across four AI topics. The authors report that the interactive condition produced significant learning gains, stronger transfer to unfamiliar AI scenarios, and more consistent self-reported confidence gains. They conclude that interactive, inquiry-driven formats are an effective way to build critical AI literacy, while cautioning that raw engagement metrics do not predict learning.","feed_headline":"Interactive sims build AI literacy that transfers","feed_subtitle":"A 605-person study finds hands-on AI tutorials beat static materials on learning and transfer.","key_machinery":"The central object is the 'Explorable': a web-based interactive article that combines explanatory text with adjustable parameters, dynamic visualizations, and real-time feedback, letting learners manipulate inputs and observe AI outputs. The paper builds on scientific discovery learning (SDL), the idea that learners acquire deeper understanding by actively testing hypotheses and observing results, which the authors argue supports the observed gains in critical AI literacy and transfer.","core_discovery":"The paper's central claim is that interactive simulations—Explorable explanations that let users adjust parameters, test hypotheses, and observe AI behavior in real time—enhance critical AI literacy. The evidence comes from a preregistered controlled study in which participants who used the Explorables improved on scenario-based assessments of AI issues, showed comparable or better performance on non-target transfer questions than controls, and reported increased confidence in their AI literacy. The authors attribute this to scientific discovery learning: engagement through experimentation and direct observation fosters conceptual understanding and transfer. They also find that the amount of interaction does not predict learning, suggesting that the quality of engagement matters more than its quantity.","pith_inferences":["A natural extension the authors do not pursue is to use the same Explorables as a reusable critical-literacy curriculum, swapping in topic-specific cases rather than building new materials each time.","A testable follow-up is to capture richer process data—such as screen recordings or verbal protocols—to identify which interaction moments actually drive transfer, since scroll counts are weak predictors.","The interaction-quality finding suggests that adaptive systems could intervene when a learner is merely clicking without experimenting, potentially boosting learning gains beyond those reported here."],"forward_implications":["AI literacy instruction should incorporate interactive, inquiry-driven materials, since interactive engagement produced consistent learning gains across topics.","Passive exposure (reading or no instruction) may not be enough to build transferable critical thinking about AI, as the no-instruction control showed weaker generalization to new scenarios.","Educational tool designs should prioritize meaningful interaction over raw activity, because interaction counts were not reliable predictors of learning.","Topics like large language models may require additional scaffolding when taught interactively, since performance on that topic declined."],"supporting_citations":[{"why":"Supplies the scientific discovery learning framework that motivates the interactive design.","marker":"[7]"},{"why":"Provides the ICAP framework used to interpret why interactive engagement should produce deeper learning.","marker":"[6]"},{"why":"Provides the Meta AI Literacy Scale used to measure self-reported AI literacy before and after the intervention.","marker":"[4]"},{"why":"Defines the AI literacy competencies the study aims to teach.","marker":"[20]"},{"why":"The actual Explorables used as the interactive treatment materials.","marker":"[11]"},{"why":"The online recruitment platform that supplied the participant sample.","marker":"[26]"}],"fun_headline_variants":["Interactive sims boost AI literacy and knowledge transfer","Hands-on AI tutorials beat static lessons in 605-person study","Simulation-based learning fosters critical AI literacy","Experimenting with AI leads to better learning and transfer"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The learning and transfer measures rely on scenario questions that were not item-calibrated, and the pre- and post-tests used different question sets, so the observed gains could reflect item difficulty differences or practice effects rather than the intervention.","fun_headline_variants_meta":{"raw":{"variants":["Interactive sims boost AI literacy and knowledge transfer","Hands-on AI tutorials beat static lessons in 605-person study","Simulation-based learning fosters critical AI literacy","Experimenting with AI leads to better learning and transfer"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000226,"raw_usage":{"total_tokens":1396,"prompt_tokens":802,"completion_tokens":594,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":418,"completion_tokens_details":{"reasoning_tokens":531}},"tokens_in":418,"tokens_out":594,"duration_ms":6195,"temperature":1.0,"reasoning_tokens":531,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:45:13.894198+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that counterbalances the scenario items across pre- and post-tests and equates item difficulty with item response theory would falsify the transfer claim if the interactive condition no longer outperformed the no-instruction control on non-target questions.","supporting_citations":[{"cited_title":"Review of educational research68(2), 179–201 (1998)","cited_arxiv_id":null,"evidence_quote":"Supplies the scientific discovery learning framework that motivates the interactive design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Meta AI Literacy Scale used to measure self-reported AI literacy before and after the intervention."},{"cited_title":"In:Proceedingsofthe2020CHIconferenceonhumanfactorsincomputingsystems","cited_arxiv_id":null,"evidence_quote":"Defines the AI literacy competencies the study aims to teach."},{"cited_title":"https://pair.withgoogle.com/explorables/ (2025), accessed: 2025-01-22","cited_arxiv_id":null,"evidence_quote":"The actual Explorables used as the interactive treatment materials."},{"cited_title":"https://www.prolific.com/ (2025), accessed: 2025-01-22","cited_arxiv_id":null,"evidence_quote":"The online recruitment platform that supplied the participant sample."}],"review_version":1}