{"id":"fd6efbfc-2559-42ff-b38c-23ab1b733ec6","arxiv_id":"2605.26010","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Survey of European secondary students identifies a confidence-competence gap in digital literacy and AI skills, with Dunning-Kruger effects, conditional gender differences, and an AI paradox.","lead":"A survey of 243 European secondary students found they greatly overestimate their digital skills in passive use and AI awareness while showing much lower actual ability in creation and algorithmic tasks. This points to a need for hands-on tech education rather than passive instruction.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Objective competence measures lack any reported validation, item details, or scoring criteria","rationale":"The reader's weakest_assumption correctly isolates the missing measurement details as the single point that blocks any verdict. The full-text placeholder does not supply the required instrument or validation information, so the concern stands and the UNVERDICTED rating is appropriate.","tokens_in":1756,"tokens_out":276,"duration_ms":17537,"concrete_test":"Release the complete item list, scoring rubric, and any reliability/validity data for the operational AI skills and technical readiness instruments; recompute the self-vs-actual comparisons and the p=0.046 gender-gap test after dropping any items with item-total correlation <0.3.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claims (Confidence-Competence Divide, AI Paradox, and classroom-specific gender gap) all require a credible gap between self-reported efficacy and independently measured 'actual technical readiness' / 'operational AI skills'. No information is supplied on instrument construction, item wording, response format, reliability statistics, pilot validation, or exclusion rules for the objective tests. Without these, the observed discrepancies could arise from poorly discriminating items, ceiling effects in the objective measures, or construct mismatch rather than a genuine Dunning-Kruger or paradox effect.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript reports a multicenter survey of N=243 European secondary students that claims to demonstrate a 'Confidence-Competence Divide' (Dunning-Kruger effect) between near-maximal self-reported efficacy in passive digital consumption and lower self-efficacy in active creation/algorithmic logic; an 'AI Paradox' in which students overestimate critical awareness of deepfakes and biases relative to operational AI skills; a technological gender gap that reaches significance (p=0.046) only inside Technology-oriented classrooms; and 76.5% student support for replacing passive instruction with hands-on creation.","tokens_in":1874,"tokens_out":493,"duration_ms":31123,"significance":"If the objective competence instruments prove to be valid, reliable, and construct-valid after full methodological disclosure, the work would supply empirical evidence against the 'digital native' assumption and could inform European secondary curricula on AI readiness and stereotype threat. The classroom-specific gender-gap finding, if robust, would be a useful qualification to broader claims about gender and technology.","major_comments":[{"comment":"Abstract and (presumably) Methods section: the headline claims (Confidence-Competence Divide, AI Paradox, classroom-specific gender gap) all rest on a credible distinction between self-reported efficacy and independently measured 'actual technical readiness' / 'operational AI skills'. No information is supplied on instrument construction, item wording, response format, reliability statistics, pilot validation, scoring criteria, or exclusion rules for the objective tests. Without these details the observed discrepancies could be artifacts of ceiling effects, poor discrimination, or construct mismatch rather than genuine effects.","section":"Abstract / Methods"},{"comment":"Results section (gender-gap analysis): the reported p=0.046 is marginal and the manuscript does not state whether it survives correction for multiple comparisons, reports an effect size, or includes sensitivity checks for classroom-type classification or sample weighting. These omissions directly affect the load-bearing claim that the gap 'emerges significantly exclusively within Technology-oriented classrooms'.","section":"Results"}],"minor_comments":[{"comment":"The abstract states 'supported by an overwhelming student demand (76.5%)' without indicating the exact survey item, response scale, or whether this percentage is conditioned on any subgroup.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"Thank you for the opportunity to respond to the referee's report. We value the feedback and will revise the manuscript to address the concerns raised regarding methodological details and statistical reporting. Below we provide point-by-point responses.","responses":[{"response":"We agree that detailed information on the objective instruments is essential for validating the reported effects. The current manuscript provides limited description in the Methods, which is insufficient. In the revision, we will expand the Methods section with comprehensive details on instrument development, including item wording, response scales, reliability (e.g., internal consistency measures), pilot testing procedures, scoring criteria, and any exclusion rules. This will allow readers to assess potential artifacts such as ceiling effects and strengthen the credibility of the confidence-competence distinctions.","revision_made":"yes","referee_comment":"[Abstract / Methods] Abstract and (presumably) Methods section: the headline claims (Confidence-Competence Divide, AI Paradox, classroom-specific gender gap) all rest on a credible distinction between self-reported efficacy and independently measured 'actual technical readiness' / 'operational AI skills'. No information is supplied on instrument construction, item wording, response format, reliability statistics, pilot validation, scoring criteria, or exclusion rules for the objective tests. Without these details the observed discrepancies could be artifacts of ceiling effects, poor discrimination, or construct mismatch rather than genuine effects."},{"response":"We acknowledge that p=0.046 is marginal and that additional statistical details are needed. In the revised manuscript, we will report appropriate effect sizes for the gender difference. We will clarify whether this was a pre-specified comparison and provide both uncorrected and multiplicity-adjusted p-values. We will also add sensitivity analyses for alternative classroom-type classifications and any sample weighting. These changes will improve the robustness assessment of the classroom-specific gender gap finding.","revision_made":"yes","referee_comment":"[Results] Results section (gender-gap analysis): the reported p=0.046 is marginal and the manuscript does not state whether it survives correction for multiple comparisons, reports an effect size, or includes sensitivity checks for classroom-type classification or sample weighting. These omissions directly affect the load-bearing claim that the gap 'emerges significantly exclusively within Technology-oriented classrooms'."}],"tokens_in":1432,"tokens_out":445,"duration_ms":18825,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing here is a survey of 243 European secondary students that finds they rate their passive digital skills high but drop off on active creation and AI logic, with an AI paradox on deepfakes and a gender gap that only shows up in technology classes. The data also shows strong student support for changing how these topics are taught.\n\nWhat stands out as new is the breakdown by classroom type for the gender difference and the specific application to AI readiness in this age group. The numbers on overestimation and the 76.5% wanting reform give concrete points for discussion.\n\nThe paper does a reasonable job of collecting multicenter data and looking at subgroups rather than just overall averages.\n\nThe soft spot is exactly what the stress-test flags: there's no information in the abstract on the objective measures of competence. No details on what the tests looked like, how they were validated, or even basic stats like reliability. That makes the central claims about the divide and the paradox rest on unexamined ground. The p-value is also right at the edge. If the full paper has a solid methods section with validated instruments, that would change things, but based on what's here the evidence for the gap is weaker than the claims suggest.\n\nThis is for researchers and educators focused on digital literacy and AI education in schools. Someone planning curriculum changes might find the demand for reform useful, but the scientific contribution on the Dunning-Kruger application is limited by the measurement issues.\n\nI'd recommend sending it for peer review. The topic matters and the data is new, so referees could help strengthen the methods reporting.","headline":"Survey of 243 European secondary students finds self-reported digital skills exceed measured ones with a classroom-specific gender gap, but the objective tests lack any reported validation or details.","tokens_in":2419,"tokens_out":403,"would_cite":false,"duration_ms":37736,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"European secondary students greatly overestimate their digital literacy and AI readiness, showing a sharp confidence-competence divide.","keywords":["digital literacy","AI readiness","Dunning-Kruger effect","gender gap","secondary students","deepfakes","algorithmic bias","confidence competence divide"],"falsifier":"Conducting a follow-up study with standardized, validated objective assessments of digital creation and AI operation skills on the same population and observing no overestimation compared to self-reports would falsify the central claims.","tokens_in":2717,"feed_emoji":"","tokens_out":600,"duration_ms":42219,"temperature":0.7,"pith_summary":"The paper challenges the idea that young people are naturally proficient with technology simply because they grew up with it. Instead, it finds that students rate their skills very high for everyday passive use of devices but much lower for actively creating technology or grasping algorithmic logic. This pattern holds across the sample of 243 students, with an added twist that gender differences in perceived tech ability show up only in technology-focused classes. Students also rate their ability to spot deepfakes and biases higher than their actual skills in using AI would support. These results matter because they point to a need for different teaching approaches to build genuine competence rather than relying on assumed natural ability.","feed_headline":"Students overestimate digital skills but lag in AI creation","feed_subtitle":"Survey of 243 European teens shows high self-ratings for passive tech use but sharp drops for creation tasks and AI operation, plus overrate","key_machinery":"The Confidence-Competence Divide, measured by contrasting self-perceived digital literacy with actual technical readiness, including the intra-pathway analysis that isolates the gender gap to specific classroom types.","core_discovery":"The central claim is that there is a severe Confidence-Competence Divide characterized by a collective Dunning-Kruger effect, with near-maximum self-efficacy in passive digital consumption declining sharply for active technological creation and algorithmic logic, a context-specific technological gender gap that emerges only in Technology-oriented classrooms, and an AI Paradox where critical awareness of deepfakes and algorithmic biases is overestimated relative to operational AI skills.","pith_inferences":["If the overestimation extends to real-world decisions, students may be more susceptible to misinformation campaigns involving deepfakes.","Implementing hands-on AI creation programs in schools could be tested to see if it reduces the reported divide.","Similar patterns might appear in non-European contexts if the digital native assumption is tested elsewhere."],"forward_implications":["Reforms should move away from passive theoretical instruction toward hands-on active technological creation.","The gender gap in technology skills is not universal but tied to formal STEM environments, suggesting targeted interventions.","Addressing the AI Paradox requires building operational skills to match critical awareness claims.","Student demand for pedagogical change at 76.5 percent indicates broad support for shifting teaching methods."],"fun_headline_variants":["Teens' tech confidence doesn't match AI creation skills","Dunning-Kruger hits students' digital literacy self-ratings","Students overestimate AI awareness relative to operational skills","Tech gender gap surfaces only in technology classes"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The measures used to assess actual technical readiness and operational AI skills are assumed to be objective and distinct from self-perception, though the study provides limited information on their construction and validation.","fun_headline_variants_meta":{"raw":{"variants":["Teens' tech confidence doesn't match AI creation skills","Dunning-Kruger hits students' digital literacy self-ratings","Students overestimate AI awareness relative to operational skills","Tech gender gap surfaces only in technology classes"]},"model":"grok-4.3","cost_usd":0.005614,"raw_usage":{"total_tokens":2612,"prompt_tokens":679,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":56140500,"prompt_tokens_details":{"text_tokens":679,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1873,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":679,"tokens_out":60,"duration_ms":14860,"temperature":1.0,"reasoning_tokens":1873,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T19:30:52.199841+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Conducting a follow-up study with standardized, validated objective assessments of digital creation and AI operation skills on the same population and observing no overestimation compared to self-reports would falsify the central claims.","supporting_citations":[],"review_version":1}