{"id":"0069a19f-ee7c-4520-a7ab-02624582529e","arxiv_id":"2608.07592","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Young adult AI chatbot users define AI dependence through chronic use, efficiency, and delegation, and report that their abilities atrophy, suggesting existing dependence scales are misaligned with lived experience.","lead":"Researchers asked 290 young adult AI chatbot users how they think about AI dependence. Participants described dependence as chronic use, efficiency-seeking, and delegation of thinking, and said they feel their abilities atrophy, which suggests current measures may miss their experience.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central 'misalignment' claim rests on an unvalidated single self-report item plus modified scales; test with a stricter criterion.","rationale":"The reader's weakest_assumption is exactly the single-item self-report criterion used to define the dependent group. I agree that this is the most load-bearing assumption: the paper's central quantitative contribution is the demonstration that self-reported dependent participants score low on existing dependence measures, and every qualitative interpretation about scale misalignment is anchored to that grouping. The reader also noted social desirability as a secondary concern; I view the unvalidated single-item criterion and the unmodified-scale issue as the sharper technical problem. The paper itself acknowledges relying on self-reports and mentions mitigation strategies, but it does not validate the criterion item. A stricter criterion would test whether the finding is robust or an artifact. Because the reader already issued a CONDITIONAL verdict, my concern does not change the verdict; it reinforces the condition. The paper's qualitative three-factor model and SDT-based implications are valuable and well-supported by quoted testimonies, and those parts do not hinge on the disputed quantitative comparison. However, the headline claim about current measures being misaligned requires the proposed verification before it should be taken as established.","tokens_in":21668,"tokens_out":3126,"duration_ms":34296,"concrete_test":"Re-run the Step 1/Step 2 analysis with the dependent group defined not by the single item alone but by a stricter composite criterion: self-identification AND endorsement of at least one functional-impairment item (e.g., 'I find it difficult to control my use of AI chatbots' or 'Using AI chatbots interferes with my other activities'), or by a validated multi-item self-report dependence scale. Independently, re-administer the original unmodified validated scales to a comparable sample and check whether the low-score pattern for self-identified dependent users persists. If the pattern disappears under the stricter criterion or with unmodified scales, the misalignment claim is a measurement artifact; if it persists, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that existing AI dependence measures misalign with lived experience—depends on splitting participants into dependent (n=74) and non-dependent (n=216) using a single Likert item, 'I consider myself dependent on AI chatbots,' with no validity evidence for that item (Methods, Analysis Step 2). The item is treated as ground truth against which multi-item validated scales are judged. If this single item captures a broad, colloquial sense of 'I rely on it a lot' rather than clinically meaningful dependence, then low scores on existing measures are expected and the 'misalignment' conclusion is an artifact of the grouping criterion. This concern is amplified by the fact that the scales were modified before administration: 'chatbot' was inserted into items, overlapping items were removed, and Likert formats were standardized (Appendix A). Such modifications can change item functioning and make direct comparison with validated scale scores unreliable. The qualitative themes (chronic use, efficiency, delegation) are internally coherent and valuable, but their interpretive weight as evidence of scale misalignment is only as strong as the unvalidated self-classification that defines the dependent group.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a mixed-methods questionnaire study of 290 US-based young adults (ages 18-25) who use AI chatbots at least weekly. Participants answered open-ended questions about their perceptions and experiences of AI dependence and closed-ended items drawn from existing scales. The authors split the sample into self-reported dependent (n=74) and non-dependent (n=216) groups using a single Likert item, and report that dependent participants did not score as dependent on existing AI dependence measures, which they interpret as a misalignment between current measures and lived experience. Qualitative thematic analysis identifies chronic use, efficiency, and delegation as perceived contributing factors, together with experienced ability atrophy and psychological effects such as inadequacy and impostor feelings. The discussion interprets these findings through self-determination theory and draws implications for reconceptualizing AI dependence, measurement, policy, and design.","tokens_in":21970,"tokens_out":5320,"duration_ms":51406,"significance":"If the central misalignment claim were established, the paper would make an important contribution to the current debate about AI dependence: it would show that existing clinical-derived scales miss a form of dependence that young adults themselves recognize, and it would provide a concrete set of user-centered factors (chronic use, efficiency, delegation) to guide new measures. The qualitative taxonomy is internally coherent and grounded in quotes from both self-identified dependent users and bystanders, and the authors are transparent about sample limitations and response bias. The two-step quantitative screening is conservative, and the paper does not overclaim on the basis of session length. However, the quantitative evidence for misalignment is currently much weaker than the qualitative contribution: it depends on an unvalidated single-item grouping and on modified scales whose psychometric properties in this sample are not reported. As a result, the headline claim should be treated as exploratory pending additional validation.","major_comments":[{"comment":"The paper's central claim that existing AI dependence measures are misaligned with users' experiences rests on splitting participants into dependent (n=74) and non-dependent (n=216) using the single Likert item 'I consider myself dependent on AI chatbots,' yet no validity evidence is offered for this item as a measure of dependence. If the item captures a broad colloquial sense of reliance rather than the clinically-derived construct operationalized by the existing scales, low scores on those scales are expected and do not indicate misalignment. Please add a validity argument or sensitivity analyses (e.g., a stricter cut-off, triangulation with behavioral indicators such as use frequency and open-ended self-descriptions) and, at minimum, soften the claim to an exploratory hypothesis about possible mismatch.","section":"Methods, Analysis Step 2; Findings, Quantitative Results"},{"comment":"The scales were modified before administration—'chatbot' was inserted, overlapping items were removed, and Likert formats were standardized—yet the paper interprets item-level responses as evidence about 'current AI dependence measures.' These modifications can change item functioning and break comparability with validated scales. Please report the exact item-level modifications, the reliability of each modified scale in this sample, and total/mean scores (with distributions) rather than selected item medians, so readers can judge whether the misalignment is substantive or an artifact of scale adaptation.","section":"Appendix A; Methods, Questionnaire"},{"comment":"The two-step screening declares significance only when the Mann-Whitney U test has p<0.05 (Bonferroni-corrected) and Cohen's D>0.7, but the results tables list only Spearman correlations and group medians. Without the effect sizes (and ideally confidence intervals) for the tested items, the reader cannot verify that the reported items actually met the stated threshold or assess the magnitude of group differences. Please include these values in the tables or supplementary material.","section":"Methods, Analysis; Tables 2 and 3"}],"minor_comments":[{"comment":"The sentence 'We first the correlations between closed-ended items and participants' self-reported dependence' appears to be missing a verb; it should read 'We first computed the correlations.'","section":"Methods, Analysis Step 1"},{"comment":"Several typos need a proofread pass: 'chabot' (participant B285), 'chabots' (participant D221), 'out mitigation strategy' in the Limitations, and 'final themes was established' in the Methods section.","section":"Throughout"},{"comment":"The heading 'Step 1: Step 2: Median' is unclear; consider splitting the columns into separate labeled headers for correlation, group medians, p-value, and effect size, and clarify the caption.","section":"Tables 2 and 3"},{"comment":"The statement that 'only five dependent participants reported using AI chatbots for companionship-related reasons' is left unexplained; if this refers to a specific survey item, the item should be identified so readers can interpret the claim.","section":"Findings, Use Motivations"}],"recommendation":"major_revision","confidential_remarks":"The qualitative findings are likely to be of broad interest to the HCI and AI-safety communities, and I would be willing to review a revised version. The main risk is that readers will take the quantitative 'misalignment' result at face value; I urge the editor to require the sensitivity analyses and scale-validation details described in the major comments before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Let me give you my read on the AI dependence paper. The genuinely new thing here is the qualitative model: young adults describe dependence in terms of chronic use, efficiency, and delegation, and they report a felt atrophy of abilities that shows up as inadequacy and impostor feelings. That comes straight from open-ended testimonials, and the themes held across self-identified dependent users, bystanders, and everyone else. The quotes are vivid and the SDT framing is used carefully, not just bolted on. This part is a solid contribution and should survive the review process.\n\nThe weak spot is the quantitative misalignment claim. The paper splits the sample into dependent and non-dependent using a single self-report item ('I consider myself dependent on AI chatbots'), with no evidence that this item is anything more than a colloquial self-assessment. Against that grouping, the fact that existing scales don't flag the 'dependent' group is the load-bearing finding. If the item captures ordinary 'I rely on it a lot' rather than pathological dependence, low scale scores are expected and the 'misalignment' conclusion becomes an artifact of the criterion. The stress-test note is right about this.\n\nOn top of that, the scales were modified — 'chatbot' inserted, overlapping items dropped, Likert formats standardized — before being used to make the comparison. That weakens any direct comparison with the validated instruments. The authors are transparent about this in Appendix A, but transparency doesn't fix the interpretive problem. They'd be on firmer ground positioning the quantitative results as exploratory and framing the qualitative model as the main contribution.\n\nThere are a couple of smaller things. The coding process is described without reliability metrics; that's common in qualitative HCI, but a couple of independent coders and an agreement statistic would quiet reviewers. And the legislation comment about session length gets a thoughtful hedge, so I'm less bothered by it than the reader was. The sample is US only and skewed toward educated women, and the limitations section says so plainly.\n\nOverall: the qualitative model is new and worth engaging with. The quantitative claim about scale misalignment is not yet supported as stated. I'd send it to review, because the qualitative contribution merits a serious referee, but I'd ask the authors to either validate the grouping item or explicitly reframe the misalignment as a hypothesis. If I were working in this space, I'd cite the three-factor model.","headline":"A genuinely useful qualitative model of AI dependence, but the headline misalignment claim rests on an unvalidated single self-report item.","tokens_in":22387,"tokens_out":2638,"would_cite":true,"duration_ms":25140,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI-dependent young adults fail today's dependence scales","keywords":["AI chatbot dependence","young adults","lived experience","self-determination theory","qualitative thematic analysis","self-report measures","ability atrophy","chronic use"],"falsifier":"A behavioral study that logs young adults' actual chatbot use (frequency across contexts, proportion of tasks delegated, and skill-use over time) could test whether self-reported dependence matches observed chronic use, efficiency-seeking, and delegation; if self-identified dependent users show no such behavioral pattern, the paper's lived-experience account is contradicted.","tokens_in":21475,"feed_emoji":"🤖","tokens_out":9772,"duration_ms":73875,"temperature":0.7,"pith_summary":"Current measures of AI chatbot dependence, built largely from clinical addiction frameworks, may be missing the way young adults actually experience dependence. This paper surveys 290 U.S. users aged 18–25 and finds that participants who self-identify as dependent on AI chatbots do not score as dependent on existing dependence scales. From open-ended testimonials it identifies three factors young adults use to define dependence: chronic use (habitual and applied across all contexts), efficiency (saving time and effort at the cost of shifting one's role), and delegation (handing learning, decision-making, and emotional processing to the chatbot). Participants also describe perceived ability atrophy that leads to feelings of inadequacy and impostor syndrome. The authors argue that AI dependence should be reconceptualized around these user-reported factors rather than translated substance-use-disorder symptoms.","feed_headline":"AI-dependent young adults fail today's dependence scales","feed_subtitle":"A 290-user survey finds chronic use, efficiency, and delegation define AI dependence—not clinical addiction items.","key_machinery":"The argument is carried by a two-step quantitative screen paired with inductive thematic analysis of open-ended testimonials. Participants who answered 'somewhat agree' or 'strongly agree' to a single self-report item ('I consider myself dependent on AI chatbots') were labeled dependent (n=74) and compared to the rest (n=216); only items with significant Spearman correlations to that item and significant group differences (Mann-Whitney U with Bonferroni correction and Cohen's D > 0.7) were retained. The qualitative coding yields the three-factor model—chronic use, efficiency, delegation—which the quantitative results corroborate (e.g., dependent users report higher frequency of use and stronger motivation to reduce effort). The lens of self-determination theory then connects these factors to unmet psychological needs: autonomy, competence, and relatedness.","core_discovery":"The central claim is that AI chatbot dependence, as lived and observed by young adults, is not what existing scales measure. Self-identified dependent participants, on average, disagreed with items such as 'I rely too much on AI chatbots' and 'I feel uneasy, anxious, or upset when I cannot use AI chatbots,' despite describing habitual and context-blind use. The qualitative analysis identifies chronic use, efficiency, and delegation as the three perceived contributing factors, and ability atrophy with psychological fallout (inadequacy, impostor feelings) as the perceived consequence. Read through self-determination theory, these factors erode autonomy, competence, and relatedness. The paper therefore proposes that future measurement start from users' own experiences rather than from clinical translation.","pith_inferences":["We infer that a longitudinal study tracking actual skill performance (e.g., coding, writing, critical-thinking tests) before and after heavy delegation would test whether the perceived ability atrophy corresponds to objective decline; the paper only has retrospective self-reports.","We infer that the single-item self-report assumption could be tested directly: if a behavioral measure (usage logs across contexts and proportion of delegated tasks) fails to corroborate self-labeled dependence, the misalignment claim would need revision.","We infer that the frequency-not-session-length finding suggests a concrete policy experiment comparing outcome measures (reported atrophy, distress on removal) under frequency-based versus session-length-based warning criteria.","We infer that because the sample is U.S.-based and skewed toward women and white participants, the three-factor model may not generalize to other cultural or demographic contexts without replication."],"forward_implications":["Existing AI dependence scales will misclassify many young adults who feel dependent, so screening tools should add items on chronic and acontextual use, efficiency-seeking, and delegation.","Frequency of turning to AI, rather than single-session length, is the time-based signal separating dependent from non-dependent users; session-length-based warnings may target the wrong behavior.","The three factors can be operationalized into a new dependence measure grounded in users' own definitions instead of clinical addiction criteria.","Policy and design should address validation-focused ('sycophantic') chatbot behavior, since participants report it encourages continued use for emotional support, and should weigh efficiency pressures against dependence risks."],"supporting_citations":[{"why":"Supplies the General Attitudes towards Artificial Intelligence Scale used in the questionnaire; dependent participants' more positive attitudes on this scale are one of the significant group differences.","marker":"Schepman and Rodway 2020"},{"why":"Provides the Problematic ChatGPT Use Scale; the 'fully absorbed' item is among the few significant items, and yet dependent participants still disagreed on average.","marker":"Yu, Chen, and Yang 2024"},{"why":"Provides the Affective Dependence Scale items (needing AI always available, inability to bear loss) that show the self-labeled dependent group disagreeing with clinical-style dependence statements.","marker":"Sirvent-Ruiz et al. 2022"},{"why":"Provides the AI Dependence Scale and AI use motivation items; items like 'I rely too much on AI chatbots' show the misalignment between self-reported dependence and scale scores.","marker":"Huang et al. 2024"},{"why":"Contributes the Cognitive-Affective-Conative framing and generative AI addiction measures (and design-feature list) that the study adapts; also used to source some dependence items.","marker":"Zhou and Zhang 2024"},{"why":"Supplies self-determination theory, the interpretive lens that maps chronic use and delegation to diminished autonomy, atrophy to diminished competence, and socioemotional substitution to diminished relatedness.","marker":"Deci and Ryan 2012"},{"why":"Provides the inductive thematic analysis method used to code the open-ended testimonials into the three-factor model.","marker":"Clarke, Braun, and Hayfield 2015"}],"fun_headline_variants":["AI dependence scales miss what young adults actually feel","Chronic use, efficiency, delegation: real AI dependence signs","Young adults: AI dependence is not anxiety without chatbots","AI use habits reveal dependence that scales overlook","Survey: AI dependence erodes abilities, breeds inadequacy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the single self-report item 'I consider myself dependent on AI chatbots' truly identifies dependent participants; if it does not, the claimed misalignment with existing scales collapses.","fun_headline_variants_meta":{"raw":{"variants":["AI dependence scales miss what young adults actually feel","Chronic use, efficiency, delegation: real AI dependence signs","Young adults: AI dependence is not anxiety without chatbots","AI use habits reveal dependence that scales overlook","Survey: AI dependence erodes abilities, breeds inadequacy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000232,"raw_usage":{"total_tokens":1443,"prompt_tokens":850,"completion_tokens":593,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":518}},"tokens_in":466,"tokens_out":593,"duration_ms":5948,"temperature":1.0,"reasoning_tokens":518,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:30:37.699052+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A behavioral study that logs young adults' actual chatbot use (frequency across contexts, proportion of tasks delegated, and skill-use over time) could test whether self-reported dependence matches observed chronic use, efficiency-seeking, and delegation; if self-identified dependent users show no such behavioral pattern, the paper's lived-experience account is contradicted.","supporting_citations":[],"review_version":1}