{"id":"f306e7a4-91c6-4676-9fcd-f8c047d65f65","arxiv_id":"2608.05561","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Software engineers report habitual, functional dependence on LLMs and some overreliance, with addiction-like behavior appearing marginal.","lead":"Software engineers who use AI coding tools report relying on them habitually for everyday tasks, but only a minority describe compulsive or addiction-like behavior. The study suggests companies should focus on trust calibration and preserving professional judgment rather than treating heavy LLM use as addiction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'addiction-related behaviors are marginal' claim rests on treating QC2 open-ended answers as more diagnostic than the structured Q9–Q18 addiction items, but QC2 asks about workflow impact, not impaired control; no decision rule is stated.","rationale":"The paper is transparent about its exploratory design, provides a data availability link, and gives a detailed methodology, which are genuine strengths. The load-bearing issue is one of internal consistency rather than external consensus. The strongest claim in Section 4.3 characterizes the entire relationship as habitual and preferential with addiction-related behaviors 'present only marginally,' yet the instrument's own addiction items, constructed from the behavioral addiction literature, show a clear majority of respondents endorsing frequent addiction-related behaviors on nearly every item (Figure 6). To reach the headline conclusion, the authors must treat the open-ended answers as more valid than the structured items, but the open-ended question they rely on (QC2) was not designed to detect impaired control, withdrawal, or negative consequences; it asks how daily tasks would be affected if LLMs were unavailable. The manuscript never states a rule for resolving conflicts between sources, nor does it provide code frequencies or a codebook allowing the reader to check whether the qualitative coding actually overrides the structured data. The proposed participant-level cross-tabulation directly settles this: it tests whether the two instruments converge or diverge. If they diverge, the central claim must be conditioned to 'addiction-related behaviors were marginal in open-ended accounts,' leaving the structured-item frequencies as an unresolved finding. Because the reader already recommended CONDITIONAL, this stress-test does not change the verdict; it strengthens the specific reason for that conditionality.","tokens_in":21840,"tokens_out":6312,"duration_ms":62318,"concrete_test":"Re-analyze the dataset at participant level: for each of the 119 respondents, compute an addiction endorsement score from Q9–Q18 (e.g., count of 'Often'/'Very Often' responses) and independently code their QC1 and QC2 open-ended responses for explicit indicators of impaired control, withdrawal, or negative consequences. Report the contingency table, in particular how many respondents with high endorsement (e.g., at least half of Q9–Q18 answered 'Often' or 'Very Often') also produced any coded addiction indicator in their open-ended answers. If high endorsers overwhelmingly gave workflow-impact answers such as 'I would switch to Google,' then the marginal-addiction conclusion is an artifact of using a workflow question to measure a clinical construct, and the paper must either justify a decision rule or weaken the claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Section 4.3 claim that addiction-related behaviors are 'present only marginally' is decided by an implicit weighting of two discordant data sources. The ten structured items Q9–Q18 were derived directly from the behavioral addiction literature (Section 3.1) and Figure 6 shows 56–71% of respondents answering 'Often' or 'Very Often' on items covering unplanned overuse (Q9), unsuccessful reduction attempts (Q14, Q15), restlessness and irritability when tools are unavailable (Q16, Q17), and negative effects on job performance (Q18). Section 4.2.3 acknowledges that these items 'captured behaviors commonly discussed in the technology addiction literature,' but then discounts them because participants' comments 'rarely indicated loss of control.' The comments come chiefly from QC2, a hypothetical question about a day without LLM access, which was designed as a contextual probe of how daily tasks would be affected, not as a diagnostic elicitation of craving, withdrawal, or impaired control. It is therefore unsurprising that many QC2 answers mention switching to search engines or slower task completion. Sections 3.5 and 3.8 say the structured and open-ended data are combined, but neither states why the open-ended source should override the structured items, nor reports the frequency of addiction-related codes. Absent that stated rule, the headline conclusion is an interpretive choice rather than a finding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an exploratory survey of 119 software practitioners recruited through Prolific and mailing lists, examining behavioral patterns associated with dependence, overreliance, and addiction-related behaviors in professional LLM use. The instrument combines structured Likert items (Q1–Q18) and two open-ended questions (QC1, QC2), analyzed with descriptive statistics and reflexive thematic analysis. The authors report that LLM use is habitual and goal-oriented: dependence manifests as functional integration into routine work, overreliance as a preference for LLMs over documentation and peers, and addiction-related behaviors are described as marginal. The central claim, stated in Section 4.3, is that engineers' relationship with LLMs is 'habitual and preferential rather than compulsive.'","tokens_in":22096,"tokens_out":5542,"duration_ms":49087,"significance":"If the central claim is supported, the study makes a useful contribution by separating functional dependence from problematic technology use in a professional software engineering population, extending prior work that focused on students or general users. The paper includes several strengths: an internationally diverse sample, explicit data-quality screening, a positionality statement, and a discussion of trust calibration as a complement to verification. However, the headline conclusion that addiction-related behaviors are marginal is not directly supported by the structured data. Figure 6 shows majorities reporting 'Often' or 'Very Often' on ten addiction-related items, and the authors' countervailing evidence consists mainly of responses to QC2, a hypothetical question about a day without LLM access that was not designed to elicit impaired control or withdrawal. Because the weighting of the two discordant data sources is not stated or justified, the significance of the paper's main finding is conditional on a revision that either provides a defensible decision rule or qualifies the claim.","major_comments":[{"comment":"The claim that addiction-related behaviors are 'present only marginally' is not warranted by the structured responses. Figure 6 shows that 56–71% of respondents answered 'Often' or 'Very Often' on Q9, Q14, Q15, Q16, Q17, and Q18, which cover unplanned overuse, unsuccessful reduction attempts, restlessness, irritability, and negative effects on job performance. Section 4.2.3 acknowledges these items 'captured behaviors commonly discussed in the technology addiction literature' but discounts them because open-ended comments 'rarely indicated loss of control.' Neither Section 3.5 nor Section 3.8 states a decision rule for weighting structured versus open-ended data, nor is the frequency of addiction-related codes in the open-ended data reported. Without such a rule, the marginal-status conclusion is an interpretive choice rather than a finding.","section":"§4.3, §4.2.3, Figure 6"},{"comment":"The use of QC2 as the main counter-evidence is problematic. QC2 asks participants to imagine a day without LLM access and describe how their daily tasks would be affected; it was designed as a contextual workflow probe, not as a diagnostic elicitation of craving, withdrawal, or impaired control. The fact that participants mention switching to search engines or slower completion therefore does not indicate the absence of loss-of-control experiences. The paper should either report codes from open-ended questions that directly probe self-regulation and control, or explicitly acknowledge that QC2 cannot speak to the addiction construct.","section":"§4.2.3, Table 2, §3.5"},{"comment":"The operationalization of the constructs is partially circular. The structured items were derived 'directly from the behavioral characteristics identified in the literature' using the same definitions that later structure the interpretation (Section 2.1, Table 1), and Section 3.2 explicitly disclaims formal psychometric validation. Consequently, the mapping from item responses to constructs is uncertain, and the pattern in Figure 6 is partly by construction. The open-ended responses are the only independent grounding, but they are used selectively to override the structured items. The paper should either treat both sources as complementary and report convergence and divergence explicitly, or clearly label the structured items as non-validated exploratory indicators and adjust the strength of the claims accordingly.","section":"§3.1, §3.2, §2.1"}],"minor_comments":[{"comment":"The text states that in Q7 'roughly half' reported 'Often' or 'Very Often' and that in Q8 a 'clear majority' did so, but Figure 4 shows 65% for Q7 and 53% for Q8; the descriptions should be corrected to match the figure.","section":"§4.2.2, Figure 4"},{"comment":"'Forty five participants' should be hyphenated as 'Forty-five.'","section":"§4.2.1"},{"comment":"Figure 7 is referenced in the text but is not present in the provided manuscript; the experience-level results for addiction-related items cannot be verified without the figure.","section":"§4.2.3, Figure 7"},{"comment":"Figure 1 (Thematic Analysis) is referenced in Section 3.5 but is not included in the provided text; the four-stage process cannot be inspected.","section":"§3.5, Figure 1"},{"comment":"In the sentence describing emotional reactions, the quotation marks around 'tedious,' 'annoying,' and 'draining' are inconsistently rendered; please use matching quotation marks.","section":"§4.2.3"},{"comment":"The description of the licensed psychologist's review would benefit from a sentence summarizing the specific feedback that led to questionnaire changes, rather than only stating that feedback was incorporated.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for the journal and addresses a timely, understudied topic in empirical software engineering. The main barrier to publication is the unstated and seemingly arbitrary weighting of discordant data sources in the addiction-related interpretation. This is fixable: the authors can either adopt an explicit, pre-specified decision rule for combining structured and open-ended evidence, or soften the 'marginal' claim to reflect the mixed evidence. I would not reject on these grounds, but the revision must address the internal inconsistency head-on."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is a readable, honest exploratory survey of 119 software engineers about their LLM use, and it fills a real gap, since most work on LLM addiction and overreliance looks at students or general users. The authors sensibly distinguish functional dependence, overreliance, and addiction-related behavior, and their central observation—that engineers mostly treat LLMs as a productivity tool, dislike working without them, often prefer them over docs or peers, but still verify generated code—is plausible and worth having on record.\n\nThe soft spot is the addiction conclusion. The ten structured items (Q9–Q18) show a majority ticking 'Often' or 'Very Often' on behaviors like spending more time than intended, failed attempts to cut back, restlessness and irritability when the tool is unavailable, and negative effects on job performance. Yet Section 4.3 concludes that addiction-related behavior is 'marginal.' That conclusion leans on open-ended comments, mostly from QC2, which asks what would happen if LLMs were unavailable. But QC2 elicits workflow impact—'I'd switch to Google'—not impaired control, craving, or withdrawal. So the two data sources are tapping different things, and the paper never explains why the open-ended responses should override the structured items, nor does it report the frequency of addiction-related codes in the qualitative data. Without a stated decision rule, 'marginal' is an interpretive choice, not a finding.\n\nTo be fair, the authors explicitly say the survey is not a psychometric instrument, and they are appropriately cautious about prevalence claims elsewhere. The problem is that they still use the frequency distributions to support a prevalence-like claim, which is where the tension bites.\n\nMy recommendation: send it to serious peer review. The population is understudied and the data are worth engaging with. A careful referee should ask for a clearer integration of the two data sources, reporting of qualitative code frequencies, and a more calibrated phrasing of the 'marginal' claim—maybe 'less prominent' rather than 'marginal.' The contribution is real even if the headline needs softening.","headline":"Useful exploratory data on an understudied population, but the headline claim that addiction-related behavior is 'marginal' rests on an unstated weighting of discordant structured and open-ended responses.","tokens_in":22594,"tokens_out":1744,"would_cite":true,"duration_ms":18409,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Survey of 119 developers finds LLM use is habitual and functional, not compulsive.","keywords":["large language models","software engineering","dependence","overreliance","addiction-related behaviors","trust calibration","qualitative survey","software practitioners"],"falsifier":"Open the paper's Figure 6 and count the structured addiction responses at face value: for most items Q9–Q18, over half of participants selected 'Often' or 'Very Often,' which contradicts the claim that addiction-related behavior is marginal if those items are treated as the primary evidence. A field test would give developers a diary or telemetry study over several weeks to see whether self-reported loss of control—failed attempts to cut back, restlessness when blocked—matches actual usage and work disruption.","tokens_in":21644,"feed_emoji":"🧑💻","tokens_out":6320,"duration_ms":51781,"temperature":0.7,"pith_summary":"This paper claims that software engineers' relationship with large language models is best described as habitual and preferential, not compulsive. Based on a survey of 119 practitioners, the authors argue that dependence and overreliance are the dominant patterns: LLMs are integrated into routine work for efficiency, and they increasingly displace documentation and peers as the first source of technical support, even while developers still verify generated outputs. Addiction-related behaviors—difficulty moderating use, emotional attachment, restlessness when tools are unavailable—appear only marginally and rarely indicate loss of control or disruption to professional work. The authors conclude that policy and organizational support should focus on trust calibration and preserving professional judgment rather than treating frequent LLM use as an addiction problem.","feed_headline":"LLM use is habitual, not addictive, for most developers","feed_subtitle":"Survey: developers rely on LLMs functionally; addiction-like behavior is rare, so focus on trust calibration.","key_machinery":"The machinery is a conceptual distinction among three constructs—dependence, overreliance, and addiction-related behaviors—operationalized through a qualitative survey. Frequency items (Q1–Q18) measure each construct, while open-ended questions ask participants to describe a concrete LLM use and to imagine a workday without these tools. Descriptive statistics summarize the structured responses and reflexive thematic analysis codes the open-ended ones; the interpretive pivot is that participants' narrative accounts of life without LLMs carry more weight than the frequency scales in deciding whether use is functional or compulsive.","core_discovery":"The central claim is that among practicing software engineers, LLM use is habitual and preferential rather than compulsive. In the paper's account, functional dependence shows up as reduced efficiency when LLMs are unavailable—slower tasks, more manual searching—rather than an inability to work. Overreliance shows up not as blind acceptance of generated code (most respondents reject deploying LLM code without review) but as LLMs becoming the default first stop for information, pushing colleagues and documentation into a fallback role. Addiction-related behaviors are the least supported by participants' own descriptions: structured items captured frequent time overruns and difficulty reducing use, but open-ended accounts rarely described impaired control or significant work disruption, so the authors characterize those behaviors as marginal.","pith_inferences":["Editorial inference: if the structured addiction items are weighted as heavily as the open-ended comments, the paper's 'marginal addiction' conclusion would likely reverse, since most Q9–Q18 items show majorities reporting frequent behavior.","Editorial inference: the data suggest a practical metric for overreliance—capturing whether an engineer's first move is an LLM rather than a colleague or documentation—could be more diagnostic than the standard 'do you verify outputs' question.","Editorial inference: a longitudinal study would test whether overreliance grows as LLMs become more reliable, since developers may verify less when errors become rarer, a trend the cross-sectional snapshot cannot capture."],"forward_implications":["Organizations should treat frequent LLM use as functional dependence rather than a sign of addiction, and support developers in maintaining skills for independent work.","Policy and training should target trust calibration: helping engineers decide when LLMs are the right source and when documentation, prior experience, or colleagues should lead.","Because overreliance appears as LLMs displacing peers and documentation as the first source of support, teams should actively preserve peer consultation and documentation habits.","Support should be tailored to career stage, since early- and mid-career developers report more LLM-centered information seeking than senior developers.","Intensive professional use should not be labeled addiction without evidence of impaired control or disruption to professional work."],"supporting_citations":[{"why":"Supplies the construct of functional dependence as routine integration of a technology because of perceived usefulness; the paper's dependence interpretation rests on it.","marker":"Fan et al., 2017"},{"why":"Defines overreliance as a misalignment between user trust and system capability; the core definition used for the overreliance items and discussion.","marker":"Passi and Vorvoreanu, 2022"},{"why":"Provides the behavioral-addiction criteria of impaired control and persistence despite negative consequences that separate addiction from dependence.","marker":"Chamberlain et al., 2016"},{"why":"Supplies the salience/tolerance/withdrawal framework used to write the addiction-related behavior items Q9–Q18.","marker":"Marchica et al., 2022"},{"why":"Justifies the qualitative survey approach, treating structured responses as interpretable evidence rather than prevalence estimates.","marker":"Braun et al., 2021"},{"why":"Guides the qualitative survey methodology and analysis, including the decision to integrate structured and open-ended responses.","marker":"Melegati et al., 2024"},{"why":"Argues against labeling intensive conversational-AI use as addiction; supports the paper's reading of reported behaviors as non-compulsive.","marker":"Ciudad-Fernández et al., 2025"},{"why":"Documents over-reliance and cognitive effects in educational AI use, the comparison literature the paper extends to practicing engineers.","marker":"Zhai et al., 2024"}],"fun_headline_variants":[],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The strongest conclusion—that addiction-related behaviors are marginal—depends on treating the open-ended comments as more diagnostic than the structured survey answers, where many addiction items drew 'Often' or 'Very Often' from majorities; the paper never states or justifies that weighting rule.","fun_headline_variants_meta":{"error":"'choices'"},"cache_creation_input_tokens":0},"created_at":"2026-08-08T10:27:57.790178+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Open the paper's Figure 6 and count the structured addiction responses at face value: for most items Q9–Q18, over half of participants selected 'Often' or 'Very Often,' which contradicts the claim that addiction-related behavior is marginal if those items are treated as the primary evidence. A field test would give developers a diary or telemetry study over several weeks to see whether self-reported loss of control—failed attempts to cut back, restlessness when blocked—matches actual usage and work disruption.","supporting_citations":[{"cited_title":"Theonline survey as a qualitative research tool","cited_arxiv_id":null,"evidence_quote":"Justifies the qualitative survey approach, treating structured responses as interpretable evidence rather than prevalence estimates."}],"review_version":1}