{"id":"08ac5304-58ad-411d-8b57-ff718dfef467","arxiv_id":"2412.13846","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A UTAUT2-based survey of 181 practitioners finds performance expectancy and habit drive fairness toolkit adoption.","lead":"This paper surveys 181 software practitioners and finds that their belief in a fairness toolkit's usefulness, plus habit, best explain whether they intend to adopt and actually use the toolkit. It applies a standard technology acceptance model to fairness tools in software engineering, yielding practical advice for organizations and toolkit vendors.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The HB→UB path (0.543) may be substantially tautological: the paper never shows the habit item wordings, and the dependent variable is a single-item frequency scale, so 'habit' might just measure past use frequency.","rationale":"The reader's weakest assumption points to the same load-bearing risk, and I agree that it is the most concrete threat to the paper's central claim. If the Habit items are standard UTAUT2 automaticity items ('natural', 'automatic', 'addicted'), the concern would not land; but the paper provides no item wordings and the hypothesis narrative itself conflates habit with consistent/regular use. That conflation, combined with a single-item frequency-based UB measure, makes the HB→UB path potentially inflated. Other issues exist—the 'intention to use large language models (LLMs)' sentence in Section V.B is an obvious copy-paste error, and the cross-sectional design limits causal language—but neither directly undermines the strongest coefficient the way item overlap would. The PE→BI path (0.465) is independent of this problem, so the paper's secondary claim about performance expectancy survives. For that reason I would not move the verdict to REJECT or UNVERDICTED; the concern is testable and addressable, which is consistent with the reader's CONDITIONAL verdict. If the concrete test confirms overlap, the authors should re-report the model with the offending item excluded and soften the habit-related implications; if the items are pure automaticity measures, the concern is resolved.","tokens_in":17708,"tokens_out":4377,"duration_ms":42613,"concrete_test":"Obtain the online appendix or survey instrument and compare the four Habit items with the single Use Behavior item. If any HB item contains wording about frequency, regularity, or amount of use (e.g., 'I use fairness toolkits often', 'I regularly use fairness toolkits') rather than automaticity ('automatic', 'natural', 'without thinking', 'must use'), re-estimate the structural model with that item removed or with UB regressed only on the non-frequency HB items. If the HB→UB coefficient (reported 0.543) drops materially, the headline claim that habit is the primary driver of actual use is not supported and the practical recommendations about habit cultivation need to be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that habit is the strongest driver of actual use behavior (HB→UB = 0.543, Table II). This claim requires that the Habit construct measures an automatic behavioral tendency, not merely the frequency with which the toolkit is already used. The paper's own hypothesis development (Section III.B) defines habit as 'familiar with certain tools and use them regularly' and argues H6b through 'practitioners who consistently integrate new tools are more inclined to use fairness toolkits consistently.' That is dangerously close to a frequency tautology. Meanwhile, Section IV.C.1 reports that UB was 'measured using a single-item frequency scale.' If any of the four HB items asks about regularity, frequency, or amount of use, then HB→UB is partly past-use-predicts-current-use, inflating both the path coefficient and the f² = 0.150 effect size. The actual item wordings are not in the main text; the online appendix reference [53] is a placeholder ('A. Authors') and was not verifiable. This is the exact assumption on which the strongest finding and the practical recommendation to 'cultivate habit' depend.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a survey-based study of 181 software practitioners recruited through Prolific, using the UTAUT2 framework and PLS-SEM to explain behavioral intention (BI) and actual use behavior (UB) regarding fairness toolkits. The model includes performance expectancy, effort expectancy, social influence, hedonic motivation, facilitating conditions, and habit as predictors. The main findings are that performance expectancy has the strongest effect on BI (path 0.465) and habit has the strongest effect on UB (path 0.543), with BI also significantly predicting UB. The authors conclude that practitioners adopt fairness toolkits primarily for performance reasons and that habitual use drives sustained adoption, leading to recommendations for organizations and tool vendors to integrate fairness toolkits into routine workflows.","tokens_in":17896,"tokens_out":4489,"duration_ms":44366,"significance":"If the findings hold, the paper would provide one of the first quantitative, theory-grounded accounts of fairness toolkit adoption, complementing existing qualitative work and offering concrete design and organizational implications. The study has methodological strengths: an a priori power analysis, iterative pilot testing, attention checks, reliability and validity reporting (Cronbach's alpha, rho_c, rho_A, AVE), bootstrapping with 10,000 subsamples, and PLSpredict benchmarking. However, the paper's central empirical claim—that habit is the primary driver of actual use—depends on the discriminant validity between the habit construct and the single-item frequency measure of use behavior, and the evidence for that discriminant validity is not currently verifiable because the item wordings are omitted and the online appendix reference is a placeholder. This gap must be resolved before the practical recommendations about cultivating habit can be accepted.","major_comments":[{"comment":"The strongest path in the model, HB→UB (0.543, Table II), is the central empirical result, but the paper does not report the wording of the four HB items or the single UB frequency item, and the cited online appendix [53] is listed as 'A. Authors' with no accessible content. The hypothesis text defines habit partly as 'use them regularly' (Section III.B) and H6b states that 'practitioners who consistently integrate new tools are more inclined to use fairness toolkits consistently.' If the HB items ask about regularity, frequency, or amount of use, then the HB→UB path is partly tautological in the sense that past use predicts current use. Please report the full item wordings, provide the UB item, and show that the HB items capture automaticity rather than mere frequency, for example by reporting a factor analysis that separates HB from UB and by re-estimating the model without any frequency-based HB indicator.","section":"Section III.B, Section IV.C.1, Table II"},{"comment":"The claim that habit and use behavior are empirically distinct rests entirely on HTMT values that are only mentioned as being in the online appendix; the appendix is not verifiable because reference [53] is a placeholder with no author and no accessible link. Given that UB is a single-item frequency scale and HB items are not shown, the reported HTMT results are insufficient to rule out the tautology concern. Provide the full HTMT matrix (or at least the HB–UB value), and ideally a robustness analysis that treats UB and HB as indicators of a single factor to show they are not empirically indistinguishable.","section":"Section V.A, 'Discriminant Validity'"},{"comment":"The practical recommendations state that organizations should 'help employees develop a habit' and that habitual use can lead to higher adoption. These causal prescriptions go beyond what a cross-sectional, self-report survey can support; the data establish associations only. The paper should either add an explicit limitation acknowledging that the direction of causality is not tested, or soften the causal language throughout the discussion and conclusion. This is directly relevant to the central claim because the practical value of the habit construct depends on whether habit can be deliberately cultivated, which the present design cannot establish.","section":"Section VI.B and Section VIII"}],"minor_comments":[{"comment":"The text reports R² values for Behavioral Intention and Use Behavior but then says 'the model successfully explains 63% of the variance in the intention to use large language models (LLMs).' This appears to be a copy-paste error; it should refer to fairness toolkits.","section":"Section V.B, 'Explanatory Power'"},{"comment":"The reference for the online appendix is incomplete: it lists 'A. Authors' and a figshare URL but no author names, title, or year-specific identifier. Please provide a complete citation and ensure the replication materials are accessible.","section":"Reference [53]"},{"comment":"For HM→BI the reported p-value is 0.94 with a negative path coefficient; this is likely a one-tailed probability for the wrong direction. Clarify whether one-tailed or two-tailed tests are used for structural paths and report the corresponding p-values consistently.","section":"Table II"},{"comment":"The survey title is written as '[\"Fairness in Software Development] Pre-Screening — Fairness Toolkit Adoption' with mismatched quotation marks; please correct the formatting.","section":"Section IV.B"},{"comment":"The internal-validity discussion does not address common method bias, which is a standard concern for single-source, self-report surveys. Consider adding a Harman's single-factor test or a marker-variable analysis, or at least discuss why this is not a concern here.","section":"Section VII, Internal Validity"},{"comment":"The claim that age, gender, and experience moderators were not significant is supported only by a reference to the online appendix. Report the relevant test statistics or model comparison in the main text to make this exclusion auditable.","section":"Section III.A"},{"comment":"The paper states that EE4 and HB2 had outer loadings below 0.70 but provides no loading values. Please include the complete loading table or add the values in the appendix rather than referring only to an inaccessible source.","section":"Section V.A, 'Indicator Reliability'"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely salvageable and could become a useful contribution to empirical software engineering on fairness tool adoption. The decisive issue is the habit/use-behavior tautology risk: the strongest claim in the paper cannot be evaluated without the measurement instruments. The missing online appendix is more than a minor formatting problem because the entire discriminant-validity argument is outsourced to it. I would advise the editor to require the authors to provide the item wordings, the HTMT matrix, and a re-estimation excluding frequency-overlapping HB items before accepting. Also note the sloppy 'LLM' copy-paste in Section V.B suggests the manuscript would benefit from a careful final proofread."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is the first to apply UTAUT2 and PLS-SEM to fairness toolkit adoption, and that new application is the main contribution. The survey work is methodologically careful: a priori power analysis, pilot tests, attention checks, and standard reliability and validity reporting. The finding that performance expectancy drives intention (PE→BI = 0.465) and habit drives actual use (HB→UB = 0.543) is plausible and broadly consistent with UTAUT2 literature. If you work on SE4AI or technology acceptance in software engineering, this is worth reading and engaging with.\n\nThe soft spots, in order of severity. First, the habit–use relationship may be partly tautological. The paper defines habit as 'using regularly' and does not show the habit item wordings in the main text. The dependent variable is a single-item frequency scale. If the habit items ask about past frequency, then the strongest path is partly past-use-predicts-current-use, and the f² = 0.150 overstates the practical lever. This is a load-bearing concern for the central claim, so the authors need to make the full questionnaire available and show that habit items tap automaticity, not frequency. Second, the online appendix reference [53] is a placeholder ('A. Authors') and the URL is not verifiable in the text; that is a concrete reproducibility gap. Third, there are clear copy-paste slips: Section V.B reports that the model explains '63% of the variance in the intention to use large language models (LLMs)' — this is a fairness toolkit paper — and a stray female-symbol glyph appears in the contribution paragraph. These are minor individually, but they suggest the manuscript was not carefully proofread.\n\nI also note that the reader's skepticism about circularity is warranted, but I do not see this as a fatal flaw. The measurement instruments are externally sourced, the statistical execution meets field standards, and the authors are transparent about the cross-sectional self-report design in their threats to validity. The paper is incremental — applying an established model to a new context — but that is still a legitimate contribution, especially given the practical implications for toolkit vendors and organizations.\n\nWho this is for: researchers studying fairness tools adoption, technology acceptance in SE, and practitioners wanting evidence on what drives toolkit use. A serious referee should engage with it, but the referee should require the replication package, the full item wordings, and a discussion of the habit/use overlap. I would recommend sending it to peer review with major revision expectations.","headline":"A competent UTAUT2 application to fairness toolkit adoption whose headline habit finding may be partly tautological, and whose manuscript shows careless copy-paste slips.","tokens_in":18443,"tokens_out":1672,"would_cite":false,"duration_ms":17594,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that software practitioners adopt fairness toolkits chiefly because they expect the tools to improve bias mitigation and because using them has become habitual, with performance expectancy the strongest driver of…","keywords":["fairness toolkits","technology adoption","UTAUT2","PLS-SEM","habit","performance expectancy","machine learning fairness","software practitioners"],"falsifier":"A follow-up measurement study in which habit is assessed with an automaticity scale rather than a frequency scale, checked against the same UTAUT2 paths, would settle the question: if the habit-to-use coefficient drops far below 0.543 or becomes non-significant, the central claim that habit formation drives adoption is not supported.","tokens_in":17509,"feed_emoji":"🧑‍💻","tokens_out":4815,"duration_ms":45449,"temperature":0.7,"pith_summary":"The paper set out to explain why software practitioners actually use fairness toolkits, which are libraries and metrics designed to detect and mitigate bias in machine-learning models. It surveyed 181 industry practitioners and analyzed their responses with a structural equation model based on the Unified Theory of Acceptance and Use of Technology (UTAUT2). The central finding is that two individual factors dominate: performance expectancy, or the belief that the toolkit will help produce fairer software, is the strongest predictor of intention to adopt, while habit is the strongest predictor of actual use behavior. Social influence, effort expectancy, facilitating conditions, and enjoyment were not significant in this sample. If the finding holds, organizations and tool vendors should promote adoption less through mandates or social pressure and more by demonstrating concrete bias-mitigation results and by embedding toolkit use into routine workflows.","feed_headline":"Payoff and habit drive fairness-toolkit adoption","feed_subtitle":"Performance expectancy steers intent; habitual use steers actual adoption, a 181-practitioner study reports.","key_machinery":"The central machinery is the UTAUT2 model, a validated survey-based framework for technology adoption that explains intention and use through constructs such as performance expectancy, effort expectancy, social influence, facilitating conditions, hedonic motivation, and habit. The study drops price value, excludes the standard demographic moderators, and links the remaining latent constructs to self-reported use via partial least squares structural equation modeling, a regression-based path analysis suited to small and non-normal samples. The habit construct, defined as the extent to which behavior becomes automatic through learning, carries the paper's strongest result because its direct path to use behavior dominates all other paths in the model.","core_discovery":"The study reports that performance expectancy and habit are the primary drivers of fairness toolkit adoption. In the estimated model, performance expectancy has the largest path to behavioral intention (coefficient 0.465), habit has the second-largest path to intention (0.323), and habit has by far the largest path to actual use behavior (0.543), with behavioral intention itself contributing a smaller path to use (0.161). The model explains 63 percent of the variance in intention and 40.7 percent of the variance in use. A mediation analysis indicated that habit's effect on use is direct rather than mediated by intention. The authors interpret this as evidence that practitioners decide to adopt fairness toolkits based on perceived effectiveness at mitigating bias, and that sustained usage comes from the tools becoming part of their regular working routines.","pith_inferences":["A natural testable extension is whether habit-forming interventions, such as scheduled fairness checks in regression pipelines, actually increase adoption over time; the paper's cross-sectional design cannot establish that habitual use causes adoption, so longitudinal or experimental studies would sharpen the claim.","The null results for effort expectancy and social influence may not generalize to less experienced practitioners or to organizations where fairness toolkits are mandatory, because the sample skews toward experienced developers and self-reported professional contexts.","The strong habit-to-use path may partly reflect measurement overlap: if the habit items ask mainly about past frequency of use, then part of that coefficient restates that past use predicts current use rather than showing an automatic tendency that managers can deliberately cultivate.","A useful replication would measure habit with an automaticity-focused scale rather than a frequency-based scale; if habit's path to use drops substantially under that measurement, the practical advice to 'build habits' would need to be reframed around workflow redesign."],"forward_implications":["If the paper is right, awareness campaigns for fairness toolkits should emphasize demonstrated bias-mitigation effectiveness and real-world success cases, since performance expectancy is the strongest driver of intention.","Tool vendors and organizations should focus on making toolkit use easy to embed in daily work, for example through well-designed APIs and integration into pipelines, because habit is the strongest predictor of actual use.","The non-significant paths for social influence, effort expectancy, facilitating conditions, and hedonic motivation imply that organizational support, peer pressure, ease of use, and enjoyment are not reliable levers for increasing adoption in this population.","The variance explained suggests that individual perceptions and routines account for a large share of adoption behavior, but that roughly 37 percent of intention and 59 percent of use remain unexplained by the factors studied.","The direct, unmediated habit-to-use path implies that forming a routine is at least as important for real adoption as strengthening a practitioner's conscious intention to use the toolkit."],"supporting_citations":[{"why":"Supplies the UTAUT2 model, its constructs including habit, and the validated questionnaire items the survey adapts.","marker":"[23]"},{"why":"Provides the PLS-SEM estimation and evaluation criteria used to assess the measurement and structural models.","marker":"[24]"},{"why":"Defines the original UTAUT constructs, including performance expectancy, behavioral intention, and use behavior, which the hypotheses extend.","marker":"[50]"},{"why":"Documents how machine-learning practitioners currently try to use fairness toolkits, establishing the adoption gap this study addresses.","marker":"[18]"},{"why":"Defines the PLSpredict procedure and the linear-model benchmark used to evaluate the model's predictive power.","marker":"[69]"},{"why":"Provides the guidelines for interpreting PLSpredict results, which the paper follows to claim strong predictive capability.","marker":"[70]"}],"fun_headline_variants":["Expectation fuels intent, habit fuels use of fairness toolkits","Payoff drives intent, habit drives adoption of fairness tools","Why adopt fairness toolkits? Perceived payoff and routine habit","Habit, not just payoff, is key to fairness-toolkit adoption","Performance expectancy and habit steer fairness-toolkit usage"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the survey's habit items measure an automatic tendency built through learning, distinct from how often practitioners say they use the toolkit; if the items mostly capture past frequency of use, the strong habit-to-use path largely restates that past use predicts current use.","fun_headline_variants_meta":{"raw":{"variants":["Expectation fuels intent, habit fuels use of fairness toolkits","Payoff drives intent, habit drives adoption of fairness tools","Why adopt fairness toolkits? Perceived payoff and routine habit","Habit, not just payoff, is key to fairness-toolkit adoption","Performance expectancy and habit steer fairness-toolkit usage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000252,"raw_usage":{"total_tokens":1556,"prompt_tokens":933,"completion_tokens":623,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":535}},"tokens_in":549,"tokens_out":623,"duration_ms":6153,"temperature":1.0,"reasoning_tokens":535,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:43:40.310764+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A follow-up measurement study in which habit is assessed with an automaticity scale rather than a frequency scale, checked against the same UTAUT2 paths, would settle the question: if the habit-to-use coefficient drops far below 0.543 or becomes non-significant, the central claim that habit formation drives adoption is not supported.","supporting_citations":[{"cited_title":"Consumer acceptance and use of information technology: Extending the unified theory of acceptance and use of technology,","cited_arxiv_id":null,"evidence_quote":"Supplies the UTAUT2 model, its constructs including habit, and the validated questionnaire items the survey adapts."},{"cited_title":"Why don’t men ever stop to ask for directions? gender, social influence, and their role in technology acceptance and usage behavior,","cited_arxiv_id":null,"evidence_quote":"Defines the original UTAUT constructs, including performance expectancy, behavioral intention, and use behavior, which the hypotheses extend."},{"cited_title":"The elephant in the room: Predictive performance of pls models,","cited_arxiv_id":null,"evidence_quote":"Defines the PLSpredict procedure and the linear-model benchmark used to evaluate the model's predictive power."}],"review_version":1}