{"id":"4a90b7ea-ab1d-4257-a844-315c8c38f79e","arxiv_id":"2505.01085","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"When government AI is framed as efficient, trust rises but perceived control falls; when it is framed as opaque, irreversible, or uncontestable, both trust and perceived control drop sharply.","lead":"A survey experiment with about 1,200 UK adults found that describing AI in government as efficient raises trust but also makes people feel they have lost control. Highlighting risks like opacity, irreversible dependence, or weak appeal sharply lowers both trust and perceived control.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Risk vignettes embed explicit evaluative conclusions, so the trust/control drops may reflect demand effects rather than the three PAT tensions; pretest manipulation checks cross-load and cannot rule this out.","rationale":"I read the paper in good faith as a preregistered factorial survey testing whether awareness of principal-agent tensions lowers institutional trust, perceived control, and support for AI in government. The preregistration, large sample, three policy domains, and consistent direction of effects are genuine strengths. The central causal inference, however, depends on clean treatment operationalizations. The risk vignettes contain explicit evaluative conclusions that closely track the outcome measures, and the pretest manipulation checks cross-load across all three constructs, so the data cannot currently distinguish between the proposed PAT mechanisms and a general demand or negativity effect. This is an internal-validity concern, not a disagreement with the field's consensus. The reader's weakest_assumption identified the same demand-effect concern; I add the cross-loading evidence from Supplementary E as a concrete supporting detail. The concern is load-bearing for all three hypothesis tests and for the 'failure-by-success' interpretation, but it is fixable by re-running with neutralized vignettes or by demonstrating discriminant validity of the manipulation checks. Because the reader already reached CONDITIONAL on essentially this basis, my assessment does not change the verdict. I do not see a basis for rejection: the descriptive findings and the benefit-condition control-loss result remain informative even if the specific tension-specific causal attribution needs further support.","tokens_in":23265,"tokens_out":4330,"duration_ms":51104,"concrete_test":"Re-run the factorial survey with risk vignettes rewritten to state only structural facts without evaluative conclusions—for example, 'the agency has reduced its human review staff and relies on the AI for most decisions' instead of 'dangerously dependent', and 'decisions are reviewed only in a second-stage appeals process' instead of 'rubber stamp' or 'dehumanised bureaucratic apparatus'. Keep the outcome items, design, and reference conditions identical, and pre-register the same hypotheses and thresholds. If the drops in trust and control relative to the human condition shrink to null or fall below the preregistered thresholds, the current estimates are attributable to demand effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal claim is that awareness of the three principal-agent tensions—assessability, dependency, contestability—causes the large drops in trust and perceived control. This requires treatments that vary only the structural feature and do not themselves assert the outcomes. The risk vignettes in Supplementary B.1 do not meet this requirement. The dependency text states that 'the state has become dangerously dependent on AI' and that 'the authority is now virtually forced to trust the AI'; the contestability text says 'the human becomes a mere rubber stamp' and 'What was intended as progress turns into a dehumanised bureaucratic apparatus'; the assessability text asserts that 'suspicions remain—independent verification is simply not possible.' These are not neutral descriptions of delegation structure; they are conclusions about loss of control, untrustworthiness, and institutional failure—close paraphrases of the dependent variables the study measures. If respondents simply echo the text's conclusions, the measured drops reflect message persuasiveness, not citizens' inferences from structural features. The pretest manipulation checks (Supplementary E) do not resolve this. Each risk condition moves its target check, but it also moves the other checks in the same negative direction: for example, the dependency condition raises the transparency check (b=0.83) and lowers the contestability check (b=-0.69), while the assessability condition raises the dependency check (b=0.84). This cross-loading pattern is exactly what a general negativity or demand account predicts. Even the AI Benefits condition raises the dependency check (b=0.81), so the 'purely beneficial' framing is not perceived as low-risk; the comparison underlying the 'even success reduces control' claim is also contaminated by text describing AI making decisions autonomously.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper applies principal-agent theory to governmental AI use, identifying three structural tensions — assessability, dependency, and contestability — and proposes a 'failure-by-success' dynamic in which efficiency gains initially raise trust while eroding perceived control, and later awareness of delegation risks undermines democratic legitimacy. To test this framework, the authors field a pre-registered factorial survey experiment in the UK (N = 1,201; 1,198 after exclusions) across three policy domains (tax, welfare, bail) with five vignette conditions: a human baseline, an AI-benefits condition, and three conditions adding a risk framing (dependency, contestability, or assessability) to the benefits text. The outcomes are institutional trust, perceived loss of control, and AI delegation preference, analyzed with linear and ordinal mixed-effects models. The AI-benefits condition raises trust (b = 0.51) while also raising perceived loss of control (b = 0.36); all three risk conditions sharply reduce trust (b around -0.45 to -0.55), increase loss of control (b around 1.2-1.3), and shift preferences toward less AI (odds ratios around 0.03). A pretest (N = 301) with manipulation checks confirms that the risk texts move the corresponding perceptions, though with substantial cross-loading across the three checks.","tokens_in":23509,"tokens_out":15244,"duration_ms":136196,"significance":"If the causal interpretation is clean, this is a strong contribution to the literature on AI governance and public opinion: the study is pre-registered with a power analysis and falsifiable success criteria; the sample is large and stratified; the mixed-effects models are pre-specified with tight confidence intervals; and code and data are posted on OSF for replication. The consistent pattern across three outcomes and the large, precisely estimated effects make the core empirical pattern credible. The principal-agent framing is a useful conceptual contribution, and the 'failure-by-success' hypothesis is falsifiable and policy-relevant. Two issues are load-bearing for the central claims, however: the risk vignettes embed explicit evaluative conclusions that closely paraphrase the dependent variables, and the paper's own domain-specific tables contradict the asserted robustness across policy domains. Both are fixable but require substantive revision of the causal framing and the robustness claims.","major_comments":[{"comment":"The central causal claim that awareness of the three principal-agent tensions causes the observed drops in trust and perceived control is not uniquely identified, because the risk vignettes directly assert the outcome constructs. The dependency text states that 'the state has become dangerously dependent on AI' and that 'the authority is now virtually forced to trust the AI'; the contestability text states that 'the human becomes a mere rubber stamp' and that 'What was intended as progress turns into a dehumanised bureaucratic apparatus'; the assessability text asserts that 'independent verification is simply not possible.' Because these phrases are close paraphrases of the dependent variables (loss of control, governmental competence, and trustworthiness), the measured effects may reflect message agreement or demand effects rather than citizens' inferences from the structural features of delegation. The pretest manipulation checks (Tables 14-16) do not resolve this: each risk condition moves not only its target check but also the other two checks in the same negative direction (for example, the dependency condition raises the transparency check by 0.83 and lowers the contestability check by -0.69), so the conditions are not clean operationalizations of single tensions. A concrete way to address this would be to re-estimate the effects with vignettes that describe the structural features without evaluative conclusions, or to reframe the claims explicitly as effects of communications about these tensions.","section":"Section 3; Supplementary B.1; Supplementary E (Tables 14-16)"},{"comment":"The introduction and Section 5 state that the effects are 'robust across all tested policy domains,' but the paper's own domain-specific models contradict this claim for the welfare domain. In Table 8, the welfare AI-Dependency coefficient on trust is -0.24 (95% CI [-0.49, 0.01], p = 0.056) and the welfare AI-Contestability coefficient is -0.20 (95% CI [-0.44, 0.05], p = 0.114), both statistically indistinguishable from zero at the conventional level. Similarly, the claimed robustness of the AI-benefits effect on loss of control is not supported in the welfare domain, where Table 9 reports 0.23 (95% CI [-0.01, 0.46], p = 0.056). The cross-domain robustness claims should be qualified to the outcomes and domains where the models actually show consistent effects, or the domain heterogeneity should be explicitly modeled and discussed.","section":"Supplementary D.2 (Tables 8-9); Sections 1 and 5"},{"comment":"The pretest analysis is offered to 'substantiate the processes underlying these results,' and the text asserts that the manipulation-check responses 'precede reductions in institutional trust and increases in perceived loss of control.' However, the reported pretest models (Tables 14-16) estimate only treatment effects on the three perception checks; they do not estimate the relationship between the checks and the outcome variables, and the pretest and the main study are separate samples. The main study, in turn, includes only a domain-recall comprehension check rather than checks of the three tension perceptions. The mediational or temporal 'precede' claim is therefore not supported by the reported analyses; it should be estimated directly (for example, with the main design including tension checks and outcome items in the same sample) or removed.","section":"Section 5 (pretest analysis); Section 4 (manipulation check)"}],"minor_comments":[{"comment":"The design section says respondents are randomly assigned to 'one of four conditions,' but five conditions are then enumerated; the count should read five.","section":"Section 3"},{"comment":"The sentence 'According to Miller Miller (2005, p. 205f.)' contains a duplicated surname and should be corrected.","section":"Section 2.2"},{"comment":"The claim that '45.4% of respondents in the human condition report feeling a lack of control (above the scale midpoint)' uses exactly the value previously reported as the share preferring 'More AI' in the human condition, and no distribution for the dichotomized control-loss variable is reported; given the human-condition mean of 4.14 on the 1-7 scale, roughly 54% would be expected above the midpoint, so the 45.4% figure should be verified.","section":"Section 5"},{"comment":"Tables 10 and 13 are both titled 'Dependent variable is loss of control' but report cumulative-link odds ratios for the ordinal AI-preference outcome; the titles should be corrected to match the models.","section":"Supplementary Tables 10 and 13"},{"comment":"The abstract and conclusion present the 'failure-by-success' dynamic as an empirical finding, but the design is a static comparison of vignette conditions and does not manipulate the temporal sequence of early success followed by risk awareness; the conclusion acknowledges this ('we could not survey future respondents directly'), and the framing should consistently treat the temporal dynamic as an interpretation rather than a tested result.","section":"Abstract; Section 6"},{"comment":"The manuscript contains numerous typos and missing spaces (for example, 'pre-registeredourdesignandhypotheses' in Section 1), and the notation for coefficients is inconsistent (b vs. beta) in Section 5; a careful copyedit is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The three hypothesis tests rest on the same treatment-validity assumption, so Major Comment 1 is the primary risk to the paper's central inference; if the editors share the demand-effect concern, the claims about the specific PAT tensions will need to be reframed or supplemented with robustness evidence. Note also that the 'AI penalty' literature cited as supporting context consists predominantly of the authors' own preprints (Jungherr et al., 2024; Jungherr & Rauchfleisch, 2025a; Rauchfleisch et al., 2025), which is acceptable but should be balanced with independent evidence. The manuscript is currently a working paper on arXiv and would benefit from a careful editorial pass for the internal inconsistencies listed in the minor comments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a genuine, pre-registered factorial survey experiment on public reactions to AI in government, and the headline result is probably right. Telling people AI makes tax, welfare, and bail decisions more efficient raises institutional trust but lowers perceived control; telling them about dependency, opacity, or lack of contestability risks drops both sharply. Second thing: the risk vignettes contain the conclusions they are supposed to elicit, so the mechanism claims—that the three PAT tensions cause the drops—are shakier than the descriptive effects.\n\nWhat is actually new: the benefits-only cell. Most prior work shows people penalize AI in high-stakes settings. This paper shows that even a purely positive framing of AI handling routine government work reduces perceived control while raising trust. That is a clean empirical anchor for the 'failure by success' idea, demonstrated with a large sample and tight confidence intervals. The PAT framing is borrowed rather than invented, but it is applied cleanly and turned into testable hypotheses. The reporting is unusually transparent for a working paper: full model tables, domain-specific analyses, manipulation checks, an ethics appendix, and a stated pre-registration. Credit is earned there.\n\nWhere it is soft, in order of severity. The demand-effect problem is real. The dependency vignette asserts 'the state has become dangerously dependent on AI' and that the authority is 'forced to trust the AI.' The contestability vignette ends with 'What was intended as progress turns into a dehumanised bureaucratic apparatus.' Respondents are being handed near-paraphrases of the outcomes. The pretest cross-loading backs this reading: the dependency condition raises the transparency check (b=0.83), and the assessability condition raises the dependency check (b=0.84). That is exactly what general negativity looks like. It does not sink the descriptive finding that people react badly when risks are spelled out, but it undercuts the attribution to the three distinct PAT mechanisms.\n\n'Robust across all tested policy domains' is an overstatement. In the welfare domain, the dependency and contestability conditions do not significantly reduce trust (p=0.056, p=0.114). The control-loss effects are robust everywhere; the trust claim is not.\n\nThe OSF link is a placeholder ('Link to data and Code'), so the analysis cannot currently be verified independently. Fixable, but it matters for a paper whose credibility rests partly on pre-registration.\n\nThe human baseline vignettes describe a slow, error-prone status quo while the AI version is uniformly rosy, so part of the 'AI raises trust' effect may be valence asymmetry rather than AI specifically. Minor. And the 'failure by success dynamic' is an interpretation imposed on a cross-sectional experiment, which the paper mostly labels as such.\n\nWho it is for: political scientists, public-administration scholars, and AI-governance researchers who care about experimental public opinion. It deserves a serious referee. With softer vignette wording or deflated mechanism claims, corrected robustness language, and posted data, it would be a solid journal article. I would send it out.","headline":"A careful, pre-registered vignette study with a real finding—efficiency-framed AI raises trust but lowers perceived control—but the risk vignettes are so strongly worded that the three PAT-tension mechanism claims don't fully survive scrutiny.","tokens_in":24074,"tokens_out":5664,"would_cite":true,"duration_ms":58189,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Even purely beneficial AI in routine government tasks lowers citizens' perceived control, and highlighting delegation risks—opacity, dependency, and weak contestability—reduces institutional trust and support for AI.","keywords":["principal-agent theory","AI governance","explainable AI","human-AI interaction","factorial survey experiment","algorithmic fairness","institutional trust","perceived control"],"falsifier":"Run a follow-up experiment that keeps the structural content identical but strips the evaluative language from the risk vignettes—for instance, stating neutrally that the AI's criteria are undisclosed, that appeal rates are low, or that former staff have left—and see whether trust and perceived control still move. A second check would observe citizens who actually interact with an AI-run government service and attempt to contest a decision, measuring trust and control before and after the interaction.","tokens_in":23075,"feed_emoji":"🤖","tokens_out":7553,"duration_ms":71065,"temperature":0.7,"pith_summary":"The paper argues that when governments hand tasks to AI, they are delegating authority in the classic principal-agent sense, and that this delegation carries three built-in tensions—assessability, dependency, and contestability—that ordinary citizens sense even when the AI performs well. In a pre-registered factorial vignette experiment run in the United Kingdom (about 1,200 participants; tax, welfare, and bail domains), a purely efficiency-focused description of AI in government raised institutional trust by roughly half a point on a 7-point scale but simultaneously increased perceived loss of control by 0.36 points. When the same benefits were paired with a paragraph highlighting opacity, lock-in, or weak appeal channels, trust fell below human-only levels, perceived control loss rose by more than a point, and between two-thirds and 70 percent of respondents said the state should use less AI. The authors interpret this as a 'failure-by-success' dynamic: early functional gains legitimise AI, but awareness of the structural risks of delegation corrodes the trust and sense of democratic control that legitimacy depends on.","feed_headline":"AI in government boosts trust while eroding perceived control","feed_subtitle":"Efficiency gains raise trust, but highlighting opacity, dependency, or weak appeals collapses trust and support.","key_machinery":"Principal-agent theory, applied to AI as a delegated agent: the state (principal) hands tasks to an AI system (agent) under information asymmetry and incomplete control. The theory supplies three named tensions—assessability, the opacity of the AI's decision logic; dependency, the irreversibility of delegation once human expertise is eroded; and contestability, the absence of effective appeal or sanction. The empirical machinery is a pre-registered factorial vignette survey (UK, N = 1,198 after attention checks; three policy domains crossed with five conditions; three vignettes per respondent), analysed with linear mixed-effects models for trust and control and cumulative-link ordinal models for AI preference. The vignettes are the operative mechanism: every AI condition pairs the same benefit paragraph with a paragraph that makes one tension concrete.","core_discovery":"The central claim is that delegating government authority to AI is a principal-agent problem, and its three structural tensions—can decisions be understood (assessability), can delegation be reversed (dependency), and can decisions be challenged (contestability)—predict citizens' trust, perceived control, and support for AI. The experiment shows that even the most favourable presentation of AI, one emphasising speed, cost savings, and error reduction, leaves citizens feeling less in control ($b = 0.36$; $p < 0.001$). When any of the three tensions is made salient, institutional trust drops by roughly 0.45–0.55 points relative to human administration, perceived loss of control rises by 1.2–1.3 points, and between 65.6% and 70.5% of respondents demand less AI use. These effects appear consistently across tax, welfare, and bail decisions. The paper concludes that awareness of delegation risks, not AI malfunction, drives the erosion of trust and control—so even successful AI implementation can undermine democratic legitimacy.","pith_inferences":["The vignette wording may mean the effects partly measure message compliance rather than reasoning from structure; a valence-controlled replication that varies emotional phrasing while holding the structural facts equal would separate the two.","The same delegation logic should apply outside government: employers, platforms, and hospitals that hand consequential decisions to opaque, hard-to-reverse, hard-to-appeal systems should see analogous drops in trust and a sense of control among affected people.","The contestability result implies a concrete policy lever: if appeal channels are visible, cheap, and staffed by humans before AI is deployed, the control-loss effect may shrink; this follows from the paper's data but is not tested there.","Because the sample is a UK online panel stratified on demographics rather than a probability sample, the effect sizes are likely to differ across political cultures; comparative replications with representative samples would test how far the failure-by-success dynamic generalises."],"forward_implications":["Any government that advertises AI's efficiency should expect an immediate legitimacy trade-off: trust rises modestly while perceived control falls.","Making delegation risks salient flips public demand: across all three risk conditions, between 65.6% and 70.5% of respondents said the state should use less AI, compared with 45.4% who wanted more AI in the human condition.","Early acceptance does not predict long-run acceptance: if awareness of AI's inner workings grows with broader use, the same public that welcomed efficiency gains will later show lower trust and control.","Explainability tools alone are unlikely to neutralise the problem, because the assessability condition—which mirrors the explainable-AI agenda—produced losses of trust and control as large as the dependency and contestability conditions."],"supporting_citations":[{"why":"Supplies the competence-control dilemma that grounds the dependency and contestability predictions.","marker":"(Abbott et al., 2020)"},{"why":"Defines the canonical principal-agent assumptions the paper maps onto AI delegation.","marker":"(Miller, 2005)"},{"why":"Establishes machine-learning opacity that motivates the assessability tension.","marker":"(Burrell, 2016)"},{"why":"Formalises information asymmetry and moral hazard, the PAT core the assessability hypothesis extends.","marker":"(Holmstrom, 1979)"},{"why":"Provides the factorial survey methodology used for the vignette experiment.","marker":"(Auspurg & Hintz, 2015)"},{"why":"Supplies the lme4 mixed-effects estimator used for the trust and control analyses.","marker":"(Bates, Mächler, Bolker, & Walker, 2015)"}],"fun_headline_variants":["AI in government: efficiency boosts trust, but control slips","Government AI: trust rises, but control falls","Efficiency gains, control losses: AI governance paradox","AI success story undermines citizens' control","When AI governs, we lose control even when it works"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The causal inference that the three tensions drive the attitude shifts assumes the vignettes manipulated only the intended structural feature and did not themselves supply the measured attitudes—yet the risk vignettes explicitly assert that the state has become 'dangerously dependent', that the human becomes 'a mere rubber stamp', and that the result is 'a dehumanised bureaucratic apparatus', so the drops in trust and control could reflect what the text urged rather than what citizens inferred.","fun_headline_variants_meta":{"raw":{"variants":["AI in government: efficiency boosts trust, but control slips","Government AI: trust rises, but control falls","Efficiency gains, control losses: AI governance paradox","AI success story undermines citizens' control","When AI governs, we lose control even when it works"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000627,"raw_usage":{"total_tokens":2912,"prompt_tokens":967,"completion_tokens":1945,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":1870}},"tokens_in":583,"tokens_out":1945,"duration_ms":14080,"temperature":1.0,"reasoning_tokens":1870,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:27:04.832504+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a follow-up experiment that keeps the structural content identical but strips the evaluative language from the risk vignettes—for instance, stating neutrally that the AI's criteria are undisclosed, that appeal rates are low, or that former staff have left—and see whether trust and perceived control still move. A second check would observe citizens who actually interact with an AI-run government service and attempt to contest a decision, measuring trust and control before and after the interaction.","supporting_citations":[{"cited_title":"Concentric Spherical GNN for 3D Representation Learning","cited_arxiv_id":"2103.10484","evidence_quote":"Defines the canonical principal-agent assumptions the paper maps onto AI delegation."},{"cited_title":"\\ Hintz, T","cited_arxiv_id":null,"evidence_quote":"Provides the factorial survey methodology used for the vignette experiment."}],"review_version":1}