{"id":"4654875d-21a8-4cf9-b416-e8d04295d938","arxiv_id":"2504.19047","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Concerns about how one is perceived, even without any monetary or reputational stakes, make people reject accurate AI recommendations and lower their performance in an image-classification task.","lead":"An online experiment shows that people use AI recommendations less and perform worse when they think a researcher will review their AI usage, even though the review has no financial consequences. The study identifies concern about being judged as a reason people underuse useful AI advice.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The treatment bundles being observed with being told one's usage will be compared to the average; this confound leaves the non-instrumental image mechanism under-identified, and a no-human-observer control is needed.","rationale":"The reader's weakest assumption correctly targets the identification of the treatment effect as non-instrumental image concerns. My pass sharpens this into a specific bundled-treatment problem: the 'compared to the average' instruction itself transmits information that may change behavior without any social evaluator. This is a load-bearing concern because the paper's central theoretical contribution is the mechanism, not just the reduced-form effect of being observed. The self-reported discomfort and satisfaction measures are suggestive but cannot discriminate between image utility and demand or norm-based channels. The proposed three-arm design would settle whether the social-observation component is necessary. Because the reader already returned CONDITIONAL, and my concern reinforces the need for additional controls rather than overturning the reduced-form result, I recommend no change to the verdict. The concern is not that the experiment is invalid, but that the specific interpretation as non-instrumental image concerns requires an additional control condition to rule out confounds inherent in the current treatment.","tokens_in":13348,"tokens_out":4217,"duration_ms":49366,"concrete_test":"Run a three-arm follow-up: (A) original treatment, (B) 'private audit' in which participants are told their AI usage will be compared to the average and the result shown only to them, with no human observer, and (C) control. If AI use falls in B relative to C, the effect is not driven by social-image utility but by norm information or self-image; if A shows a larger drop than B, the social-observation channel is confirmed. Also elicit post-task beliefs about the researcher's expectations in all arms to quantify demand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that anticipated social evaluation, with no monetary consequence, reduces AI use. The treatment, however, differs from control in two bundled ways: (1) a researcher will see the participant's AI usage, and (2) that usage will be compared to the average participant's behavior (Section 2.2; Appendix A.3). Component (2) can operate without any social-image mechanism: it may convey a descriptive norm, activate a reference point, or signal the researcher's hypothesis about appropriate AI use. The paper's defense in Section 2.5 that 'social pressure' experimenter demand is 'precisely the object of study' conflates approval-seeking from an experimenter with the real-world construct of non-instrumental image concerns among peers. If participants reduce AI use because they infer the researcher wants them to, the result is a demand effect, not evidence about how perceived intelligence or effort shapes AI adoption. The self-reported discomfort measure (Section 3.3) is consistent with image concerns but is also consistent with mere evaluation apprehension, which may be instrumental (e.g., desire to avoid awkwardness) or a general reaction to being monitored. Thus the design does not isolate non-instrumental image utility from information about norms or from experimenter demand.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a pre-registered online experiment (Prolific, N=220) examining whether non-instrumental image concerns reduce the use of AI recommendations. Participants completed 50 rounds of an image classification task with an AI recommendation offered after their initial choice; one round was randomly selected for a performance-based bonus. In the treatment condition, participants were told that a research team member would review their AI usage and compare it with the average participant's behavior during a mandatory end-of-study video call; in the control condition the video call only revealed the bonus outcome. The treatment reduced AI recommendation use by 4.5 percentage points (10%) and final accuracy by 2.7 percentage points (3.3%), with no significant effects on initial choices or response times. The paper interprets these effects as evidence that non-instrumental image concerns cause underutilization of AI, and reports exploratory heterogeneity showing stronger effects for minority participants.","tokens_in":13518,"tokens_out":6584,"duration_ms":62443,"significance":"The paper addresses an important question: why human-AI collaboration underperforms despite accurate AI recommendations. It has notable strengths: the experiment is pre-registered, treatment groups are balanced on observables (Table 1), the main regressions include image fixed effects and participant-clustered standard errors, and the two pre-specified hypotheses (lower AI use, lower accuracy) receive statistically significant support. The distributional shift (first-order stochastic dominance in Figure 1) and the back-of-envelope calculation that treated participants substitute 85%-accurate AI advice with own answers correct only about 25% of the time add credibility to the behavioral result. However, the strong mechanistic claim of non-instrumental image concerns is not uniquely identified by the design, and the paper's defense in Section 2.5 conflates experimenter demand with the intended construct. The underlying behavioral effect is likely real, but a more cautious interpretation or an additional control condition is needed before the mechanism claim can be accepted.","major_comments":[{"comment":"The treatment bundles human observation with explicit comparison-to-average information. Participants in the treatment group are told that the research team member will review their AI usage and compare it to the average participant's behavior (Section 2.2); the control condition contains neither element. The comparison component can operate independently of any social-image channel: it may convey a descriptive norm, activate a reference point, or signal the experimenter's expectations about appropriate AI use. For instance, a participant may infer from the fact that a comparison is being made that heavy reliance on AI is unusual or undesirable, and reduce usage accordingly. The paper's statement in Section 2.2 that 'the key is not the actual comparison but the moment participants anticipate it' does not resolve this confound, because anticipation of comparison can itself trigger norm-based conformity. A design that separates the two components (e.g., a treatment with human observation but no comparison, or a comparison without a human observer) is needed to attribute the effect to non-instrumental image concerns rather than to information about norms. Without such a control, the estimated effect is an upper bound on the image-concern channel.","section":"Section 2.2 and Appendix A.3"},{"comment":"The claim that social-pressure experimenter demand is 'precisely the object of study' conflates two distinct mechanisms. If participants reduce AI use because they infer the experimenter's expectations and want to comply, that is an experimenter demand effect, not a non-instrumental image concern about how they are perceived by peers or by a neutral observer. The experimenter is not a neutral stranger: the participant knows the experimenter designed the study and can potentially evaluate the participant's behavior in light of the research question. The paper even describes the experimenter as 'an authority figure' (Section 2.5). Under that description, approval-seeking is an instrumental response to a perceived authority, with no necessary connection to the real-world construct of image concerns among colleagues, clients, or supervisors. To make the target construct credible, the paper must either argue more carefully that the experimenter's neutral language precludes demand (which it currently does not) or provide auxiliary evidence that the effect is not explained by belief about the experimenter's desired behavior.","section":"Section 2.5"},{"comment":"The self-reported discomfort measure is offered as support for the mechanism, but it cannot discriminate between non-instrumental image concerns and more general evaluation apprehension or discomfort about being monitored. The 55% increase in discomfort (42 vs. 27 participants reporting 'agree' or 'strongly agree') is consistent with the image account, but it is also exactly what one would expect if participants simply dislike being observed by a researcher, regardless of any concern about the researcher's perception of their AI use. Moreover, the second survey question ('I believe I made good use of AI recommendations') is a retrospective self-assessment and may reflect post-treatment rationalization rather than a direct measure of perceived constraint. The paper should acknowledge these limitations explicitly or supplement the survey evidence with a more targeted measure, such as beliefs about what the researcher would think of them for using AI, or an index that separates shame from mere discomfort.","section":"Section 3.3"}],"minor_comments":[{"comment":"There are several typographical errors: 'ashe-donic motives' should read 'as hedonic motives'; 'asocial pressure' should read 'a social pressure'; and the pre-registration identifier is rendered inconsistently (the acknowledgments cite #220425, while Section 2.2 gives a URL that does not contain this number).","section":"Introduction and Section 2.5"},{"comment":"The reference for de Quidt et al. (2018) is incorrect: the title 'Stigma in welfare programs' belongs to a different paper (e.g., Friedrichsen et al. or Celhay et al.), not to the demand-effects review actually cited in the text.","section":"References"},{"comment":"The sentence 'Regressing initial correct answers on Response time (controlling for image-level noise and participant fixed effects)...' is confusing: it is unclear whether the dependent variable is initial accuracy or response time, and the direction of the regression should be stated precisely.","section":"Section 3.1"},{"comment":"The paper would benefit from a discussion of multiple testing for the six pre-registered outcomes and for the exploratory heterogeneity analyses in Table B.1; the reported p-values are not adjusted, and the significant minority interaction is described as 'striking' despite the multiple comparisons.","section":"Section 3.2 and Table B.1"},{"comment":"Replication data and code are not provided; given the simplicity of the design, posting them would strengthen the paper's reproducibility and allow readers to verify the clustered standard errors and the construction of the AI-use measures.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reports a carefully executed pre-registered experiment with clear main effects, but the causal interpretation as non-instrumental image concerns is not uniquely identified. I would encourage the editor to request either a new control condition that separates human observation from comparison-to-average information, or a substantial reframing of the paper's claims to focus on the effect of anticipated evaluation on AI use, with the image-concern mechanism presented as suggestive. The current version's Section 2.5 defense does not adequately address the experimenter-demand concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a carefully run online experiment with a believable main result. Treated participants, told a researcher would review their AI usage on a one-time video call, followed AI recommendations 4.5 p.p. less and were 2.7 p.p. less accurate, with no financial incentive to do so. That's a new, policy-relevant finding for the human-AI collaboration literature. The design is genuinely clean: pre-registered, balanced groups, image fixed effects, participant-level clustering, and the null on initial choices rules out a simple effort-channel story. I also like the discomfort measure; it points in the right direction.\n\nWhere I'd push back: the treatment is a bundle. Participants in the treatment group learned not just that their usage would be seen but that it would be compared to the average participant. That second element could work through a descriptive norm or a signal about the researcher's expectations rather than non-instrumental image utility. The paper's reply in Section 2.5—that social pressure is 'precisely the object of study'—conflates experimenter demand with a peer-image construct. I don't think this is fatal: the comparison only bites because a person sees it, and the discomfort results are more consistent with social-evaluative anxiety than with mere information. But a control condition that gives comparison information without human observation would tighten the mechanism claim. Right now the contribution is real but the specific psychological channel is under-identified.\n\nOther soft spots are minor. No replication data or code is posted, so the exact estimates can't be checked. The mechanism questions are self-reported. The minority heterogeneity is post hoc, although the paper marks it exploratory and the interaction is sizable.\n\nOverall: the main effect is credible and worth taking seriously. I'd send this to a good referee. The paper needs a sharper mechanism control and a replication package, but it deserves careful review rather than a desk reject. I'd also bring it to a reading group—it will generate a useful debate about what counts as 'non-instrumental' in online experiments.","headline":"A clean pre-registered experiment showing that anticipating a stranger's review of your AI usage reduces following AI advice and performance; the image-concern interpretation is plausible but the treatment also includes an explicit social comparison, so the mechanism is not perfectly isolated.","tokens_in":14064,"tokens_out":4179,"would_cite":true,"duration_ms":41803,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Being seen using AI makes people ignore accurate advice.","keywords":["AI recommendations","non-instrumental image concerns","human-AI collaboration","algorithmic aversion","social image","online experiment","AI adoption"],"falsifier":"A control condition with identical instructions except that the AI-usage review is conducted by an automated system with no human observer, while the video call still reveals the bonus, should eliminate or greatly shrink the treatment effect on AI use and accuracy. If the reduction persists in that condition, the paper's attribution to non-instrumental image concerns—rather than to a general effect of being monitored—would be weakened.","tokens_in":13099,"feed_emoji":"👀","tokens_out":8143,"duration_ms":76577,"temperature":0.7,"pith_summary":"This paper argues that a purely social motive—concern about how one is perceived by a stranger, even when that perception carries no monetary or reputational consequences—makes people turn down accurate AI recommendations. In a pre-registered online experiment, participants who were told that a researcher would review their AI usage on a required video call followed AI advice 4.5 percentage points less often and answered 2.7 percentage points fewer questions correctly, despite bonuses tied only to performance. The claim matters because human–AI collaboration often underperforms in field settings such as medicine, courts, and hiring, and this mechanism offers an explanation for why people override algorithms even when doing so is costly. The finding also implies that the social visibility of AI use is part of what determines whether people benefit from these tools.","feed_headline":"Being watched using AI cuts its use and lowers accuracy","feed_subtitle":"Told a researcher would review their AI use, participants followed it 10% less often and got 3.3% fewer answers right.","key_machinery":"The central mechanism is the video-call review: a scripted, one-time interaction in which a researcher will look at the participant's AI-usage rate and compare it with others, while payoffs depend only on the accuracy of the final answer. This design isolates non-instrumental image concerns by making AI reliance observable to a person who can form an impression but has no power over the participant. The two-stage task (initial choice, then AI recommendation with an option to switch) is what makes the behavior measurable: 'using AI' is defined as changing one's answer to match the recommendation, and because the recommendation arrives after the initial choice, the act of following it is unambiguously attributable to the participant. The image-fixed-effects regressions with participant-level clustered standard errors are the statistical vehicle, but the conceptual engine is the contrast between an accuracy-only incentive scheme and a one-time human review.","core_discovery":"The paper's central finding is that non-instrumental image concerns—worry about being seen in a certain light by an observer, with no consequences attached to that perception—reduce the use of AI recommendations and lower task performance. In the experiment, all participants completed fifty rounds of image classification, chose an initial answer, saw an AI recommendation, and could switch; a $5 bonus was paid for a randomly selected correct final answer. Everyone also joined a brief video call to learn their bonus, but only treated participants were told that the call would include a review of their AI usage compared with the average participant. This single change reduced AI-following by 4.5 percentage points overall (and by 6.9 points when a switch was possible) and cut final accuracy by 2.7 percentage points; initial-choice accuracy, response times, and effort did not shift. The paper reads this as evidence that the anticipated gaze of another person—not any change in incentives, information, or effort—led participants to sacrifice accuracy to avoid appearing reliant on AI.","pith_inferences":["Pith editor's inference: the paper's design implies that the image cost is about the act of relying on AI, not about the gender or age of the observer; varying the observer's identity (peer vs stranger, human vs automated) would test whether the effect is specific to human judgment or generalizes to any 'audience.'","Pith editor's inference: the result suggests a behavioral explanation for racial gaps in AI adoption beyond access: if minority workers anticipate harsher judgment, they may avoid useful tools even when they are free, a mechanism that could be tested with a stereotype-threat manipulation.","Pith editor's inference: the durable digital trace of AI use is the feature that makes avoidance possible; an exact replication with a setting where reliance is harder to attribute (such as a verbal suggestion) would clarify whether the effect is about AI specifically or about any assistance that leaves an audit trail.","Pith editor's inference: the paper's self-report measure of discomfort is suggestive but not incentive-compatible; a sharper test would make participants choose between a private and a public AI-use setting with a small monetary trade-off, revealing the shadow price of being seen."],"forward_implications":["In workplaces where coworkers, managers, or clients can see how often someone relies on AI, underutilization of accurate recommendations should be stronger than in the anonymous experiment, because the image stakes are higher.","The observed 2.7-point accuracy loss implies that when participants switched away from the AI, they were replacing an 85%-accurate recommendation with their own answers, so reducing image concerns could raise decision quality without changing payoffs or information.","Making AI use less observable—for example, by aggregating usage statistics, delaying feedback about following rates, or routing review through an automated system—should increase recommendation following and accuracy.","Because minority participants showed larger treatment effects, fairness assessments of AI-assisted decisions should include how people respond to being observed using AI, not only the algorithm's outputs."],"supporting_citations":[{"why":"Supplies the image dataset, ground-truth labels, noise levels, and VGG-19 AI predictions used in the task.","marker":"Steyvers et al. (2022)"},{"why":"Defines algorithmic aversion, the baseline phenomenon the paper distinguishes from image concerns.","marker":"Dietvorst et al. (2015)"},{"why":"Establishes the video-call-with-stranger method used to induce non-instrumental image concerns.","marker":"Dellavigna et al. (2017)"},{"why":"Provides the conceptual framework for social image and hedonic motives that the paper applies to AI use.","marker":"Bursztyn and Jensen (2017)"},{"why":"Separates signaling from shaming in stigma; the paper contrasts its non-instrumental setting with their signaling-driven result.","marker":"Chandrasekhar et al. (2019)"},{"why":"Supplies evidence and terminology for experimenter demand effects that the design must rule out.","marker":"de Quidt et al. (2018)"},{"why":"Shows people under-report AI use for social reasons, motivating the image-avoidance mechanism.","marker":"Ling and Imas (2025)"}],"fun_headline_variants":["Watching people use AI makes them ignore it and perform worse","Non-instrumental image worries lead workers to reject AI help","Just knowing someone will review your AI use lowers accuracy","Observer gaze without consequences reduces AI trust and performance","People avoid AI advice when watched, sacrificing accuracy for image"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result stands or falls on the assumption that telling participants a researcher will look at their AI use changes nothing except how they feel about being seen using AI—not what they think the task is really about, not what they expect to learn about their own ability, and not a desire to please the experimenter.","fun_headline_variants_meta":{"raw":{"variants":["Watching people use AI makes them ignore it and perform worse","Non-instrumental image worries lead workers to reject AI help","Just knowing someone will review your AI use lowers accuracy","Observer gaze without consequences reduces AI trust and performance","People avoid AI advice when watched, sacrificing accuracy for image"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000687,"raw_usage":{"total_tokens":3045,"prompt_tokens":810,"completion_tokens":2235,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":426,"completion_tokens_details":{"reasoning_tokens":2155}},"tokens_in":426,"tokens_out":2235,"duration_ms":17199,"temperature":1.0,"reasoning_tokens":2155,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:02:05.196609+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A control condition with identical instructions except that the AI-usage review is conducted by an automated system with no human observer, while the video call still reveals the bonus, should eliminate or greatly shrink the treatment effect on AI use and accuracy. If the reduction persists in that condition, the paper's attribution to non-instrumental image concerns—rather than to a general effect of being monitored—would be weakened.","supporting_citations":[{"cited_title":"Bayesian modeling of human–ai complementarity","cited_arxiv_id":null,"evidence_quote":"Supplies the image dataset, ground-truth labels, noise levels, and VGG-19 AI predictions used in the task."},{"cited_title":"List, Ulrike Malmendier, and Gautam Rao","cited_arxiv_id":null,"evidence_quote":"Establishes the video-call-with-stranger method used to induce non-instrumental image concerns."},{"cited_title":"Social image and economic behavior in the field: Identifying, understanding, and shaping social pressure","cited_arxiv_id":null,"evidence_quote":"Provides the conceptual framework for social image and hedonic motives that the paper applies to AI use."},{"cited_title":"Stigma in welfare programs","cited_arxiv_id":null,"evidence_quote":"Supplies evidence and terminology for experimenter demand effects that the design must rule out."},{"cited_title":"Underreporting of ai use: The role of social desirability bias","cited_arxiv_id":null,"evidence_quote":"Shows people under-report AI use for social reasons, motivating the image-avoidance mechanism."}],"review_version":1}