Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Labeling Synthetic Content: User Perceptions of Warning Label Designs for AI-generated Content on Social Media

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Warning labels make users believe social media images are AI-made, but do not change how they engage with them.

desk verdict A useful design-space paper with a credible null on engagement, but the headline trust-by-design claim rests on an unreported between-group ANOVA. read the letter →

arxiv 2503.05711 v1 pith:3TIEGJF3 submitted 2025-02-14 cs.HC cs.AIcs.CYcs.ET

classification cs.HCcs.AIcs.CYcs.ET
keywords warninglabeldesignAI-generatedcontentdeepfakeuserperceptionsocialmediaengagementtrustinlabelshuman-computerinteractionlabeling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to define a design space for warning labels on AI-generated social media content and to test whether such labels work. It derives ten label prototypes from four dimensions—sentiment, iconography and color, position, and level of detail—and runs a 911-participant experiment with a control group. The central result is that any label made users more likely to believe an image was AI-generated or edited, while trust in the label itself depended on its design. At the same time, labels did not significantly change likes, comments, or shares, although political content drew more engagement than entertainment content. The authors conclude that labels inform belief and build platform trust but are not a standalone fix for engagement with synthetic media.

What carries the argument

The carrying instrument is a four-dimensional design space for AI-content warning labels—label sentiment (hazardous, neutral without AI, neutral with AI), iconography and color (warning icon in red, neutral icon, no icon), position (on content obscuring, above content, on content non-obscuring), and level of detail (simple text versus detailed provenance). From this space the authors built ten prototype labels and embedded them into eight realistic social-media image posts, four AI-generated or edited and four real, split evenly between political and entertainment content. The experiment randomly assigned 911 participants to the control or one of the ten treatments and measured belief that content was AI-generated, trust in the label, trust in the platform, and like, comment, and share intentions using five-point Likert items. This design lets the paper attribute differences in belief and trust to label design while using the control group to isolate the mere presence of a label.

What would settle it

Re-analyze the posted dataset with linear mixed models that include random intercepts for participant and image; if the belief effect (reported as F(9,6468)=4.05, p<.001) becomes non-significant or the trust differences between label designs disappear, then the headline findings are artifacts of the independence assumption. A field experiment with real like, comment, and share data would also settle whether the null engagement finding holds outside survey intentions.

Watch

Extended reading notes

Core claim

The paper's core claim is that warning labels for AI-generated content are effective as transparency signals but weak as behavior-change tools. Across ten label designs tested against a no-label control, the presence of a label significantly raised users' belief that content was AI-generated, deepfake, or AI-edited regardless of their familiarity with the content, with the strongest effects for labels reading 'Made with AI' with a neutral diamond icon. Trust in labels varied by design: the Content Credentials-style detailed label earned the highest trust, while red 'Deepfake' warnings scored lower and were read as ambiguous. Label trust and platform trust were strongly correlated, yet engagement—liking, commenting, sharing—did not differ significantly between labeled and unlabeled content; content category, not labeling, drove engagement differences. The authors present this as evidence that label design choices matter for belief and trust, and that labels should be part of a broader intervention strategy rather than a standalone solution.

Load-bearing premise

The statistical conclusion that labels change belief rests on treating each of the eight image ratings by the same participant as independent observations, even though those ratings are correlated within participant and within image; the authors acknowledge in Section 5.4 that repeated-measures analysis can affect significance and plan linear mixed models instead.

Editorial extensions

If this is right

  • Platforms that adopt any of the ten tested label designs can expect users to be more likely to believe an image is AI-generated or edited, even when the image is real.
  • Label wording matters: designs using 'Made with AI' with a neutral icon produced the strongest belief effect, while designs using the word 'Deepfake' were less clear to users and scored lower on trust.
  • Trust in the label transfers to trust in the platform: label trust and platform trust correlated at r(85)=.73, p<.001.
  • Labels alone will not reduce engagement: no significant difference in like, comment, or share appeared between labeled and unlabeled conditions, but political content drew more reactions and comments than entertainment content.
  • Familiarity with the content does not change whether users believe the label's assertion that content is AI-generated.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the engagement null holds in real-world settings, regulators should not expect transparency labels alone to slow the spread of AI-generated misinformation; complementary measures such as friction or accuracy prompts would be needed.
  • The finding that labels made users believe even real images were AI-generated suggests a false-positive cost: over-labeling could cultivate blanket skepticism toward authentic content, an effect the paper does not test.
  • A testable extension is to vary label prevalence across a feed to see whether trust erodes or the belief effect weakens under warning fatigue, and to measure actual clicks and shares rather than Likert intentions.
  • The design space could be combined with provenance metadata to test whether a two-layer label—simple 'Made with AI' text plus click-through details—captures both belief and trust, since the most trusted design was not the one that best communicated that content was AI-generated.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a four-dimensional design space for warning labels on AI-generated social media content (label sentiment, icon/color, position, level of detail) and evaluates ten prototype labels derived from it in a randomized between-subjects experiment with 911 Prolific participants. Participants rated eight images (four political, four entertainment; half AI-generated/deepfake, half real) either without a label (control) or with one of ten label designs. The authors report that labels significantly increased belief that content was AI-generated (RQ2), that trust in the label varied by design (RQ3), that label trust correlated with platform trust (RQ4), and that labels did not significantly change like/comment/share intentions (RQ5), while content type did affect engagement. The paper contributes a design-space framework and empirical evidence on label efficacy, with data and code shared on OSF.

Significance. This is a timely and well-scoped empirical contribution to the emerging literature on labeling AI-generated content. The randomized design, the use of realistic social-media-style stimuli, and the open sharing of materials and code are clear strengths. The design space itself is a useful synthesis of platform practice and prior work, and the null finding for engagement (F5) is a valuable counterpoint to the misinformation-labeling literature, suggesting that AI-content labels may not reduce sharing and liking in the way traditional fact-check labels do. If the trust-by-design finding were properly substantiated, the results would give platform designers concrete guidance. However, the statistical support for the headline trust-by-design claim and the trust-platform correlation is currently incomplete, and the repeated-measures structure of the data is not appropriately modeled in most analyses. These issues are fixable within the scope of the manuscript, but they need to be addressed before the central claims are reliable.

major comments (3)
  1. [§4.3.1, Table 3, Finding F3] The reported analysis does not support the claim that trust in the label significantly varied based on the label design. Table 3 reports one-way ANOVAs run separately within each treatment group across the eight images; a significant F for a given group means trust in that label differed across the eight images, not that Design 1 was trusted differently from Design 10. The between-group ANOVA is mentioned only parenthetically ('To confirm the results in trust with the design samples, we performed a one-way ANOVA between the groups, it showed that TRT10 with warning label design 10 had the highest'), with no F statistic, degrees of freedom, p-value, or effect size reported in the main text, and the pairwise comparisons are deferred to a supplementary file. The abstract's statement that 'their trust in the label significantly varied based on the label design' and Finding F3 ('Design Sample 10 ... elicited the highest trust level') therefore lack a demonstrated statistical basis in the manuscript. Please report the full between-group ANOVA with pairwise comparisons in the main text, or temper the claim to what the within-group analysis actually shows.
  2. [§4.2.1, §4.2.2, §4.5.1–§4.5.4, §5.4] The ANOVAs treat the roughly 7,000 image-level ratings as independent observations, even though each of the 911 participants rated eight images within a single condition. This nesting within participant and within image can inflate test statistics and shrink p-values. The authors acknowledge this in Section 5.4 ('Our repeated statistical analysis may affect the significance test') and state an intention to use linear mixed models, but the results presented in the main text are all based on the independence assumption. Until mixed models or some appropriate correction are reported, the significance of the belief effect (F1) and the between-treatment comparisons in Section 4.2.2 should be treated as provisional. At minimum, the paper should clarify the unit of analysis, justify the independent-observation assumption, or replace the affected ANOVAs with models that account for the repeated-measures structure.
  3. [§4.4.2, Finding F4] The correlation between trust in the label and trust in the platform is reported as r(85) = .73, p < .001, but the degrees of freedom do not match any obvious unit of analysis. The study has 10 treatment groups and 911 participants; a group-level correlation would yield df = 8, while a participant-level correlation would yield df ≈ 900. The reported df = 85 is inconsistent with both. Please clarify the unit of analysis on which the correlation was computed and report the corresponding sample size. If the correlation was computed on aggregated group means, the claim that individual users who trust the label also trust the platform is much weaker than presented.
minor comments (5)
  1. [§4.2.1–§4.2.2] The between-group ANOVAs in these two subsections report nearly identical F(9, 6468) values and p-values, yet Section 4.2.1 describes a comparison of control vs. treatment groups while Section 4.2.2 describes a comparison among the 10 treatment groups. It is unclear whether these are the same test presented twice or two different tests with coincidentally identical statistics; please clarify the actual models and report the N used in each analysis.
  2. [§4.5.1–§4.5.3] The text refers to 'control group (n = 648)' and 'treatment group (n = 6,593)' for the engagement analyses, but these are image-level observations, not numbers of participants. This is confusing for readers and should be explicitly labeled as image-level observations (e.g., '648 image ratings from 81 control participants').
  3. [Figure 9 caption] The caption contains a typo: 'ward-cloud' should be 'word cloud'.
  4. [References] Several references are incomplete, e.g., [12] lists no author names and no publication year, and a few other entries lack venue or retrieval details. Please complete the reference list.
  5. [§4.2.1] The effect size for the treatment effect on belief is reported as eta-squared = 0.006. This is a very small effect, and while the paper acknowledges small effect sizes in the Discussion, the abstract's wording ('a significant effect') may overstate practical importance; a brief note in the abstract or results about the small magnitude would help calibrate reader expectations.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity (score 0): the paper is a randomized between-subjects experiment whose design dimensions were fixed a priori and whose findings are empirical comparisons of direct Likert measures, with no fitted parameters and no load-bearing self-citations.

full rationale

This paper is an empirical study with no formal derivation, no fitted parameters, and no load-bearing self-citations, so its results cannot reduce to their inputs by construction. The design space (Section 3.1) was articulated a priori by expert brainstorming, prior research, and platform examples ('Four researchers who are experts in mis/disinformation and user-centered design articulated the initial dimensions'), and the ten prototypes were selected from regulatory and platform considerations before data collection; none of the four dimensions (sentiment, icon/color, position, detail) is defined in terms of the outcome variables. Dependent variables are direct single-item Likert measures (Q3 belief, Q8 trust-in-label, Q9 trust-in-platform, Q4-Q6 engagement), not composites or model outputs, so the reported findings are empirical estimates from a randomized between-subjects comparison (911 participants, 11 conditions) rather than restatements of inputs. The reference list contains no papers by the present authors, so no self-citation chain is load-bearing. The two vulnerabilities noted in review are statistical-validity issues, not circularity: (1) Table 3 reports within-group ANOVAs across eight images, whereas the RQ3 claim that trust varies by design rests on a between-group ANOVA reported only parenthetically, and (2) the eight image-level ratings per participant are treated as independent; the paper explicitly acknowledges this ('Our repeated statistical analysis may affect the significance test') and plans linear mixed models. Per the review rules, an unsupported or mis-specified statistical claim is a correctness risk, not a circular step, and the paper's own limitation statement is in scope and weighs toward caution on those findings without implying circularity. The design-space contribution is explicitly framed as an initial framework subject to refinement, with the subjective selection of samples acknowledged ('the choice of any given design will be subjective'), which further precludes any claim that the framework's dimensions are validated by the same data that produced them.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central empirical claims rest on domain assumptions about self-report validity, image ground truth, prototype representativeness, and statistical independence. There are no fitted free parameters and no invented entities.

assumptions (4)
  • domain assumption Self-reported Likert responses measure belief, trust, and engagement intentions sufficiently for the conclusions.
    The experiment relies on Q3, Q8-Q11 Likert scales as direct measures of belief, trust, and engagement, without validation against actual behavior.
  • domain assumption The eight images are correctly classified as real versus AI-generated/deepfake and are representative of political and entertainment social media content.
    Ground truth is taken from online sources and platform examples; no independent verification or pilot rating is reported.
  • ad hoc to paper The 10 label prototypes embody the design dimensions without uncontrolled confounds.
    Labels vary on multiple dimensions at once, and the authors explicitly treat each design as a separate treatment rather than a factorial design, so differences between samples cannot be attributed to a single dimension.
  • domain assumption Image-level observations from the same participant are treated as independent in the ANOVAs.
    The paper uses one-way and two-way ANOVAs on roughly 7,000 image-level ratings; repeated measures by participant and image are not modeled, which the paper acknowledges in Section 5.4.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Labeling Synthetic Content: User Perceptions of Warning Label Designs for AI-generated Content on Social Media." pith.science (2026). https://pith.science/paper/3TIEGJF3

@misc{pith2026250305711,
  author       = {Pith},
  title        = {Pith review of: Labeling Synthetic Content: User Perceptions of Warning Label Designs for AI-generated Content on Social Media},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3TIEGJF3}},
  note         = {Machine review of arXiv:2503.05711}
}
read the original abstract

In this research, we explored the efficacy of various warning label designs for AI-generated content on social media platforms e.g., deepfakes. We devised and assessed ten distinct label design samples that varied across the dimensions of sentiment, color/iconography, positioning, and level of detail. Our experimental study involved 911 participants randomly assigned to these ten label designs and a control group evaluating social media content. We explored their perceptions relating to 1. Belief in the content being AI-generated, 2. Trust in the labels and 3. Social Media engagement perceptions of the content. The results demonstrate that the presence of labels had a significant effect on the users belief that the content is AI generated, deepfake, or edited by AI. However their trust in the label significantly varied based on the label design. Notably, having labels did not significantly change their engagement behaviors, such as like, comment, and sharing. However, there were significant differences in engagement based on content type: political and entertainment. This investigation contributes to the field of human computer interaction by defining a design space for label implementation and providing empirical support for the strategic use of labels to mitigate the risks associated with synthetically generated media.

Figures

Figures reproduced from arXiv: 2503.05711 by the authors.

Figure 1
Figure 1. Examples design choices adopted for AI-generated content warning labels, highlighted by blue arrow annotation, on different [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Experiment Procedure: 911 samples qualified, all of them took the same pre-survey, then randomly assigned to one of the treatment [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Design Samples derived from design space, showing design options chosen for each in parentheses. [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Mean scores for (a) likelihood to engage; and (b) support for making informed decisions to engage [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Belief-Of-Content control vs. treatment 4.2.2 Which Warning Label Design Led to Higher Belief that Content was AI-generated? To understand which label design had the most effect on users’ belief that the content was AI-generated, we examined our results using a between…
Figure 6
Figure 6. Figure 6: User trust levels for (a) the warning labels; and (b) for the Social Media platform [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Mean engagement likelihood by content category for (a) “Like” reaction; (b) commenting; and (c) sharing [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: Mean scores for (a) likelihood to engage; and (b) support for making informed decisions to engage [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: Ward-cloud capturing how Users encounter deepfake or AI generated content—they found mostly in the form of Videos, images and [PITH_FULL_IMAGE:figures/full_fig_p017_9.png]
Figure 10
Figure 10. Figure 10: The images used to evaluate the user perceptions [PITH_FULL_IMAGE:figures/full_fig_p024_10.png]
Figure 11
Figure 11. Figure 11: Types of social media use by the participants with the frequency of user [PITH_FULL_IMAGE:figures/full_fig_p026_11.png]
Figure 12
Figure 12. Figure 12: Relationship Between warning Label trust and trust to the Platform showed a significant correlation [PITH_FULL_IMAGE:figures/full_fig_p032_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Signals of Provenance: Practices & Challenges of Navigating Indicators in AI-Generated Media for Sighted and Blind Individuals

    cs.HC 2025-05 conditional novelty 6.0 of 10

    Both sighted and blind/low-vision users frequently overlook platform AI labels and rely on titles, comments, and other content cues, with blind users further hindered by inaccessible label design.

Reference graph

Works this paper leans on

54 extracted references · 50 canonical work pages · cited by 1 Pith paper

  1. [1]

    S Andı and J Akesson. 2021. Nudging away false news: Evidence from a social norms experiment. Digital Journalism, 9 (1), 106-125

  2. [2]

    Announcement. [n. d.]. New Disclosures and Labels for Generative AI Content on YouTube . https://support.google.com/youtube/thread/264550152/new- disclosures-and-labels-for-generative-ai-content-on-youtube?hl=en

  3. [3]

    Md Momen Bhuiyan, Hayden Whitley, Michael Horning, Sang Won Lee, and Tanushree Mitra. 2021. Designing transparency cues in online news platforms to promote trust: Journalists’& consumers’ perspectives. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–31

  4. [4]

    Md Momen Bhuiyan, Kexin Zhang, Kelsey Vick, Michael A Horning, and Tanushree Mitra. 2018. FeedReflect: A tool for nudging users to assess news credibility on twitter. In Companion of the 2018 ACM conference on computer supported cooperative work and social computing . 205–208

  5. [5]

    Leticia Bode and Emily K Vraga. 2015. In related news, that was wrong: The correction of misinformation through related stories functionality in social media. Journal of Communication 65, 4 (2015), 619–638

  6. [6]

    Olivia Burrus, Amanda Curtis, and Laura Herman. 2024. Unmasking AI: Informing Authenticity Decisions by Labeling AI-Generated Content. Interactions 31, 4 (2024), 38–42

  7. [7]

    Coalition for Content Provenance and Authenticity

    c2pa 2024. Coalition for Content Provenance and Authenticity . Retrieved Sept 2024 from https://c2pa.org/

  8. [8]

    Content Credintials - How it works

    CCredential 2024. Content Credintials - How it works . Retrieved Sept 2024 from https://contentauthenticity.org/how-it-works

Show all 54 references
  1. [9]

    Man-pui Sally Chan, Christopher R Jones, Kathleen Hall Jamieson, and Dolores Albarracín. 2017. Debunking: A meta-analysis of the psychological efficacy of messages countering misinformation. Psychological science 28, 11 (2017), 1531–1546

  2. [10]

    Content Authenticity Initiative

    ContentAUthen 2022. Content Authenticity Initiative. Retrieved Sept 2024 from https://contentauthenticity.org/

  3. [11]

    Shawn Domgaard and Mina Park. 2021. Combating misinformation: The effects of infographics in verifying false vaccine news. Health Education Journal 80, 8 (2021), 974–986

  4. [12]

    Ziv Epstein, Antonio Alonso Arechar, and David Rand. [n. d.]. What label should be applied to content produced by generative AI? ([n. d.])

  5. [13]

    Daniel Ershov and Juan S Morales. 2024. Sharing News Left and Right: Frictions and Misinformation on Twitter. The Economic Journal (2024), 27

  6. [14]

    Artificial Intelligence Act of the EU Parliment

    EuAIact 2024. Artificial Intelligence Act of the EU Parliment . Retrieved Sept 2023 from https://www.europarl.europa.eu/doceo/document/TA-9-2024- 0138_EN.pdf

  7. [15]

    Lisa Fazio. 2020. Pausing to consider why a headline is true or false can help reduce the sharing of false news. Harvard Kennedy School Misinformation Review 1, 2 (2020)

  8. [16]

    KJ Kevin Feng, Nick Ritchie, Pia Blumenthal, Andy Parsons, and Amy X Zhang. 2023. Examining the Impact of Provenance-Enabled Media on Trust and Accuracy Perceptions. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2 (2023), 1–42

  9. [17]

    Four Corners Project

    forcorner 2022. Four Corners Project. 2022. An initiative to increase authorship and credibility in visual media. Retrieved Sept 2024 from https://fourcornersproject. org/en/

  10. [18]

    Mingkun Gao, Ziang Xiao, Karrie Karahalios, and Wai-Tat Fu. 2018. To label or not to label: The effect of stance and credibility labels on readers’ selection and perception of news articles. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 1–16

  11. [19]

    Michael Geers, Briony Swire-Thompson, Philipp Lorenz-Spreen, Stefan M Herzog, Anastasia Kozyreva, and Ralph Hertwig. 2023. The online misinformation engagement framework. Current Opinion in Psychology (2023), 101739. 22 User Perceptions of Warning Label Designs for AI-generate...

  12. [20]

    Chen Guo, Nan Zheng, and Chengqi Guo. 2023. Seeing is not believing: a nuanced view of misinformation warning efficacy on video-sharing social media platforms. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2 (2023), 1–35

  13. [21]

    Katrin Hartwig, Frederic Doell, and Christian Reuter. 2024. The Landscape of User-centered Misinformation Interventions-A Systematic Literature Review. Comput. Surveys 56, 11 (2024), 1–36

  14. [22]

    Sarah Hawa, Lanita Lobo, Unnati Dogra, and Vijaya Kamble. 2021. Combating misinformation dissemination through verification and content driven recommendation. In 2021 Third International Conference on Intelligent Communication Technologies and Virtual Mobile Networks (ICICV) ....

  15. [23]

    Melissa Heikkilä. 2024. Google is finally taking action to curb non-consensual deepfakes. (2024)

  16. [24]

    Farnaz Jahanbakhsh, Amy X Zhang, and David R Karger. 2022. Leveraging structured trusted-peer assessments to combat misinformation. Proceedings of the ACM on Human-computer Interaction 6, CSCW2 (2022), 1–40

  17. [25]

    Chenyan Jia, Alexander Boltz, Angie Zhang, Anqing Chen, and Min Kyung Lee. 2022. Understanding effects of algorithmic vs. community label on perceived accuracy of hyper-partisan misinformation. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2 (2022), 1–27

  18. [26]

    Anastasia Kozyreva, Sam Wineburg, Stephan Lewandowsky, and Ralph Hertwig. 2023. Critical ignoring as a core competence for digital citizens. Current Directions in Psychological Science 32, 1 (2023), 81–88

  19. [27]

    Mateusz Łabuz. 2024. Deep fakes and the Artificial Intelligence Act—An important signal or a missed opportunity? Policy & Internet (2024)

  20. [28]

    Stephan Lewandowsky, John Cook, Ullrich Ecker, Dolores Albarracin, Panayiota Kendeou, Eryn J Newman, Gordon Pennycook, Ethan Porter, David G Rand, David N Rapp, et al. 2020. The debunking handbook 2020. (2020)

  21. [29]

    Andrew Lewis, Patrick Vu, Raymond M Duch, and Areeq Chowdhury. 2022. Do content warnings help people spot a deepfake? Evidence from two experiments. OSF Preprint (2022)

  22. [30]

    Gionnieve Lim and Simon T Perrault. 2023. Effects of automated misinformation warning labels on the intents to like, comment and share posts. In Proceedings of the 11th International Conference on Human-Agent Interaction . 299–305

  23. [31]

    Cameron Martel and David G Rand. 2023. Misinformation warning labels are widely effective: A review of warning effects and their moderating features. Current Opinion in Psychology (2023), 101710

  24. [32]

    Paul Mena. 2020. Cleaning up social media: The effect of warning labels on likelihood of sharing false news on Facebook. Policy & internet 12, 2 (2020), 165–183

  25. [33]

    Label AI content on Facebook

    Meta 2024. Label AI content on Facebook . Retrieved Nov 2024 from https://www.facebook.com/help/7434563519957988/iphone-app-help/?helpref=platform_ switcher&cms_platform=iphone-app&_rdr

  26. [34]

    Sebasti ao Miranda, David Nogueira, Afonso Mendes, Andreas Vlachos, Andrew Secker, Rebecca Garrett, Jeff Mitchel, and Zita Marinho. 2019. Automated fact checking in the news room. In The world wide web conference . 3579–3583

  27. [35]

    Fake news

    Maria D Molina, S Shyam Sundar, Thai Le, and Dongwon Lee. 2021. “Fake news” is not simply false information: A concept explication and taxonomy of online content. American behavioral scientist 65, 2 (2021), 180–212

  28. [36]

    In Transparency We Trust? Evaluating the Effectiveness of Watermarking and Labeling AI-Generated Content

    Mozilla 2024. In Transparency We Trust? Evaluating the Effectiveness of Watermarking and Labeling AI-Generated Content . Retrieved Sept 2023 from https://foundation.mozilla.org/en/research/library/in-transparency-we-trust/research-report/

  29. [37]

    Meta’s approach to labeling AI generated and manupiliative media

    Mozilla 2024. Meta’s approach to labeling AI generated and manupiliative media . Retrieved April 2025 from https://about.fb.com/news/2024/04/metas- approach-to-labeling-ai-generated-content-and-manipulated-media/

  30. [38]

    Jack Nassetta and Kimberly Gross. 2020. State media warning labels can counteract the effects of foreign misinformation. Harvard Kennedy School Misinformation Review (2020)

  31. [39]

    Global Affairs Nick Clegg, President. [n. d.]. Labeling AI-Generated Images on Facebook, Instagram and Threads . https://about.fb.com/news/2024/02/labeling- ai-generated-images-on-facebook-instagram-and-threads/

  32. [40]

    Anne Oeldorf-Hirsch, Mike Schmierbach, Alyssa Appelman, and Michael P Boyle. 2020. The ineffectiveness of fact-checking labels on news memes and articles. Mass Communication and Society 23, 5 (2020), 682–704

  33. [41]

    Abby Ohlheiser. 2016. What Facebook hasn’t said about its plan to fight fake news. Washingtonpost. com (2016)

  34. [42]

    Andrew Otis. 2024. The effects of transparency cues on news source credibility online: an investigation of ‘opinion labels’. Journalism 25, 1 (2024), 198–217

  35. [43]

    Orestis Papakyriakopoulos and Ellen Goodman. 2022. The impact of Twitter labels on misinformation spread and user engagement: Lessons from Trump’s election tweets. In Proceedings of the ACM web conference 2022 . 2541–2551

  36. [44]

    Partnership on AI (PAI)

    PartnershipAI 2024. Partnership on AI (PAI). Retrieved Sept 2024 from https://partnershiponai.org/about/

  37. [45]

    Gordon Pennycook, Adam Bear, Evan T Collins, and David G Rand. 2020. The implied truth effect: Attaching warnings to a subset of fake news headlines increases perceived accuracy of headlines without warnings. Management science 66, 11 (2020), 4944–4957

  38. [46]

    Gordon Pennycook, Ziv Epstein, Mohsen Mosleh, Antonio A Arechar, Dean Eckles, and David G Rand. 2021. Shifting attention to accuracy can reduce misinformation online. Nature 592, 7855 (2021), 590–595

  39. [47]

    Gordon Pennycook and David G Rand. 2022. Nudging social media toward accuracy. The Annals of the American Academy of Political and Social Science 700, 1 (2022), 152–164

  40. [48]

    Fabian Prochazka and Wolfgang Schweiger. 2019. How to measure generalized trust in news media? An adaptation and test of scales. Communication methods and measures 13, 1 (2019), 26–42

  41. [49]

    Haeseung Seo, Aiping Xiong, and Dongwon Lee. 2019. Trust it or not: Effects of machine-learning warnings in helping individuals mitigate misinformation. In Proceedings of the 10th ACM Conference on Web Science . 265–274

  42. [50]

    Christian Stransky, Dominik Wermke, Johanna Schrader, Nicolas Huaman, Yasemin Acar, Anna Lena Fehlhaber, Miranda Wei, Blase Ur, and Sascha Fahl

  43. [51]

    Chloe Wittenberg, Ziv Epstein, Adam J Berinsky, and David G Rand. 2024. Labeling AI-Generated Content: Promises, Perils, and Future Directions. (2024)

  44. [52]

    Manuel Wörsdörfer. 2024. Biden’s Executive Order on AI and the EU’s AI Act: A Comparative Computer-Ethical Analysis. Philosophy & Technology 37, 3 (2024), 74

  45. [53]

    Danni Xu, Shaojing Fan, and Mohan Kankanhalli. 2023. Combating misinformation in the era of generative AI models. In Proceedings of the 31st ACM International Conference on Multimedia . 9291–9298. 23 CHI ’25, April 26-May 1, 2025, Yokohama, Japan Dilrukshi Gamage, Dilki Sewwan...

  46. [2021]

    In Seventeenth Symposium on Usable Privacy and Security (SOUPS 2021)

    On the Limited Impact of Visualizing Encryption: Perceptions of{E2E} Messaging Security. In Seventeenth Symposium on Usable Privacy and Security (SOUPS 2021). 437–454

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.