Pith. sign in

REVIEW 3 major objections 7 minor 1 cited by

Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing

T0 review · 3 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Disclosing AI assistance lowers writing ratings, and only machine raters show identity-dependent preferences that vanish when disclosure is present.

desk verdict Honest, pre-registered study with a solid disclosure-penalty result, but the headline LLM demographic interaction is confounded and model-specific, so treat it as hypothesis-generating. read the letter →

arxiv 2507.01418 v1 pith:SUTQMMU7 submitted 2025-07-02 cs.CY cs.AI

classification cs.CYcs.AI
keywords AIDisclosureLargeLanguageModelsLLMratersStigmatizationpenaltyAuthoridentityRaceGender
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the growing demand for AI-use disclosure penalizes some writers more than others. Using one identical human-written news article, the authors varied the author's perceived race and gender and the presence of an AI disclosure statement, then had 1,970 human raters and 2,520 ratings from two large language models score the piece. Both humans and LLMs gave lower ratings when AI assistance was disclosed, confirming a small but consistent 'AI discount.' Only the LLM raters showed demographic interactions: one model favored Black authors and another favored woman authors when no disclosure was present, and those advantages vanished when AI use was disclosed. The paper matters because AI systems increasingly grade essays, screen job applications, and pick manuscripts, so identity-contingent machine judgments can translate into unequal access to opportunities.

What carries the argument

The machinery is a controlled evaluation design: one human-written news article, presented with a systematically varied author profile (photo, name, pronouns, identity-linked expertise) and an AI-disclosure banner, scored on a four-item perception index (trustworthiness, comprehensiveness, writing quality, and shareability, each on a 1–7 scale). Linear models with disclosure-by-identity interaction terms isolate whether the AI penalty changes with perceived race or gender, and post-hoc multiple-comparison tests identify where demographic gaps appear. The paper's named mechanism, 'vanishing alignment,' is the observation that LLM preferences favoring historically marginalized authors in the control condition disappear when the disclosure statement is added, suggesting that fairness-aligned model behavior is conditional on context rather than stable.

What would settle it

Run the same eighteen-condition experiment with bios that keep expertise wording identical across demographic conditions, changing only photo, name, and pronouns; if the LLM preferences for Black or woman authors disappear, the reported demographic interaction is an artifact of bio content rather than perceived identity. Likewise, a higher-powered human arm or a within-subjects design that reveals identity-by-disclosure interactions would falsify the paper's claim that only LLM raters exhibit these effects.

Watch

Extended reading notes

Core claim

The central claim is that disclosure of AI assistance lowers perceived writing quality for all authors, but only machine raters let the author's identity modulate that penalty. In a pre-registered 2 by 3 by 3 between-subjects experiment (disclosure by race by gender), human participants rated the same health-news article and showed a significant disclosure penalty with no significant identity-by-disclosure interactions. Two vision-language models given the same eighteen conditions also penalized disclosure, but in the no-disclosure condition one model scored Black authors higher than Asian authors and the other scored woman authors higher than man authors; post-hoc tests with multiple-comparison correction showed these demographic gaps shrank or disappeared when AI assistance was disclosed. The authors interpret this as 'vanishing alignment': fairness-oriented preferences that emerge from alignment training are fragile and can be suppressed by a single contextual cue.

Load-bearing premise

The load-bearing assumption is that the demographic manipulation changes only perceived race and gender, but in practice the bios also varied identity-linked expertise, so expertise or relevance—not identity—could drive the effects attributed to author demographics.

Editorial extensions

If this is right

  • A writer who honestly discloses AI assistance should expect a measurable, though small, drop in how their work is rated by both humans and LLMs.
  • Because LLM demographic preferences appear only without disclosure, the same article by a Black or woman author can be rated comparatively higher when AI use is hidden and comparatively lower when it is revealed.
  • The 'vanishing alignment' result implies that alignment-trained models do not hold stable fairness preferences; a contextual cue about AI involvement can remove them.
  • In high-stakes settings where automated systems screen writing, the identity-contingent machine judgments documented here could affect whose work is surfaced, hired, or promoted.
  • The absence of demographic interaction effects in human raters, if it holds, means the machine raters are not simply mirroring human bias in this task; human and algorithmic evaluation patterns diverge.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the paper's mechanism would swap only the bio's expertise phrase while keeping photo, name, and pronouns fixed; if LLM preferences follow the expertise topic rather than identity, the 'vanishing alignment' interpretation would need revision.
  • The small magnitude of the human penalty suggests real-world effects may depend on accumulation across many evaluations, and a more naturalistic writing task—longer text or multiple articles—might reveal human interaction effects this single-article design could not detect.
  • The two models favored different groups, which implies the favored group is not a universal property of LLMs but is tied to each model's training and alignment data; comparing more models would map how alignment choices create these conditional preferences.
  • Since only one phrasing of AI disclosure was tested, the paper leaves open whether different disclosures (e.g., 'edited by AI' versus 'drafted by AI') would shift the penalty or the vanishing-alignment effect.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This manuscript reports a pre-registered 2×3×3 between-subjects experiment in which 1,970 human raters and 2,520 LLM-based evaluations (GPT-4o-mini; Qwen2.5-7B-Instruct) rated an identical human-written health-news article while the author's race (Asian, Black, White), gender (man, woman, non-binary), and the presence of an AI-assistance disclosure were varied. The central claims are that both human and LLM raters impose a statistically significant but modest penalty on articles with disclosed AI assistance (roughly 0.1 to 0.15 points on a 7-point scale), and that only the LLM raters show demographic interaction effects: GPT-4o-mini rates Black authors higher and Qwen rates woman authors higher in the no-disclosure condition, with these advantages largely disappearing when AI assistance is disclosed. The authors interpret this pattern as 'vanishing alignment,' i.e., fairness-oriented preferences that are unstable under contextual cues. The human arm found no significant disclosure-by-demographic interactions. The manuscript is a short workshop paper with appendices documenting the survey interface, the design rationale, and data availability.

Significance. If the demographic interaction effects are genuine identity effects, the finding is significant for the growing practice of using LLMs to evaluate writing in hiring, publishing, and educational settings: it would imply that machine evaluators apply identity-dependent preferences that switch off precisely when AI use is disclosed, creating an asymmetrical burden of transparency. The paper has genuine methodological strengths: the pre-registration, the between-subjects deception design (only 1.1% of participants guessed the demographic-bias purpose), the reported manipulation checks, the open data and code, and the parallel human/LLM arms on identical stimuli. The disclosure-penalty main effect replicates across three evaluator types and is a solid empirical contribution. The 'vanishing alignment' interpretation is offered as a hypothesis, which is appropriate, but the identification of the demographic effects is currently undermined by the confound between identity signals and bio-expertise content documented in Appendix B; that point drives my recommendation.

major comments (3)
  1. [Appendix B / §3] The load-bearing identification problem is in Appendix B: to make demographic signals salient, the authors 'enhanced demographic signals by adding pronouns and identity-related expertise (e.g., inclusive healthcare practices within Black and LGBTQIA+ communities).' The race and gender manipulation therefore changes not only perceived identity but also the expertise content of the bio, and for a health-news article a bio mentioning Black community health is more topically matched and competence-relevant than the generic alternatives. The §3 finding that GPT-4o-mini favors Black authors and Qwen favors woman authors only in the control condition is thus equally consistent with the models reacting to topical relevance or explicit competence cues rather than to identity per se; the disappearance of the advantage under disclosure is also explainable by cue competition, in which the disclosure banner dominates the bio cue. The manuscript needs either a control condition that holds expertise content constant across demographic conditions or a robustness analysis (e.g., a bio condition that varies identity markers while keeping expertise phrasing identical) showing that the effects survive when the bio text is expertise-neutral.
  2. [§3] The LLM evaluation protocol is under-specified to the point that the central result cannot be fully verified. The paper does not report the exact prompt template, the response format, the number of prompt combinations (the text says 1,260 evaluations 'one per condition and prompt combination,' which implies multiple prompts, but neither the count nor the content is given), the sampling temperature, the number of stochastic runs, or how model outputs were aggregated into the perception scores shown in Fig. 1. In addition, Qwen2.5-7B-Instruct is described as a vision-language model that processes 'article content and author photos simultaneously,' but the standard Qwen2.5-7B-Instruct release accepts text only; if the authors used a vision-capable variant, the model name in §1 and §3 must be corrected, and if they did not, the Qwen arm never received the photo cue, which materially changes what the gender interaction claim means for that model. These details matter because the demographic manipulation includes visual identity cues and because the data availability statement invites reproduction.
  3. [§2] The human null result is asserted without a power or equivalence analysis. With 1,970 participants across 18 cells (about 109 per cell), the report of non-significant disclosure-by-identity interactions is uninterpretable without effect sizes or confidence bounds; the paper reports only the main disclosure-effect p-value (p = 0.021) and the statement that the interactions are not significant. Since the abstract's central contrast is that 'only LLM raters exhibit demographic interaction effects,' the human arm needs a sensitivity power analysis or equivalence test to support an absence claim. The cross-rater comparison should also acknowledge that the human and LLM arms differ in response noise and design (single noisy human judgments versus repeated model calls), which affects the power to detect interactions in each arm.
minor comments (7)
  1. [Fig. 1] The error bars in Fig. 1 are not defined (standard error versus confidence interval), and the by-race and by-gender panels are difficult to read; please add a legend, define the error bars, and state the reference levels used in each panel.
  2. [§2] The four Likert items that form the perception score are described, but their reliability (e.g., Cronbach's alpha) is not reported.
  3. [Throughout] Typos: 'AI assitance' in the Fig. 1 caption; 'Ithaka' for Ithaca in the affiliation block; 'University of Chicago, , WA, USA' for Mina Lee; 'photos and names were not good enough to remain impressions' in Appendix B; 'A vailable' in reference [4]; and the garbled sentence in Appendix B beginning 'This created a tension...' should be rewritten.
  4. [Table 1 / Abstract] The control condition is not a true no-disclosure baseline: it contains its own disclosure line ('Statistical information updated as of Oct. 11, 2024'). The abstract, §2, and Fig. 1 should consistently describe the manipulation as 'AI-assistance disclosure' versus 'update-only disclosure' rather than presence versus absence of disclosure.
  5. [§2 / §3] The main-effect p-values for the three rater groups (p = 0.021, p = 0.002, p = 0.026) are reported without multiple-comparison adjustment, while the post-hoc comparisons use Tukey's adjustment; please state the full family of tests and the adjustment applied.
  6. [References] Reference [17] embeds submission-status text ('under the submission to COLM 2025') inside the citation, and reference [15] is cited in §4 for 'alignment-driven over-correction' although it concerns geographic-inclusivity fine-tuning; a more directly relevant citation would strengthen that claim.
  7. [§2 / §4] The study uses a single news article in a single genre; §4 mentions genre as a possible moderator, but the abstract's word 'consistently' should be qualified to this stimulus set.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper's claims are direct experimental estimates, not derivations from their own inputs.

full rationale

This paper reports a controlled experiment rather than a formal derivation, so the circularity patterns enumerated in the review criteria do not apply. The central findings—that human and LLM raters penalize disclosed AI use, and that only LLM raters show demographic interaction effects—are estimated directly from experimental data using linear models with interaction terms. No parameter is fitted to a subset of data and then relabeled as a prediction; no 'first-principles result' is derived; no uniqueness theorem is invoked; and no ansatz is smuggled in via citation. The only self-citations are references to the authors' prior work on 'perceptual harms' (reference [11]) and on LLM research practices (reference [17]), and neither is load-bearing for the empirical claims. Appendix B's disclosure that the demographic manipulation included 'identity-related expertise (e.g., inclusive healthcare practices within Black and LGBTQIA+ communities)' is a legitimate potential confound for the interpretation of the demographic effects, because the bio wording may independently affect perceived competence or topical relevance. However, this is a validity threat, not a circularity: the estimated effects remain empirical outputs of the experiment rather than being equivalent to an input assumption by construction. The 'vanishing alignment' interpretation is a post-hoc explanatory label for observed interaction patterns, but labeling an observed pattern is not circular. The paper's result does not reduce to its inputs; it is an independent measurement, so the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters were fitted to make a derivation work; regression coefficients are estimated results. The assumptions listed are domain assumptions about the experimental setup, and no new theoretical entities are introduced beyond the interpretive label 'vanishing alignment,' which is not a mechanistic entity.

assumptions (4)
  • domain assumption A single health-news article can stand in for 'writing' when estimating disclosure and demographic effects.
    The experiment varies only the author bio and disclosure banner around one identical article, so reported effects may be specific to this article and genre.
  • domain assumption Perceived author race and gender can be manipulated through photo, name, pronouns, and bio without changing actual content.
    The manipulation relies on visual and textual cues; 52.7% of participants recalled all three elements correctly, so the cues are partially effective.
  • ad hoc to paper Identity-related expertise inserted into bios does not independently affect ratings beyond signaling identity.
    The bio statements vary with demographic condition and may alter perceived expertise, a confound not controlled in the analysis.
  • domain assumption LLM ratings obtained through unspecified prompts are comparable to human perceptual judgments for the same task.
    The paper does not include the exact prompts, temperature, number of runs, or aggregation method for GPT-4o-mini and Qwen2.5-7B-Instruct.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing." pith.science (2026). https://pith.science/paper/SUTQMMU7

@misc{pith2026250701418,
  author       = {Pith},
  title        = {Pith review of: Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SUTQMMU7}},
  note         = {Machine review of arXiv:2507.01418}
}
read the original abstract

As AI integrates in various types of human writing, calls for transparency around AI assistance are growing. However, if transparency operates on uneven ground and certain identity groups bear a heavier cost for being honest, then the burden of openness becomes asymmetrical. This study investigates how AI disclosure statement affects perceptions of writing quality, and whether these effects vary by the author's race and gender. Through a large-scale controlled experiment, both human raters (n = 1,970) and LLM raters (n = 2,520) evaluated a single human-written news article while disclosure statements and author demographics were systematically varied. This approach reflects how both human and algorithmic decisions now influence access to opportunities (e.g., hiring, promotion) and social recognition (e.g., content recommendation algorithms). We find that both human and LLM raters consistently penalize disclosed AI use. However, only LLM raters exhibit demographic interaction effects: they favor articles attributed to women or Black authors when no disclosure is present. But these advantages disappear when AI assistance is revealed. These findings illuminate the complex relationships between AI disclosure and author identity, highlighting disparities between machine and human evaluation patterns.

Figures

Figures reproduced from arXiv: 2507.01418 by the authors.

Figure 1
Figure 1. Perception scores across 9 evaluation conditions, comparing human raters (top row) with large language models GPT-4o-mini [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. (a) Participants read a health-news article with author photo, name, pronouns, bio, and the AI-disclosure banner. (b) While [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. More Human or More AI? Visualizing Human-AI Collaboration Disclosures in Journalistic News Production

    cs.HC 2026-01 conditional novelty 6.0 of 10

    Disclosure visualization format systematically shifts readers' perceptions of human vs AI contribution: role-based timelines amplify perceived AI role in mostly human articles, while task-based timelines make mostly A...

Reference graph

Works this paper leans on

23 extracted references · 20 canonical work pages · cited by 1 Pith paper

  1. [1]

    Susanne Althoff. 2016. Algorithms Could Save Book Publishing—But Ruin Novels. Wired (2016). https://www.wired.com/2016/09/bestseller-code/

  2. [2]

    Tae Hyun Baek, Jungkeun Kim, and Jeong Hyun Kim. 2024. Effect of disclosing AI-generated content on prosocial advertising evaluation. International Journal of Advertising (2024), 1–22

  3. [3]

    Karim Benharrak, Tim Zindulka, Florian Lehmann, Hendrik Heuer, and Daniel Buschek. 2024. Writer-defined AI personas for on-demand feedback generation. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems . 1–18

  4. [4]

    Jordan Boyd-Graber, Naoaki Okazaki, and Anna Rogers. 2023. ACL 2023 policy on AI writing assistance. A vailable under: https://2023. aclweb. org/blog/ACL-2023-policy/[15.11. 2023] (2023)

  5. [5]

    Daniel Buschek. 2024. Collage is the New Writing: Exploring the Fragmentation of Text and User Interfaces in AI Tools. In Proceedings of the 2024 ACM Designing Interactive Systems Conference . 2719–2737

  6. [6]

    Won Ik Cho, Eunjung Cho, and Kyunghyun Cho. 2023. PaperCard for reporting machine assistance in academic writing. arXiv preprint arXiv:2310.04824 (2023)

  7. [7]

    Katherine M Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E Zhang, Tan Zhi-Xuan, Mark Ho, Vikash Mansinghka, et al. 2024. Building machines that learn and think with people. Nature human behaviour 8, 10 (2024), 1851–1863. Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgment...

  8. [8]

    Dayeon Eom, Amanda L Molder, Helen A Tosteson, Emily L Howell, Meredith DeSalazar, Elliot Kirschner, Sarah S Goodwin, and Dietram A Scheufele. 2025. Race and gender biases persist in public perceptions of scientists’ credibility. Scientific Reports 15, 1 (2025), 11021

Show all 23 references
  1. [9]

    Jane Hanson. 2023. AI Is Replacing Humans In The Interview Process - What You Need To Know To Crush Your Next Video Interview. Forbes (sep 2023). https://www.forbes.com/sites/janehanson/2023/09/30/ai-is-replacing-humans-in-the-interview-processwhat-you-need-to-know-to-crush- y...

  2. [10]

    Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. AI generates covertly racist decisions about people based on their dialect. Nature 633, 8028 (2024), 147–154

  3. [11]

    Kowe Kadoma, Danaë Metaxa, and Mor Naaman. 2024. Generative AI and Perceptual Harms: Who’s Suspected of using LLMs? arXiv preprint arXiv:2410.00906 (2024)

  4. [12]

    Peters Keaton. 2024. Texas will use computers to grade written answers on this year’s STAAR tests. The Texas Tribune (apr 2024). https: //www.texastribune.org/2024/04/09/staar-artificial-intelligence-computer-grading-texas/

  5. [13]

    Taewan Kim, Donghoon Shin, Young-Ho Kim, and Hwajung Hong. 2024. DiaryMate: Understanding User Perceptions and Experience in Human-AI Collaboration for Personal Journaling. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems . 1–15

  6. [14]

    Elena Klaas and Mark Boukes. 2022. A woman’s got to write what a woman’s got to write: The effect of journalist’s gender on the perceived credibility of news articles. Feminist media studies 22, 3 (2022), 571–587

  7. [15]

    K Shantha Kumari, S Udhaya Shree, and C Calarany. 2024. Dynamic Region-Aware Fine-Tuning: Enhancing Geographic Inclusivity in Large Language Models. In 2024 IEEE 16th International Conference on Computational Intelligence and Communication Networks (CICN) . IEEE, 1418–1423

  8. [16]

    Mina Lee, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A Alghamdi, et al. 2024. A design space for intelligent and interactive writing assistants. In Proceedings of the 2024 ...

  9. [17]

    Zhehui Liao, Maria Antoniak, Inyoung Cheong, Evie Yu-Yen Cheng, Ai-Heng Lee, Kyle Lo, Joseph Chee Chang, and Amy X. Zhang. 2024. LLMs as Research Tools: A Large Scale Survey of Researchers’ Usage and Perceptions. arXiv:2411.05025 [cs.CL] https://arxiv.org/abs/2411.05025 under ...

  10. [18]

    Sue Lim and Ralf Schmälzle. 2024. The effect of source disclosure on evaluation of AI-generated messages. Computers in Human Behavior: Artificial Humans 2, 1 (jan 2024), 100058. https://doi.org/10.1016/j.chbah.2024.100058

  11. [19]

    Ingrid Lunden. 2024. Inkitt, the self-publishing platform using AI to develop bestsellers, nabs $37M. TechCrunch (feb 2024). https://techcrunch.com/ 2024/02/26/inkitt-ai-publishing-37-million/

  12. [20]

    David Rice. 2024. How AI Is Transforming Employee Performance Reviews. People Managing People (oct 2024). https://peoplemanagingpeople. com/articles/ai-performance-review/

  13. [21]

    Chirag Shah and Emily M Bender. 2022. Situating search. In Proceedings of the 2022 Conference on Human Information Interaction and Retrieval . 221–232

  14. [22]

    It Felt Like Having a Second Mind

    Qian Wan, Siying Hu, Yu Zhang, Piaohong Wang, Bo Wen, and Zhicong Lu. 2024. “It Felt Like Having a Second Mind”’: Investigating Human-AI Co-creativity in Prewriting with Large Language Models. Proceedings of the ACM on Human-Computer Interaction 8, CSCW1 (2024), 1–26

  15. [23]

    Ruyuan Wan, Simret Araya Gebreegziabher, Toby Jia-Jun Li, and Karla Badillo-Urquiola. 2024. CoCo Matrix: Taxonomy of Cognitive Contributions in Co-writing with Intelligent Agents. In Proceedings of the 16th Conference on Creativity & Cognition . 504–511. A Survey Interface Our...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.