Pith. sign in

REVIEW 5 major objections 6 minor 24 references

What's Behind the Magic? Audiences Seek Artistic Value in Generative AI's Contributions to a Live Dance Performance

T0 review · 5 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper argues that audiences grant more artistic merit to generative-AI visuals when they learn only after the performance that AI was used, and that this disclosure timing effect should shape how AI art is presented and explained.

desk verdict A small but honest case study with a clean 2x2 design; the headline effect is plausible but rests on one uncorrected test in non-randomized groups. read the letter →

arxiv 2508.00239 v1 pith:KGN6N66R submitted 2025-08-01 cs.HC cs.AI

classification cs.HCcs.AI
keywords generativeAIartaudienceperceptiondisclosuretimingartisticmeritlivedanceperformanceexplainablehuman-computerinteraction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that audiences' knowledge of generative AI's role changes how they value an artwork, independent of the artwork itself. It compares four live dance performances built from the same choreography: the visuals and sound were driven either by generative AI or by traditional digital tools, and viewers were told about the technology either before the show or after filling out the survey. The reported difference is that viewers who learned about the AI after the survey rated the projected visuals' artistic merit higher than those told beforehand (M=4.11 vs M=2.83), while the told-before group was more likely to call the visuals random. The paper takes this as evidence that disclosure timing shapes aesthetic judgment, and argues that explaining AI in the arts should address AI's presence and social context, not just its internal mechanics.

What carries the argument

The load-bearing device is a 2x2 between-subjects disclosure design: two versions of a twelve-minute dance performance, identical choreography and structure, differ only in whether generative AI or a technologist made the creative visual and sound mapping decisions, and audiences learn which version they saw either before the performance or after the survey. The measurement instrument is a 29-item Likert-scale survey probing whether the projected visuals and sound seemed creative, meaningful, random, distracting, or artistically meritorious, with pairwise group differences analyzed by Mann-Whitney tests. The design's power is that it holds the performed artwork constant and varies only the viewer's knowledge of the technology's role, so any difference between tell-before and tell-after groups is attributable, in the paper's logic, to that knowledge.

What would settle it

A preregistered replication with random assignment to disclosure timing and a single pre-specified measure of artistic merit—analyzed with a correction for multiple comparisons—would settle the claim: if the tell-after group still outrates the tell-before group, disclosure timing is the active ingredient; if the gap vanishes, group composition or chance was responsible.

Watch

Extended reading notes

Core claim

The study's central claim is that the same generative-AI visuals are judged differently depending on when viewers are told AI was involved. In the AI condition, participants who were informed after completing the survey agreed more strongly that 'The projected visuals demonstrated artistic merit' than participants informed before the performance, and the before group agreed more strongly that the visuals 'appeared random.' The paper interprets these differences as a bias against known AI involvement: once viewers know a machine contributed, they search for intention differently and grant less artistic credit. It also reports that among told-after audiences, the AI performance triggered more curiosity about how the visuals were made, while the non-AI version was rated higher on the sound complementing the performance. This is presented as a case study, not a general law, and the paper calls on explainable-AI work to include the viewer's prior knowledge as part of the explanation.

Load-bearing premise

The central claim assumes the rating gap between the AI/Tell-Before and AI/Tell-After groups was caused by when disclosure happened, rather than by their pre-existing differences—the groups attended different performances, were not randomly assigned, differed in mean age, and 29 survey items were tested without a multiple-comparison correction.

Editorial extensions

If this is right

  • Artists and curators who disclose AI involvement before a performance should expect lower artistic-merit ratings from audiences than if the same information comes after the experience, all else equal.
  • User studies of AI art that announce AI use upfront may systematically understate the aesthetic value audiences would otherwise assign to the work.
  • Explainable-AI practice in the arts should treat when and how AI's presence is revealed as part of the explanation, alongside any account of model mechanics.
  • The higher curiosity ratings in the AI/Tell-After group suggest that undisclosed AI can make audiences more inquisitive about the work's creation, not merely more approving.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Left implicit in the paper is a packaging dilemma: if disclosure after the experience raises artistic-merit ratings, artists and platforms face a transparency trade-off between honest upfront labeling and maximizing perceived value, and the field will need normative guidance on which should win.
  • The design mixes a large-language model and separately trained neural networks under one 'AI' label, so an obvious extension is to isolate whether the disclosure effect is driven by audience beliefs about AI as a category or by the specific generation mechanism used.
  • A testable extension would carry the same disclosure manipulation to static media such as images or music to see whether the effect is tied to live embodied performance or generalizes across art forms.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper reports a 2x2 between-subjects field study (N=39) in which audience members watched one of four live dance performances that differed in whether the visuals and sound were created with generative AI or with traditional digital tools, and in whether the technology was disclosed before the performance or after the survey. The central claim is that participants who were told about AI after completing the survey rated the AI-generated visuals as having greater artistic merit than those told before (M = 4.11 vs. 2.83, Z = -2.501, p = .012, Table 1). The authors interpret this as evidence that awareness of AI involvement lowers perceived artistic value and argue that explainable AI research should focus on transparency about AI's presence and capabilities rather than only on algorithmic mechanics.

Significance. The research question is timely and relevant to HCI, XAI, and arts practice. The study's main strength is that it uses real, professionally performed dance pieces with identical choreography, and the authors provide a detailed appendix of prompts, system architecture, and the survey instrument. If the finding were robust, it would have practical implications for artists and curators and theoretical implications for how disclosure timing shapes value judgments. However, the central result currently rests on a small, non-randomized sample and a single uncorrected statistical test among many, so the contribution is best viewed as an exploratory case study rather than a demonstration of the claimed effect.

major comments (5)
  1. [Section 3 / Table 1] The headline comparison (AI/Tell After vs. AI/Tell Before) is one of 29 Mann-Whitney tests on individual Likert items, but no multiple-comparison correction is applied. The key p-value of .012 would not survive a Bonferroni threshold of 0.05/29 = .0017. Under the global null, the expected number of p<.05 findings among 29 independent tests is about 1.45, and the probability of at least one is roughly 0.77, so this single significant result is not strong evidence against chance. Please report adjusted p-values (e.g., Holm-Bonferroni or Benjamini-Hochberg) or, preferably, test a pre-specified composite hypothesis about artistic merit.
  2. [Section 2.1 / Section 2.3] Participants were not randomly assigned to conditions; they attended one of four scheduled performances. The AI/Tell Before group differs from the AI/Tell After group in mean age (29.0 vs. 23.89) and group size (12 vs. 9). The observed difference in artistic merit could therefore be due to session effects, self-selection, or demographic differences rather than disclosure timing. The paper should report comparability of all groups on age, gender, and prior AI/art experience, or include age as a covariate in a regression/ANCOVA. Random assignment within each performance session would be the most direct fix.
  3. [Table 1] One row in Table 1 is corrupted: "Being informed about the production of a piece of art would impact its monetary valueThe projected visuals appeared random (Z = -2.698, p = .007...)" appears to concatenate two separate survey items. This raises doubts about the accuracy of the table as a whole. The authors should reconstruct the table directly from their analysis output and verify every row, and fix the malformed "p = <.001" entry.
  4. [Section 3 / Table 1] No effect sizes are reported for any comparison, despite very small group sizes (n = 12 and n = 9 for the key comparison). With samples this small, a one-point mean difference can be driven by one or two participants. Please report rank-biserial correlation or Cliff's delta with 95% confidence intervals, and consider a leave-one-out sensitivity analysis to assess the stability of the headline result.
  5. [Section 2.2 / Appendix B.2] The paper states that the Tell After condition added "six additional Likert scale questions," but the survey appendix lists several additional items that are not standard 5-point Likert scales: the two "Which of the following..." items use 1-to-5 anchors with qualitatively different endpoint labels, the "I noticed a mapping..." item has Yes/No/I don't recall responses, and one item is open-ended. Clarify exactly which six items were included in the Mann-Whitney analyses and how the non-Likert items were handled, since this affects the reproducibility of the statistical tests.
minor comments (6)
  1. [Abstract] The phrase "we uncovered the mixed opinions" overstates the contribution of a small exploratory survey; consider phrasing that matches the tentative nature of the evidence.
  2. [Section 2.3] The procedure says informed consent documents were distributed after the performance, which means audience members did not consent to participate before being exposed to the experimental manipulation. Please clarify the IRB-approved consent procedure, including any debriefing process.
  3. [Table 1] The Non-AI comparison row contains a stray "])." at the start of the entry for "I tried to make a connection..."; this typo should be corrected.
  4. [Appendix B.1] The "Used Responses for Scene 3" bullets appear to describe wave mappings ("amplitude of the waves," "wave's direction") rather than stars, and one bullet refers to "the third dance move's resemblance probability." Please verify that these are indeed the responses used for Scene 3 and correct any mislabeling.
  5. [References] Reference [17] contains a URL with spaces and line breaks; clean it up for the camera-ready version.
  6. [Section 3] The paper does not report internal consistency (e.g., Cronbach's alpha) for the survey items used as dependent variables, which is relevant given that items are analyzed individually and many appear to tap similar constructs.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: this is an empirical survey study whose headline claim is an observed group difference, not a quantity derived from its own inputs.

full rationale

The paper makes no derivation that reduces to its inputs. Its central claim is that audiences attributed more artistic merit to GenAI visuals when disclosure of AI use was withheld until after the survey, supported by a Mann-Whitney comparison of a Likert item between the AI/Tell Before and AI/Tell After groups. The survey questions, including 'The projected visuals demonstrated artistic merit,' are defined independently of the outcome, and the disclosure conditions were manipulated before data collection. There is no fitted parameter later relabeled as a prediction, no equation that defines the result in terms of itself, and no load-bearing appeal to the authors' own prior work. The study does have serious threats to causal validity, including non-random assignment, demographic differences between groups, and uncorrected multiple comparisons across 29 Likert items, but those are concerns about inference and chance, not circularity. Accordingly, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The paper is an empirical study and introduces no fitted numerical parameters or new theoretical entities. The central claim depends on the statistical validity of the Mann-Whitney tests as applied, which is captured in the single axiom listed.

assumptions (1)
  • ad hoc to paper The 29 Likert items are analyzed independently with uncorrected Mann-Whitney tests; the paper implicitly assumes the significant differences are not false positives from multiple testing.
    This assumption is specific to the paper's analysis choices in Section 3 and Table 1, and it directly affects the load-bearing claim about artistic merit.

how reviews work

0 comments
Cite this review

Pith. "Pith review of What's Behind the Magic? Audiences Seek Artistic Value in Generative AI's Contributions to a Live Dance Performance." pith.science (2026). https://pith.science/paper/KGN6N66R

@misc{pith2026250800239,
  author       = {Pith},
  title        = {Pith review of: What's Behind the Magic? Audiences Seek Artistic Value in Generative AI's Contributions to a Live Dance Performance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KGN6N66R}},
  note         = {Machine review of arXiv:2508.00239}
}
read the original abstract

With the development of generative artificial intelligence (GenAI) tools to create art, stakeholders cannot come to an agreement on the value of these works. In this study we uncovered the mixed opinions surrounding art made by AI. We developed two versions of a dance performance augmented by technology either with or without GenAI. For each version we informed audiences of the performance's development either before or after a survey on their perceptions of the performance. There were thirty-nine participants (13 males, 26 female) divided between the four performances. Results demonstrated that individuals were more inclined to attribute artistic merit to works made by GenAI when they were unaware of its use. We present this case study as a call to address the importance of utilizing the social context and the users' interpretations of GenAI in shaping a technical explanation, leading to a greater discussion that can bridge gaps in understanding.

Figures

Figures reproduced from arXiv: 2508.00239 by the authors.

Figure 1
Figure 1. System architecture of the Non-AI performance version. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. System architecture of the AI performance version. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. The dancer rehearsing during an experiment [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Outline of the choreography for Scene 2. [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

24 extracted references · 18 canonical work pages

  1. [1]

    Brian A Anderson, Patryk A Laurent, and Steven Yantis. 2011. Learned value magnifies salience-based attentional capture. PloS one 6, 11 (2011), e27926

  2. [2]

    David Bayles and Ted Orland. 2023. Art & fear: Observations on the perils (and rewards) of artmaking. Souvenir Press

  3. [3]

    Z Epstein, A Hertzmann, L Herman, R Mahari, MR Frank, M Groh, H Schroeder, A Smith, M Akten, J Fjeld, et al. 2023. Art and the science of generative AI: A deeper dive (arXiv: 2306.04141). arXiv

  4. [4]

    Frank, Matthew Groh, Laura Her- man, Neil Leach, Robert Mahari, Alex “Sandy” Pentland, Olga Russakovsky, Hope Schroeder, and Amy Smith

    Ziv Epstein, Aaron Hertzmann, the Investigators of Human Creativity, Memo Akten, Hany Farid, Jessica Fjeld, Morgan R. Frank, Matthew Groh, Laura Her- man, Neil Leach, Robert Mahari, Alex “Sandy” Pentland, Olga Russakovsky, Hope Schroeder, and Amy Smith. 2023. Art and the science of genera- tive AI. Science 380, 6650 (2023), 1110–1111. doi:10.1126/science....

  5. [5]

    Stefan Feuerriegel, Jochen Hartmann, Christian Janiesch, and Patrick Zschech

  6. [6]

    Rebecca Fiebrink and Perry R Cook. 2010. The Wekinator: a system for real-time, interactive machine learning in music. InProceedings of The Eleventh International Society for Music Information Retrieval Conference (ISMIR 2010)(Utrecht) , Vol. 3. Citeseer, 2–1

  7. [7]

    Fiona Fui-Hoon Nah, Ruilin Zheng, Jingyuan Cai, Keng Siau, and Langtao Chen

  8. [8]

    Danah Henriksen, Nicole Oster, Punya Mishra, and Lindsey McCaleb. 2024. Gen- erative AI, Creativity, Culture, and the Future of Learning: a Conversation with Mairéad Pratschke. TechTrends (2024), 1–7

Show all 24 references
  1. [9]

    IBM. [n. d.]. What is generative AI? https://www.ibm.com/think/topics/ generative-ai. Accessed: 1-4-2025

  2. [10]

    Jiang, Lauren Brown, Jessica Cheng, Mehtab Khan, Abhishek Gupta, Deja Workman, Alex Hanna, Johnathan Flowers, and Timnit Gebru

    Harry H. Jiang, Lauren Brown, Jessica Cheng, Mehtab Khan, Abhishek Gupta, Deja Workman, Alex Hanna, Johnathan Flowers, and Timnit Gebru. 2023. AI Art and its Impact on Artists. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (Montréal, QC, Canada) (AI...

  3. [11]

    Caroline A Jones, Huma Gupta, and Matthew Ritchie. 2024. Visual artists, tech- nological shock, and generative AI. (2024)

  4. [12]

    Kostas Karpouzis. 2024. Plato’s Shadows in the Digital Cave: Controlling Cultural Bias in Generative AI. Electronics 13, 8 (2024), 1457

  5. [13]

    Jeong Hyun Kim, Jungkeun Kim, Jooyoung Park, Changju Kim, Jihoon Jhang, and Brian King. 2025. When ChatGPT gives incorrect answers: the impact of inaccurate information by generative AI on tourism decision-making. Journal of Travel Research 64, 1 (2025), 51–73

  6. [14]

    Rensis Likert. 1932. A technique for the measurement of attitudes. Archives of Psychology (1932)

  7. [15]

    Vivian Liu and Lydia B Chilton. 2022. Design guidelines for prompt engineering text-to-image generative models. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–23

  8. [16]

    Henry B Mann and Donald R Whitney. 1947. On a test of whether one of two random variables is stochastically larger than the other. The annals of mathematical statistics (1947), 50–60

  9. [17]

    L Mineo. 2023. If it wasn’t created by a human artist, is it still art. The Harvard Gazette. Artikkeli. Luettavissa: https://news. harvard. edu/gazette/story/2023/08/is- art-generated-byartificial-intelligence-real-art (2023)

  10. [18]

    OpenAI. 2024. ChatGPT (May 13 version) [Large language model]. https://chat. openai.com. Accessed: 2025-05-07

  11. [19]

    Jonas Oppenlaender, Rhema Linder, and Johanna Silvennoinen. 2024. Prompting AI art: An investigation into the creative skill of prompt engineering.International Journal of Human–Computer Interaction (2024), 1–23

  12. [20]

    James Prather, Brent N Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S Randri- anasolo, Brett A Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The widening gap: The benefits and harms of generative ai for novice programmers. In Proceedings of the 2024 ACM Conferen...

  13. [21]

    Dirk HR Spennemann. 2024. Generative artificial intelligence, human agency and the future of cultural heritage. Heritage 7, 7 (2024), 3597

  14. [22]

    If I had to make petals of a painting move according to the respiratory rate of a dancer how should I map this respiratory rate to the petals?

    Frank Wilcoxon. 1992. Individual comparisons by ranking methods. In Break- throughs in statistics: Methodology and distribution . Springer, 196–202. A Performance Development Imagery Figure 1. System architecture of the Non-AI performance version. XAIxArts 2025, June 23, 2025,...

  15. [2023]

    277–304 pages

    Generative AI and ChatGPT: Applications, challenges, and AI-human collaboration. 277–304 pages

  16. [2024]

    Business & Information Systems Engineering 66, 1 (2024), 111–126

    Generative ai. Business & Information Systems Engineering 66, 1 (2024), 111–126

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.