REVIEW 5 major objections 6 minor 24 references
What's Behind the Magic? Audiences Seek Artistic Value in Generative AI's Contributions to a Live Dance Performance
T0 review · 5 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper argues that audiences grant more artistic merit to generative-AI visuals when they learn only after the performance that AI was used, and that this disclosure timing effect should shape how AI art is presented and explained.
desk verdict A small but honest case study with a clean 2x2 design; the headline effect is plausible but rests on one uncorrected test in non-randomized groups. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is a 2x2 between-subjects disclosure design: two versions of a twelve-minute dance performance, identical choreography and structure, differ only in whether generative AI or a technologist made the creative visual and sound mapping decisions, and audiences learn which version they saw either before the performance or after the survey. The measurement instrument is a 29-item Likert-scale survey probing whether the projected visuals and sound seemed creative, meaningful, random, distracting, or artistically meritorious, with pairwise group differences analyzed by Mann-Whitney tests. The design's power is that it holds the performed artwork constant and varies only the viewer's knowledge of the technology's role, so any difference between tell-before and tell-after groups is attributable, in the paper's logic, to that knowledge.
What would settle it
A preregistered replication with random assignment to disclosure timing and a single pre-specified measure of artistic merit—analyzed with a correction for multiple comparisons—would settle the claim: if the tell-after group still outrates the tell-before group, disclosure timing is the active ingredient; if the gap vanishes, group composition or chance was responsible.
Extended reading notes
Core claim
The study's central claim is that the same generative-AI visuals are judged differently depending on when viewers are told AI was involved. In the AI condition, participants who were informed after completing the survey agreed more strongly that 'The projected visuals demonstrated artistic merit' than participants informed before the performance, and the before group agreed more strongly that the visuals 'appeared random.' The paper interprets these differences as a bias against known AI involvement: once viewers know a machine contributed, they search for intention differently and grant less artistic credit. It also reports that among told-after audiences, the AI performance triggered more curiosity about how the visuals were made, while the non-AI version was rated higher on the sound complementing the performance. This is presented as a case study, not a general law, and the paper calls on explainable-AI work to include the viewer's prior knowledge as part of the explanation.
Load-bearing premise
The central claim assumes the rating gap between the AI/Tell-Before and AI/Tell-After groups was caused by when disclosure happened, rather than by their pre-existing differences—the groups attended different performances, were not randomly assigned, differed in mean age, and 29 survey items were tested without a multiple-comparison correction.
Editorial extensions
If this is right
- Artists and curators who disclose AI involvement before a performance should expect lower artistic-merit ratings from audiences than if the same information comes after the experience, all else equal.
- User studies of AI art that announce AI use upfront may systematically understate the aesthetic value audiences would otherwise assign to the work.
- Explainable-AI practice in the arts should treat when and how AI's presence is revealed as part of the explanation, alongside any account of model mechanics.
- The higher curiosity ratings in the AI/Tell-After group suggest that undisclosed AI can make audiences more inquisitive about the work's creation, not merely more approving.
Reading between the lines
- Left implicit in the paper is a packaging dilemma: if disclosure after the experience raises artistic-merit ratings, artists and platforms face a transparency trade-off between honest upfront labeling and maximizing perceived value, and the field will need normative guidance on which should win.
- The design mixes a large-language model and separately trained neural networks under one 'AI' label, so an obvious extension is to isolate whether the disclosure effect is driven by audience beliefs about AI as a category or by the specific generation mechanism used.
- A testable extension would carry the same disclosure manipulation to static media such as images or music to see whether the effect is tied to live embodied performance or generalizes across art forms.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a 2x2 between-subjects field study (N=39) in which audience members watched one of four live dance performances that differed in whether the visuals and sound were created with generative AI or with traditional digital tools, and in whether the technology was disclosed before the performance or after the survey. The central claim is that participants who were told about AI after completing the survey rated the AI-generated visuals as having greater artistic merit than those told before (M = 4.11 vs. 2.83, Z = -2.501, p = .012, Table 1). The authors interpret this as evidence that awareness of AI involvement lowers perceived artistic value and argue that explainable AI research should focus on transparency about AI's presence and capabilities rather than only on algorithmic mechanics.
Significance. The research question is timely and relevant to HCI, XAI, and arts practice. The study's main strength is that it uses real, professionally performed dance pieces with identical choreography, and the authors provide a detailed appendix of prompts, system architecture, and the survey instrument. If the finding were robust, it would have practical implications for artists and curators and theoretical implications for how disclosure timing shapes value judgments. However, the central result currently rests on a small, non-randomized sample and a single uncorrected statistical test among many, so the contribution is best viewed as an exploratory case study rather than a demonstration of the claimed effect.
major comments (5)
- [Section 3 / Table 1] The headline comparison (AI/Tell After vs. AI/Tell Before) is one of 29 Mann-Whitney tests on individual Likert items, but no multiple-comparison correction is applied. The key p-value of .012 would not survive a Bonferroni threshold of 0.05/29 = .0017. Under the global null, the expected number of p<.05 findings among 29 independent tests is about 1.45, and the probability of at least one is roughly 0.77, so this single significant result is not strong evidence against chance. Please report adjusted p-values (e.g., Holm-Bonferroni or Benjamini-Hochberg) or, preferably, test a pre-specified composite hypothesis about artistic merit.
- [Section 2.1 / Section 2.3] Participants were not randomly assigned to conditions; they attended one of four scheduled performances. The AI/Tell Before group differs from the AI/Tell After group in mean age (29.0 vs. 23.89) and group size (12 vs. 9). The observed difference in artistic merit could therefore be due to session effects, self-selection, or demographic differences rather than disclosure timing. The paper should report comparability of all groups on age, gender, and prior AI/art experience, or include age as a covariate in a regression/ANCOVA. Random assignment within each performance session would be the most direct fix.
- [Table 1] One row in Table 1 is corrupted: "Being informed about the production of a piece of art would impact its monetary valueThe projected visuals appeared random (Z = -2.698, p = .007...)" appears to concatenate two separate survey items. This raises doubts about the accuracy of the table as a whole. The authors should reconstruct the table directly from their analysis output and verify every row, and fix the malformed "p = <.001" entry.
- [Section 3 / Table 1] No effect sizes are reported for any comparison, despite very small group sizes (n = 12 and n = 9 for the key comparison). With samples this small, a one-point mean difference can be driven by one or two participants. Please report rank-biserial correlation or Cliff's delta with 95% confidence intervals, and consider a leave-one-out sensitivity analysis to assess the stability of the headline result.
- [Section 2.2 / Appendix B.2] The paper states that the Tell After condition added "six additional Likert scale questions," but the survey appendix lists several additional items that are not standard 5-point Likert scales: the two "Which of the following..." items use 1-to-5 anchors with qualitatively different endpoint labels, the "I noticed a mapping..." item has Yes/No/I don't recall responses, and one item is open-ended. Clarify exactly which six items were included in the Mann-Whitney analyses and how the non-Likert items were handled, since this affects the reproducibility of the statistical tests.
minor comments (6)
- [Abstract] The phrase "we uncovered the mixed opinions" overstates the contribution of a small exploratory survey; consider phrasing that matches the tentative nature of the evidence.
- [Section 2.3] The procedure says informed consent documents were distributed after the performance, which means audience members did not consent to participate before being exposed to the experimental manipulation. Please clarify the IRB-approved consent procedure, including any debriefing process.
- [Table 1] The Non-AI comparison row contains a stray "])." at the start of the entry for "I tried to make a connection..."; this typo should be corrected.
- [Appendix B.1] The "Used Responses for Scene 3" bullets appear to describe wave mappings ("amplitude of the waves," "wave's direction") rather than stars, and one bullet refers to "the third dance move's resemblance probability." Please verify that these are indeed the responses used for Scene 3 and correct any mislabeling.
- [References] Reference [17] contains a URL with spaces and line breaks; clean it up for the camera-ready version.
- [Section 3] The paper does not report internal consistency (e.g., Cronbach's alpha) for the survey items used as dependent variables, which is relevant given that items are analyzed individually and many appear to tap similar constructs.
Circularity Check
No circularity: this is an empirical survey study whose headline claim is an observed group difference, not a quantity derived from its own inputs.
full rationale
The paper makes no derivation that reduces to its inputs. Its central claim is that audiences attributed more artistic merit to GenAI visuals when disclosure of AI use was withheld until after the survey, supported by a Mann-Whitney comparison of a Likert item between the AI/Tell Before and AI/Tell After groups. The survey questions, including 'The projected visuals demonstrated artistic merit,' are defined independently of the outcome, and the disclosure conditions were manipulated before data collection. There is no fitted parameter later relabeled as a prediction, no equation that defines the result in terms of itself, and no load-bearing appeal to the authors' own prior work. The study does have serious threats to causal validity, including non-random assignment, demographic differences between groups, and uncorrected multiple comparisons across 29 Likert items, but those are concerns about inference and chance, not circularity. Accordingly, the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (1)
- ad hoc to paper The 29 Likert items are analyzed independently with uncorrected Mann-Whitney tests; the paper implicitly assumes the significant differences are not false positives from multiple testing.
Cite this review
Pith. "Pith review of What's Behind the Magic? Audiences Seek Artistic Value in Generative AI's Contributions to a Live Dance Performance." pith.science (2026). https://pith.science/paper/KGN6N66R
@misc{pith2026250800239,
author = {Pith},
title = {Pith review of: What's Behind the Magic? Audiences Seek Artistic Value in Generative AI's Contributions to a Live Dance Performance},
year = {2026},
howpublished = {\url{https://pith.science/paper/KGN6N66R}},
note = {Machine review of arXiv:2508.00239}
}
read the original abstract
With the development of generative artificial intelligence (GenAI) tools to create art, stakeholders cannot come to an agreement on the value of these works. In this study we uncovered the mixed opinions surrounding art made by AI. We developed two versions of a dance performance augmented by technology either with or without GenAI. For each version we informed audiences of the performance's development either before or after a survey on their perceptions of the performance. There were thirty-nine participants (13 males, 26 female) divided between the four performances. Results demonstrated that individuals were more inclined to attribute artistic merit to works made by GenAI when they were unaware of its use. We present this case study as a call to address the importance of utilizing the social context and the users' interpretations of GenAI in shaping a technical explanation, leading to a greater discussion that can bridge gaps in understanding.
Figures
Reference graph
Works this paper leans on
-
[1]
Brian A Anderson, Patryk A Laurent, and Steven Yantis. 2011. Learned value magnifies salience-based attentional capture. PloS one 6, 11 (2011), e27926
work page 2011
-
[2]
David Bayles and Ted Orland. 2023. Art & fear: Observations on the perils (and rewards) of artmaking. Souvenir Press
work page 2023
-
[3]
Z Epstein, A Hertzmann, L Herman, R Mahari, MR Frank, M Groh, H Schroeder, A Smith, M Akten, J Fjeld, et al. 2023. Art and the science of generative AI: A deeper dive (arXiv: 2306.04141). arXiv
work page Pith review arXiv 2023
-
[4]
Ziv Epstein, Aaron Hertzmann, the Investigators of Human Creativity, Memo Akten, Hany Farid, Jessica Fjeld, Morgan R. Frank, Matthew Groh, Laura Her- man, Neil Leach, Robert Mahari, Alex “Sandy” Pentland, Olga Russakovsky, Hope Schroeder, and Amy Smith. 2023. Art and the science of genera- tive AI. Science 380, 6650 (2023), 1110–1111. doi:10.1126/science....
-
[5]
Stefan Feuerriegel, Jochen Hartmann, Christian Janiesch, and Patrick Zschech
-
[6]
Rebecca Fiebrink and Perry R Cook. 2010. The Wekinator: a system for real-time, interactive machine learning in music. InProceedings of The Eleventh International Society for Music Information Retrieval Conference (ISMIR 2010)(Utrecht) , Vol. 3. Citeseer, 2–1
work page 2010
-
[7]
Fiona Fui-Hoon Nah, Ruilin Zheng, Jingyuan Cai, Keng Siau, and Langtao Chen
-
[8]
Danah Henriksen, Nicole Oster, Punya Mishra, and Lindsey McCaleb. 2024. Gen- erative AI, Creativity, Culture, and the Future of Learning: a Conversation with Mairéad Pratschke. TechTrends (2024), 1–7
work page 2024
Show all 24 references
-
[9]
IBM. [n. d.]. What is generative AI? https://www.ibm.com/think/topics/ generative-ai. Accessed: 1-4-2025
2025
-
[10]
Jiang, Lauren Brown, Jessica Cheng, Mehtab Khan, Abhishek Gupta, Deja Workman, Alex Hanna, Johnathan Flowers, and Timnit Gebru
Harry H. Jiang, Lauren Brown, Jessica Cheng, Mehtab Khan, Abhishek Gupta, Deja Workman, Alex Hanna, Johnathan Flowers, and Timnit Gebru. 2023. AI Art and its Impact on Artists. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (Montréal, QC, Canada) (AI...
2023
-
[11]
Caroline A Jones, Huma Gupta, and Matthew Ritchie. 2024. Visual artists, tech- nological shock, and generative AI. (2024)
2024
-
[12]
Kostas Karpouzis. 2024. Plato’s Shadows in the Digital Cave: Controlling Cultural Bias in Generative AI. Electronics 13, 8 (2024), 1457
2024
-
[13]
Jeong Hyun Kim, Jungkeun Kim, Jooyoung Park, Changju Kim, Jihoon Jhang, and Brian King. 2025. When ChatGPT gives incorrect answers: the impact of inaccurate information by generative AI on tourism decision-making. Journal of Travel Research 64, 1 (2025), 51–73
2025
-
[14]
Rensis Likert. 1932. A technique for the measurement of attitudes. Archives of Psychology (1932)
1932
-
[15]
Vivian Liu and Lydia B Chilton. 2022. Design guidelines for prompt engineering text-to-image generative models. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–23
2022
-
[16]
Henry B Mann and Donald R Whitney. 1947. On a test of whether one of two random variables is stochastically larger than the other. The annals of mathematical statistics (1947), 50–60
1947
-
[17]
L Mineo. 2023. If it wasn’t created by a human artist, is it still art. The Harvard Gazette. Artikkeli. Luettavissa: https://news. harvard. edu/gazette/story/2023/08/is- art-generated-byartificial-intelligence-real-art (2023)
2023
-
[18]
OpenAI. 2024. ChatGPT (May 13 version) [Large language model]. https://chat. openai.com. Accessed: 2025-05-07
2024
-
[19]
Jonas Oppenlaender, Rhema Linder, and Johanna Silvennoinen. 2024. Prompting AI art: An investigation into the creative skill of prompt engineering.International Journal of Human–Computer Interaction (2024), 1–23
2024
-
[20]
James Prather, Brent N Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S Randri- anasolo, Brett A Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The widening gap: The benefits and harms of generative ai for novice programmers. In Proceedings of the 2024 ACM Conferen...
2024
-
[21]
Dirk HR Spennemann. 2024. Generative artificial intelligence, human agency and the future of cultural heritage. Heritage 7, 7 (2024), 3597
2024
-
[22]
If I had to make petals of a painting move according to the respiratory rate of a dancer how should I map this respiratory rate to the petals?
Frank Wilcoxon. 1992. Individual comparisons by ranking methods. In Break- throughs in statistics: Methodology and distribution . Springer, 196–202. A Performance Development Imagery Figure 1. System architecture of the Non-AI performance version. XAIxArts 2025, June 23, 2025,...
1992
-
[2023]
277–304 pages
Generative AI and ChatGPT: Applications, challenges, and AI-human collaboration. 277–304 pages
-
[2024]
Business & Information Systems Engineering 66, 1 (2024), 111–126
Generative ai. Business & Information Systems Engineering 66, 1 (2024), 111–126
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.