REVIEW 3 major objections 5 minor 44 references
From Creation to Curriculum: Examining the role of generative AI in Arts Universities
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Arts universities should make generative AI image tools a core part of their curricula, this paper argues.
desk verdict A candid practitioner case report on teaching Stable Diffusion to arts students; the workshop material is genuinely useful, but the urgent 'act now' recommendation rests on an unsupported five-year industry forecast. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the open-source Stable Diffusion pipeline as operated through a browser-based interface on cloud notebooks, with checkpoint models, seed, steps, samplers, weighted prompts, and negative prompts as the primary controls. The two extensions that do the heaviest lifting are LoRA, a lightweight fine-tuning method that adds style or subject-specific weights to an existing model, and ControlNet, a conditioning network that lets artists steer composition through edges, depth, or pose. The pedagogical machinery is constructionism—learning by making—mapped onto photography's rapid feedback loop: generate many images, review them as contact sheets, refine, and exhibit a selected piece. That pairing of a controllable open-source toolchain with project-based, exhibition-driven instruction is what carries the claim that AI image generation is teachable and should enter the curriculum.
What would settle it
A systematic annual review of entry-level creative-industry job postings over the next five years would settle the forecast: if by 2028 generative-AI skills are not widely listed as required or preferred in those postings, or if copyright rulings have removed major text-to-image tools from common use, the paper's five-year transformation claim is contradicted.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the complete current text-to-image workflow—checkpoint models, seed values, steps, samplers, weighted prompts, LoRA fine-tuning, ControlNet conditioning, and upscaling—can be taught to beginners with no technical background, and that the learning follows a constructionist pattern of iterative making, contact-sheet review, and exhibition-driven refinement. Three student case studies carry this claim: a printmaking graduate student trained a style-specific LoRA from her own silk-screen prints; a fourth-year student trained a LoRA on Buddhist and Shinto imagery for a series about machine spirituality; and an architecture student combined 200 generated images into one large composite. The paper further claims that this teachability, together with an expected full adoption of these tools by creative industries within four to five years, makes immediate curricular integration a responsibility of arts universities.
Load-bearing premise
The load-bearing premise is the unverified forecast that creative industries will fully adopt generative AI by the time current freshmen graduate and be completely transformed within five years, which is what makes immediate curriculum change seem urgent.
Editorial extensions
If this is right
- Arts universities can run serious AI-art coursework with only cloud-hosted notebooks, shared storage, and open-source software, so students are not required to own expensive GPUs.
- Students who complete a constructionist AI-art module should be able to calibrate prompts, train or apply LoRAs, condition compositions with ControlNet, and produce exhibition-quality large-format prints.
- The teacher's role shifts from transmitting fixed knowledge to coaching, curating feedback, and providing technical scaffolding, because the open-source community supplies documentation and support.
- If the paper's industry forecast is right, graduates without generative-AI literacy will enter a shrinking entry-level creative job market already reshaped by these tools.
- Teaching through open-source tools keeps the curriculum aligned with the artist-developer-researcher loop that continues to drive the technology forward.
Reading between the lines
- Beyond the paper, the same constructionist, exhibition-driven format could transfer to generative music, video, 3D, and design courses, since the underlying pattern of iterative generation plus curation is medium-independent.
- A testable extension of the paper's pedagogy would be a controlled comparison of a constructionist AI-art module against a tutorial-driven module, measuring portfolio quality, tool fluency, and transfer to unseen models; the paper's feasibility claim predicts the constructionist group will do at least as well.
- Because the paper's urgency argument rests on an empirical forecast, a reader can treat that forecast as a testable hypothesis: if adoption takes longer than five years, the case for acting now weakens even though teaching the tools may still be justified on creative-education grounds alone.
- The paper's use of the x/y/z plot script as an automated contact sheet suggests a concrete assessment artifact—students submit a grid of parameter variations plus a curated final work—that other institutions could adopt without new infrastructure.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that generative AI, particularly Stable Diffusion, is a transformative technology with profound implications for arts education. It provides a technical overview of text-to-image models and the open-source ecosystem, reports on two workshops and a follow-up exhibition at Kyoto Seika University, interprets these experiences through Papert's constructionism, and concludes with an urgent recommendation that university administrators and teachers immediately integrate generative AI into curricula. The central normative claim is that universities must act now because creative industries will be 'completely transformed' by the time current freshmen graduate.
Significance. The paper's practical materials, such as the workshop design, the discussion of LoRA and ControlNet workflows, and the open-source tooling, are genuinely useful for educators considering similar courses. The constructionist framing provides a credible pedagogical lens for hands-on AI art instruction, and the explicit focus on open-source software is a strength in an area often dominated by proprietary services. However, the paper's broader significance is limited by its thin empirical basis: the workshop observations are anecdotal, the survey data were too few to report, and the urgent policy recommendation rests on an unsupported industry-transformation forecast. As a practitioner report it is informative, but as a research contribution it needs substantial reframing and support.
major comments (3)
- [Section 4, paragraph 4] The recommendation to 'take action now' rests squarely on the empirical prediction that 'By time this year's freshmen graduate, creative industries will have fully adopted these technologies. In 5 years time, these industries be completely transformed.' This is a load-bearing assertion, yet no labor-market, industry, or technology-adoption evidence is cited, and the terms 'fully adopted' and 'completely transformed' are undefined, making the forecast unfalsifiable. The prediction is not derived from the Section 3 case studies, which concern a small, self-selected group of students. If adoption is slower, partial, or checked by copyright rulings, the urgency claim collapses, even though teaching AI tools could still be justified on constructionist or digital-literacy grounds. The authors should either support the timeline with concrete data (e.g., industry reports, hiring statistics, adoption curves) or revise the argument to a conditional recommendation that does not depend on an unverified five-year horizon.
- [Section 3.3, 'Results of the workshops' and Section 3.5] The paper acknowledges that the post-workshop survey had too few respondents to present quantitative data, but it then proceeds to draw general conclusions from qualitative observations, informal discussions, and only three exhibition participants. The sample is self-selected (47 enrolled, 3 exhibited), and the four student profiles are not representative of a typical university cohort. The statement that students 'generally had no problem using the software' and the positive-engagement observations are offered without a systematic coding scheme, comparison group, or pre/post assessment. Moreover, the constructionist interpretation in Section 3.5 is applied to data generated within the author's own advocacy framework, making the case study self-confirming: students were recruited, trained, supported, and exhibited, and their participation is then read as evidence that AI should be adopted. This circularity does not invalidate the workshop account, but it means the paper cannot support its broad curricular recommendation. The authors should reframe the Section 3 material as an exploratory case study and explicitly discuss its limitations before making general claims about effective curriculum integration.
- [Section 2.2.1 and 2.2.2] The technical history contains inaccuracies that, while not central to the recommendation, undermine the credibility of the overview. The claim that 'the first notable attempt at text-to-image synthesis was in 2014 by a team of researchers from the University of Montreal, who proposed a model called Generative Adversarial Networks (GANs)' is contradicted two sentences later by the admission that 'GANs are not technically considered TTI models.' Earlier text-to-image approaches existed before 2014, and the 2014 GAN paper was not a text-to-image model at all. Similarly, the statement that 'OpenAI introduced a new approach to text-to-image synthesis using the diffusion model which was first introduced in 2015' conflates the 2015 diffusion-model paper with OpenAI's later application of diffusion to text-to-image; the direct precedence for Stable Diffusion is the latent diffusion work of Rombach et al. (2021). These errors should be corrected to provide an accurate foundation for the non-technical readers the paper targets.
minor comments (5)
- [Figure 5] The caption reads 'Enter Caption', which appears to be a placeholder that must be replaced with an actual descriptive caption.
- [Section 2.2.2] The text refers to 'Fig X below' before Figure 2; this should be updated to 'Figure 2'.
- [Throughout] There are several typos and inconsistent spellings, including 'curriculua' (§1), 'Stable Diffuion' (§2), 'ControNet' (§2.4.2), and 'Diffusion' in the section title 'Diffusion Models' (§2.2.2). A careful proofread is needed.
- [Section 4] The claim that 'GPT-4 was trained on the output of ChatGPT' is stated without a citation or source. If this is speculative, it should be labeled as such; if it is reported information, a reference should be provided.
- [Section 2.2.4] The sentence 'The original latent diffusion model was further developed and trained on the LAION-5B image dataset [25] with the support of Stability.AI' could be clarified: Stability.AI provided training resources rather than developing the model itself, and the lineage is better described as a collaboration between CompVis, the LMU group, and Stability.AI.
Circularity Check
No significant circularity: the paper's technical overview and workshop case studies are self-contained, and the unsupported five-year adoption forecast is an evidence weakness, not a derivation that reduces to its own inputs.
full rationale
The paper's argument chain is: (1) an external technical overview of Stable Diffusion and the open-source ecosystem, (2) a report of July 2023 workshops and an exhibition with qualitative observations, and (3) a recommendation that arts universities adopt generative AI tools. The technical section relies on external sources such as Rombach et al., ControlNet, LoRA, and OpenAI's CLIP release, and it does not depend on any claim defined in terms of the paper's own conclusion. The workshop evidence is primary qualitative data reported by the author; no parameter is fitted to any subset and then repackaged as a prediction. The constructionist analysis applies Papert's framework to the observed learning process, but the paper does not define its recommendation in terms of that framework, nor does it present the framework's vocabulary as the evidence for the recommendation. The strongest rhetorical premise, stated in Section 4, is that 'By time this year's freshmen graduate, creative industries will have fully adopted these technologies. In 5 years time, these industries be completely transformed...' This is an unsupported empirical forecast with undefined terms like 'fully adopted' and 'completely transformed,' and it is not derived from the workshop data. That is a correctness and evidence limitation, not circularity: the forecast is an independent premise, not a restatement of the recommendation or a fitted parameter. The paper also explicitly admits that survey responses were too few for quantitative data, which further demonstrates that it is not disguising fitted inputs as predictions. There is no self-citation chain, no imported uniqueness theorem, and no ansatz smuggled in via citation. Incomplete elements such as Figure 5's 'Enter Caption' and the placeholder '[cite multiple references]' are manuscript completeness issues, not circular reasoning. Overall, the paper's central pedagogical claim is self-contained against the evidence it reports, and the unsupported adoption timeline should be treated as an evidentiary gap rather than a circular step.
Assumptions & free parameters
assumptions (3)
- domain assumption The AI revolution is one of the most significant changes in human history, comparable to the printing press or the wheel.
- domain assumption Creative industries will be fully transformed within five years, with far fewer new positions.
- domain assumption Constructionism is the ideal pedagogical framework for teaching generative AI art.
Cite this review
Pith. "Pith review of From Creation to Curriculum: Examining the role of generative AI in Arts Universities." pith.science (2026). https://pith.science/paper/URKA7R5C
@misc{pith2026241216531,
author = {Pith},
title = {Pith review of: From Creation to Curriculum: Examining the role of generative AI in Arts Universities},
year = {2026},
howpublished = {\url{https://pith.science/paper/URKA7R5C}},
note = {Machine review of arXiv:2412.16531}
}
read the original abstract
The age of Artificial Intelligence (AI) is marked by its transformative "generative" capabilities, distinguishing it from prior iterations. This burgeoning characteristic of AI has enabled it to produce new and original content, inherently showcasing its creative prowess. This shift challenges and requires a recalibration in the realm of arts education, urging a departure from established pedagogies centered on human-driven image creation. The paper meticulously addresses the integration of AI tools, with a spotlight on Stable Diffusion (SD), into university arts curricula. Drawing from practical insights gathered from workshops conducted in July 2023, which culminated in an exhibition of AI-driven artworks, the paper aims to provide a roadmap for seamlessly infusing these tools into academic settings. Given their recent emergence, the paper delves into a comprehensive overview of such tools, emphasizing the intricate dance between artists, developers, and researchers in the open-source AI art world. This discourse extends to the challenges and imperatives faced by educational institutions. It presents a compelling case for the swift adoption of these avant-garde tools, underscoring the paramount importance of equipping students with the competencies required to thrive in an AI-augmented artistic landscape.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Foster, Generative deep learning
D. Foster, Generative deep learning. ” O’Reilly Media, Inc.”, 2022
work page 2022
-
[2]
Generative adversarial networks,
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, “Generative adversarial networks,”Communications of the ACM, vol. 63, no. 11, pp. 139–144, 2020
2020
-
[3]
Unsupervised representation learning with deep convolutional generative adversarial networks,
A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” arXiv preprint arXiv:1511.06434, 2015
arXiv 2015
-
[4]
Attngan: Fine-grained text to image generation with attentional generative adversarial networks,
T. Xu, P. Zhang, Q. Huang, H. Zhang, Z. Gan, X. Huang, and X. He, “Attngan: Fine-grained text to image generation with attentional generative adversarial networks,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1316–1324, 2017. 15 From Creation to Curriculum: Examining the role of generative AI in Arts Universities
work page 2018
-
[5]
L. P. Cinelli, M. A. Marins, E. A. B. Da Silva, and S. L. Netto, Variational Methods for Machine Learning with Applications to Deep Networks. Springer, 2021
work page 2021
-
[6]
Auto-encoding variational bayes,
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” CoRR, vol. abs/1312.6114, 2013
arXiv 2013
-
[7]
beta- vae: Learning basic visual concepts with a constrained variational framework,
I. Higgins, L. Matthey, A. Pal, C. P. Burgess, X. Glorot, M. M. Botvinick, S. Mohamed, and A. Lerchner, “beta- vae: Learning basic visual concepts with a constrained variational framework,” in International Conference on Learning Representations, 2016
work page 2016
- [8]
Show all 44 references
-
[9]
Deep unsupervised learning using nonequilibrium thermodynamics,
J. N. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” ArXiv, vol. abs/1503.03585, 2015
2015 arXiv
-
[10]
Diffusion models in vision: A survey,
F.-A. Croitoru, V . Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” ArXiv, vol. abs/2209.04747, 2022
2022 arXiv
-
[11]
Diffusion models: A comprehensive survey of methods and applications,
L. Yang, Z. Zhang, S. Hong, R. Xu, Y . Zhao, Y . Shao, W. Zhang, M.-H. Yang, and B. Cui, “Diffusion models: A comprehensive survey of methods and applications,”ArXiv, vol. abs/2209.00796, 2022
2022
-
[12]
Pareidolia - wikipedia
“Pareidolia - wikipedia.”
-
[13]
Diffusion models beat gans on image synthesis,
P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” ArXiv, vol. abs/2105.05233, 2021
2021 arXiv
-
[14]
High-resolution image synthesis with latent dif- fusion models,
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent dif- fusion models,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10674– 10685, 2021
2022
-
[15]
Fixed forward diffusion process,
N. D. Blog, “Fixed forward diffusion process,” 2022
2022
-
[16]
Clip: Connecting text and images,
A. Radford, I. Sutskever, J. Kim, G. Krueger, and S. Agarwal, “Clip: Connecting text and images,” 2021
2021
-
[17]
Zero-shot text-to- image generation,
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. V oss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to- image generation,” ArXiv, vol. abs/2102.12092, 2021
2021 arXiv
-
[18]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” ArXiv, vol. abs/2010.11929, 2020
2010 arXiv
-
[19]
Generating images from caption and vice versa via clip- guided generative latent space search,
F. A. Galatolo, M. G. C. A. Cimino, and G. Vaglini, “Generating images from caption and vice versa via clip- guided generative latent space search,” ArXiv, vol. abs/2102.01645, 2021
2021 arXiv
-
[20]
Hierarchical text-conditional image generation with clip latents,
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen, “Hierarchical text-conditional image generation with clip latents,” ArXiv, vol. abs/2204.06125, 2022
2022 arXiv
-
[21]
Transformers in vision: A survey,
S. H. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” ACM Computing Surveys (CSUR), vol. 54, pp. 1 – 41, 2021
2021
-
[22]
Clip: Contrastive language-image pretraining,
OpenAI, “Clip: Contrastive language-image pretraining,” 2021
2021
-
[23]
High-resolution image synthesis with latent diffusion models,
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” 2021
2021
-
[24]
Revolutionizing image generation by ai: Turning text in . . . - lmu munich,
L. Munich, “Revolutionizing image generation by ai: Turning text in . . . - lmu munich,” September 2022
2022
-
[25]
Laion- 5b: An open large-scale dataset for training next generation image-text models,
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev, “Laion- 5b: An open large-scale dataset for training next generation image...
-
[26]
Stability ai homepage,
S. AI, “Stability ai homepage,” 2023
2023
-
[27]
Stable diffusion public release,
S. AI, “Stable diffusion public release,” 2022
2022
-
[28]
Stable diffusion 2.0 release,
S. AI, “Stable diffusion 2.0 release,” 2022
2022
-
[29]
Art isn’t dead, it’s just machine-generated,
G. Appenzeller, M. Bornstein, M. Casado, and Y . Li, “Art isn’t dead, it’s just machine-generated,” 2022
2022
-
[30]
Dall-e 2,
OpenAI, “Dall-e 2,” 2022
2022
-
[31]
Midjourney ai,
Midjourney, “Midjourney ai,” 2023
2023
-
[32]
Ai generative art tools
pharmapsychotic, “Ai generative art tools.”
-
[33]
Welcome to colaboratory - colaboratory
Google, “Welcome to colaboratory - colaboratory.”
-
[34]
The ai community building the future,
H. Face, “The ai community building the future,” 2023. 16 From Creation to Curriculum: Examining the role of generative AI in Arts Universities
2023
-
[35]
Gradio: Hassle-free sharing and testing of ml models in the wild,
A. Abid, A. Abdalla, A. Abid, D. Khan, A. Alfozan, and J. Y . Zou, “Gradio: Hassle-free sharing and testing of ml models in the wild,” ArXiv, vol. abs/1906.02569, 2019
1906 arXiv
-
[36]
Stable diffusion web ui,
AUTOMATIC1111, “Stable diffusion web ui,” 2023
2023
-
[37]
Lora: Low-rank adaptation of large language models,
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” 2021
2021
-
[38]
Adding conditional control to text-to-image diffusion models,
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” 2023
2023
-
[39]
Constructionism (learning theory),
Wikipedia, “Constructionism (learning theory),” 2023
2023
-
[40]
Constructionism, a learning theory and a model for maker education,
C. Flores, “Constructionism, a learning theory and a model for maker education,” 2023
2023
-
[41]
Piaget’s constructivism, papert’s constructionism: What’s the difference,
E. Ackermann, “Piaget’s constructivism, papert’s constructionism: What’s the difference,” Future of learning group publication, vol. 5, no. 3, p. 438, 2001
2001
-
[42]
The learning value of personalization in children’s reading recommendation systems: What can we learn from constructionism?,
N. Kucirkova, “The learning value of personalization in children’s reading recommendation systems: What can we learn from constructionism?,”International Journal of Mobile and Blended Learning (IJMBL), vol. 11, no. 4, pp. 80–95, 2019
2019
-
[43]
B. G. Wilson, Constructivist learning environments: Case studies in instructional design. Educational Technol- ogy, 1996
1996
-
[44]
Full interview:
60 Minutes, “Full interview: ”godfather of artificial intelligence” talks impact and potential of ai,” 2023. 17
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.