REVIEW 3 major objections 5 minor 30 references
Rethinking Citation of AI Sources in Student-AI Collaboration within HCI Design Education
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Based on documentation from 35 student teams in a UX course, this paper finds that over half of AI citations are general mentions that hide design judgment, and argues that AI citation should become a reflective pedagogical practice.
desk verdict Useful empirical snapshot of AI citation behavior under permissive course policy, but the headline majority claim is computed per citation instance, not per team or student. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a three-dimensional codebook built from qualitative content analysis. Content type distinguishes text generation from image generation; design phase records whether AI was used in literature review, ideation, prototyping, documentation, or other activities such as interview questions and video scripting; and citation classification sorts each reference into explicit citation, general mention, indirect mention, or no citation. The distribution tabulated from these codes carries the conclusion that existing citation norms fall short, and it also frames the paper's proposed alternative of process-aware, reflective citation models such as AI contribution statements.
What would settle it
A concrete check would be to take the same 35 project documents, have independent coders blind to the original analysis re-classify every AI reference with a reported reliability measure, and compare the resulting distribution. If general mention is no longer the majority category, or if coder agreement is low, the paper's distributional conclusion fails; likewise, if real-time observation of the teams during the project showed AI use substantially beyond what the documents report, the artifacts would not support the inadequacy claim.
Extended reading notes
Core claim
The central empirical claim is a distribution of citation behaviors among students given unrestricted AI use. In 70 coded AI-use references across 35 team projects, 51.43% were general mentions that named a tool but omitted version, date, or detailed rationale, 31.43% were indirect mentions in reflections or appendices, and 17.14% were explicit citations with full metadata; five teams made no mention of AI at all. The paper reads this distribution as evidence that current AI citation practices are inadequate for capturing the complexities of student–AI collaboration in HCI design education, because vague mentions obscure the traces of design judgment that instructors need for assessment and feedback. Its constructive claim is that citation should become a reflective documentation practice—like sketching or critique—that records when AI was used, what content it generated, how the tool was engaged, and why outputs were used, adapted, or discarded.
Load-bearing premise
The load-bearing premise is that students' written documentation and group reflections are a faithful and complete record of their AI usage and reasoning; if students omitted, downplayed, or overstated their AI use, the observed citation distribution and the conclusion that current frameworks are inadequate would be weakened.
Editorial extensions
If this is right
- Instructors cannot reliably trace students' design judgment from current documentation, putting the validity of process-based assessment at risk.
- Formal metadata-only citation guidelines will continue to miss how students actually use AI, so new scaffolds such as AI contribution statements are needed.
- Reframing citation as reflective documentation would give studio critique and feedback richer material about how design decisions were reached.
- Students who document prompts, evaluations, and adaptations of AI output are expected to develop stronger self-awareness of their design decisions.
Reading between the lines
- Inference: Because the study intentionally gave no citation format, the observed predominance of general mention likely reflects default behavior; a redesigned intervention with reflective scaffolds could be tested against a control section to see whether documentation depth improves.
- Inference: The 'general mention' habit suggests students treat AI as a utility like a search engine, which points to a need for explicit instruction on AI as a design collaborator rather than merely stricter citation rules.
- Inference: If adopted, process-aware AI citation would have a natural extension beyond coursework: professional UX designers could use similar contribution statements to communicate AI-assisted design reasoning to stakeholders.
- Inference: An untested corollary of the pedagogical argument is that reflection quality, not citation completeness, should be the assessed outcome; a rubric measuring how well a citation explains why AI output was used or discarded could operationalize this.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a case study of AI citation practices in an undergraduate UX design course. The authors analyzed 35 team projects and 175 student reflections, coding 70 instances of AI reference into four categories (explicit citation, general mention, indirect mention, no citation). They found that 51.43% of instances were general mentions, and argue that this shows current AI citation practices are inadequate for capturing student-AI collaboration. The paper proposes rethinking AI citation as a reflective pedagogical practice, with suggestions such as AI contribution statements and process-aware citation models.
Significance. If the empirical claims hold, the paper addresses a timely and important gap: how students attribute AI use in design education. Its strengths include a naturalistic setting with a permissive AI policy, a clearly defined coding scheme, and a substantial dataset of 70 coded instances. The proposed pedagogical reframing is valuable and builds on relevant prior work. However, the central quantitative claim is currently undermined by a unit-of-analysis mismatch, and the coding reliability is not reported; both issues require attention before the paper's main conclusion can be accepted.
major comments (3)
- [Section 5.1, Table 1] The claim that 'the predominant citation practice adopted by the students was the general mention approach (51.43%)' is based on percentages calculated over 70 coded instances, as reported in Table 1, not over the 35 teams or 175 students. Teams contribute multiple instances (e.g., T19a, T19b, T19c in Section 4.0.1), so a team with several general mentions inflates the instance-level percentage, and a team can appear in multiple categories. Consequently, the instance-level 51.43% does not establish that 'most of the students simply mentioned that they used AI' (Section 5.1) or that general mention was the predominant practice at the team or student level. Please report the distribution of teams by their dominant (or exclusive) citation category, or alternatively qualify the central claim to refer to instances rather than students/teams. This is a load-bearing issue because the paper's motivation for rethinking citation frameworks rests on this predominance claim.
- [Section 3.4.3] The coding was performed by the three course instructors and two faculty members, with no inter-rater reliability metric reported. Because the coders also graded the projects and were familiar with the teams, the coding of whether a mention is 'general' versus 'indirect' could be subject to interpretive bias. Please report a reliability check, such as independent coding of a subsample with a measure of agreement and a clear process for resolving disagreements. Also specify how 'instances of reference to AI' were delimited and segmented, since the choice of unit directly determines the percentages in Tables 1 and 2.
- [Section 7] The concluding statement that 'current AI citation practices are inadequate for capturing the complexities of student-AI collaboration' is an interpretive extension of the descriptive distribution; the study did not measure the consequences of the observed citation practices for assessment or learning outcomes. Consider softening the conclusion to reflect that the observed variability is in tension with the authors' design-argumentation framework, rather than asserting empirical inadequacy.
minor comments (5)
- [Section 2.1] The sentence beginning 'This tension about the suitability of old citation practices...' is a run-on; consider splitting it into two sentences for readability.
- [References] Reference [30] has formatting errors in the author list (e.g., 'Wu, Wang , Jiatong' with stray spaces and comma placements). Please correct the reference entry.
- [Figure 1] The text refers to Figure 1 as an example of explicit citation, but the figure's content is not described in the body text. Provide a brief explanation in the text or a more detailed caption.
- [Table 1] The table would benefit from a footnote stating that the five teams with 'No Citation Mentioned' were excluded from the totals, since the text mentions this but the table alone is ambiguous.
- [Section 4.0.2] The phrase 'where most of the students simply mentioned that they used AI' appears before the team-level analysis; consider moving the team-level interpretation to the discussion where it can be properly qualified.
Circularity Check
No circularity: the citation-practice percentages are descriptive counts of coded artifacts, and the self-citations are non-load-bearing background support.
full rationale
The paper's central empirical claim is the distribution of coded citation practices (12 explicit, 36 general mention, 22 indirect, 5 no-citation codes) reported in Section 4 and Tables 1-2. These numbers are counts of instances extracted from student documentation (Section 3.4.3), not quantities derived from a fitted model or from the paper's conclusions. No parameter is fitted and then renamed as a prediction; the claim that general mention was 'predominant' (51.43%) is a direct arithmetic summary of the coded instance counts. The interpretive conclusion that current AI citation practices are inadequate is an argument, not a quantity derived from the data by construction. The self-citations (Parsons and Gray [22]; Shukla et al. [26]) appear in the background section to support general premises about design-studio transparency and concerns about AI obscuring design thinking; the empirical findings do not depend on those premises being true, and the cited works are not used to define any coding category or percentage. The skeptical concern that 51.43% is an instance-level figure being interpreted at team/student level is a unit-of-analysis validity issue, not circularity: the percentage would be the same regardless of whether the conclusion is warranted, and it is not an input to its own calculation. Likewise, the acknowledged limitation that written documentation may not fully reflect AI use (Section 6) weakens external validity but does not make the reported distributions circular. Under the hard rule requiring a quoted reduction of a claim to its own inputs, no such reduction exists here.
Assumptions & free parameters
assumptions (3)
- domain assumption Students' written documentation and reflections provide a valid and complete record of their AI usage and reasoning.
- domain assumption The four citation categories (explicit, general mention, indirect, no citation) are mutually exclusive and collectively exhaustive for coding AI references.
- domain assumption A single course at one institution is representative enough to support broad claims about gaps in existing citation frameworks.
Cite this review
Pith. "Pith review of Rethinking Citation of AI Sources in Student-AI Collaboration within HCI Design Education." pith.science (2026). https://pith.science/paper/3TQUSDFY
@misc{pith2026250608467,
author = {Pith},
title = {Pith review of: Rethinking Citation of AI Sources in Student-AI Collaboration within HCI Design Education},
year = {2026},
howpublished = {\url{https://pith.science/paper/3TQUSDFY}},
note = {Machine review of arXiv:2506.08467}
}
read the original abstract
The growing integration of AI tools in student design projects presents an unresolved challenge in HCI education: how should AI-generated content be cited and documented? Traditional citation frameworks -- grounded in credibility, retrievability, and authorship -- struggle to accommodate the dynamic and ephemeral nature of AI outputs. In this paper, we examine how undergraduate students in a UX design course approached AI usage and citation when given the freedom to integrate generative tools into their design process. Through qualitative analysis of 35 team projects and reflections from 175 students, we identify varied citation practices ranging from formal attribution to indirect or absent acknowledgment. These inconsistencies reveal gaps in existing frameworks and raise questions about authorship, assessment, and pedagogical transparency. We argue for rethinking AI citation as a reflective and pedagogical practice; one that supports metacognitive engagement by prompting students to critically evaluate how and why they used AI throughout the design process. We propose alternative strategies -- such as AI contribution statements and process-aware citation models that better align with the iterative and reflective nature of design education. This work invites educators to reconsider how citation practices can support meaningful student--AI collaboration.
Figures
Reference graph
Works this paper leans on
-
[1]
Nicole Alea Albada and Vanessa E. Woods. 2024. Giving Credit Where Credit is Due: An Artificial Intelligence Contribution Statement for Research Methods Writing Assignments. Teaching of Psychology (June 2024), 00986283241259750. https://doi.org/10.1177/00986283241259750
-
[2]
Erik Borg. 2000. Citation practices in academic writing. Patterns and perspectives: Insights into EAP writing practice (Jan. 2000), 27–45
work page 2000
-
[3]
Cecilia Ka Yuk Chan and Tom Colloton. 2024. Generative AI in Higher Educa- tion: The ChatGPT Effect (1 ed.). Routledge, London. https://doi.org/10.4324/ 9781003459026
work page 2024
-
[4]
Vigneshkumar Chellappa and Yan Luximon. 2024. Understanding the percep- tion of design students towards ChatGPT. Computers and Education: Artificial Intelligence 7 (Dec. 2024), 100281. https://doi.org/10.1016/j.caeai.2024.100281
arXiv 2024
-
[5]
Ricky Chen, Mychajlo Demko, Daragh Byrne, and Marti Louw. 2021. Probing Documentation Practices: Reflecting on Students’ Conceptions, Values, and Ex- periences with Documentation in Creative Inquiry. In Creativity and Cognition. ACM, Virtual Event Italy, 1–14. https://doi.org/10.1145/3450741.3465391
-
[6]
Debby R. E. Cotton, Peter A. Cotton, and J. Reuben Shipway. 2024. Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International 61, 2 (March 2024), 228–239. https: //doi.org/10.1080/14703297.2023.2190148
arXiv 2024
-
[7]
Peter Dalsgaard, Christian Dindler, and Jonas Fritsch. 2013. Design argumentation in academic design education. Nordes 1, 5 (June 2013). https://archive.nordes. org/index.php/n13/article/view/331 Number: 5
work page 2013
-
[8]
Deanna P. Dannels. 2005. Performing Tribal Rituals: A Genre Analysis of “Crits” in Design Studios. Communication Education 54, 2 (April 2005), 136–160. https://doi.org/10.1080/03634520500213165 Publisher: NCA Website _eprint: https://doi.org/10.1080/03634520500213165
Show all 30 references
-
[9]
Satu Elo and Helvi Kyngäs. 2008. The qualitative content analysis process. Journal of Advanced Nursing 62, 1 (2008), 107–115. https://doi.org/10.1111/j.1365- 2648.2007.04569.x _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1365- 2648.2007.04569.x
2008
-
[10]
Fazilatfar, S
Ali M. Fazilatfar, S. E. Elhambakhsh, and Hamid Allami. 2018. An Investigation of the Effects of Citation Instruction to Avoid Plagiarism in EFL Academic Writing Assignments. SAGE Open 8, 2 (April 2018), 2158244018769958. https://doi.org/ 10.1177/2158244018769958 Publisher: SA...
2018 doi
-
[11]
Jane Forman and Laura Damschroder. 2007. Qualitative Content Analysis. In Empirical Methods for Bioethics: A Primer . Vol. 11. Emerald Group Publishing Limited, 39–62. https://doi.org/10.1016/S1479-3709(07)11003-7 ISSN: 1479-3709
2007 doi
-
[12]
Colin M. Gray. 2018. Narrative Qualities of Design Argumentation. InEducational Technology and Narrative: Story and Instructional Design, Brad Hokanson, Gregory Clinton, and Karen Kaminski (Eds.). Springer International Publishing, Cham, 51–64. https://doi.org/10.1007/978-3-31...
2018 doi
- [13]
-
[14]
Mohammad Hosseini, David B Resnik, and Kristi Holmes. 2023. The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscripts. Research Ethics 19, 4 (Oct. 2023), 449–465. https://doi.org/10. 1177/17470161231180449
2023
-
[15]
Hsiu-Fang Hsieh and Sarah E. Shannon. 2005. Three Approaches to Qualitative Content Analysis. Qualitative Health Research 15, 9 (Nov. 2005), 1277–1288. https://doi.org/10.1177/1049732305276687 Publisher: SAGE Publications Inc
2005 doi
-
[16]
Olimpius Istrate. 2023. How to cite ChatGPT and AI products.Revista de Pedagogie Digitala 2, 1 (2023), 3–8. https://doi.org/10.61071/RPD.2347
2023 doi
-
[17]
Scott Lanning. 2016. A modern, simplified citation style and student response. Reference Services Review 44, 1 (Feb. 2016), 21–37. https://doi.org/10.1108/RSR- 10-2015-0045
2016 doi
-
[18]
Seo-young Lee, Matthew Law, and Guy Hoffman. 2025. When and How to Use AI in the Design Process? Implications for Human-AI Design Collaboration. International Journal of Human–Computer Interaction (Jan. 2025). https://www. tandfonline.com/doi/full/10.1080/10447318.2024.2353451...
2025
-
[19]
Jason Lively, James Hutson, and Elizabeth Melick. 2023. Integrating AI-Generative Tools in Web Design Education: Enhancing Student Aesthetic and Creative Copy Capabilities Using Image and Text-Based AI Generators. DS Journal of Artificial Intelligence and Robotics 1, 1 (Sept. ...
2023 doi
-
[20]
Timothy McAdoo. 2023. How to cite ChatGPT. https://apastyle.apa.org/blog/ how-to-cite-chatgpt
2023
-
[21]
Eleftherios Papachristos, Yavuz Inal, Carlos Vicient Monllaó, Eivind Arnstein Johansen, and Mari Hermansen. 2024. Integrating AI into Design Ideation: Assessing ChatGPT’s Role in Human-Centered Design Education. https://doi. org/10.36227/techrxiv.171656320.09963657/v1
2024
-
[22]
Paul C Parsons and Colin M Gray. 2022. Separating Grading and Feedback in UX Design Studios. In EduCHI’22: 4th Annual Symposium on HCI Education . New Orleans, LA, 9
2022
-
[23]
Paul C Parsons, Prakash Chandra Shukla, Ali Baigelenov, and Colin M Gray. 2023. Developing Framing Judgment Ability: Student Perceptions from a Graduate UX Design Program. In Proceedings of the 5th Annual Symposium on HCI Education . ACM, Hamburg Germany, 23–32. https://doi.or...
2023
-
[24]
Hauke Sandhaus, Quiquan Gu, Maria Teresa Parreira, and Wendy Ju. 2024. Student Reflections on Self-Initiated GenAI Use in HCI Education. http: //arxiv.org/abs/2410.14048 arXiv:2410.14048 [cs]
2024 arXiv
-
[25]
Yang Shi, Tian Gao, Xiaohan Jiao, and Nan Cao. 2023. Understanding Design Collaboration Between Designers and Artificial Intelligence: A Systematic Liter- ature Review. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2 (Sept. 2023), 1–35. https://doi.org/10.1145/3610217
2023 doi
-
[26]
Prakash Shukla, Phuong Bui, Sean S Levy, Max Kowalski, Ali Baigelenov, and Paul Parsons. 2025. De-skilling, Cognitive Offloading, and Misplaced Responsibilities: Potential Ironies of AI-Assisted Design. In Proceedings of the Extended Abstracts of the CHI Conference on Human Fa...
2025
-
[27]
Prakash Shukla, Suchismita Naik, Ike Obi, Phuong Bui, and Paul Parsons. 2024. Communication Challenges Reported by UX Designers on Social Media: An Anal- ysis of Subreddit Discussions. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ...
2024
-
[28]
Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. 2024. The Metacognitive Demands and Opportunities of Generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems . ACM, Honolulu ...
2024
-
[29]
Yadi Wang and Susan R. Fussell. 2024. They May Have Seen My ChatGPT Tab: Exploring Social Perceptions of AI-Assisted Writing for ESL Students. In Proceedings of the Third Workshop on Intelligent and Interactive Writing Assistants . ACM, Honolulu HI USA, 13–15. https://doi.org/...
2024
-
[30]
Junqi Wu, Wang , Jiatong, Lei , Shuang, Wu , Feiyan, , and Xingyu Gao. 2025. The impact of metacognitive scaffolding on deep learning in a GenAI-supported learning environment. Interactive Learning Environments 0, 0 (2025), 1–18. https: //doi.org/10.1080/10494820.2025.2479162
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.