Pith. sign in

REVIEW 3 major objections 5 minor 30 references

Rethinking Citation of AI Sources in Student-AI Collaboration within HCI Design Education

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Based on documentation from 35 student teams in a UX course, this paper finds that over half of AI citations are general mentions that hide design judgment, and argues that AI citation should become a reflective pedagogical practice.

desk verdict Useful empirical snapshot of AI citation behavior under permissive course policy, but the headline majority claim is computed per citation instance, not per team or student. read the letter →

arxiv 2506.08467 v2 pith:3TQUSDFY submitted 2025-06-10 cs.HC

classification cs.HC
keywords AIcitationstudent-AIcollaborationdesigneducationHCIjudgmentmetacognitionqualitativecontentanalysisgenerativein
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks what happens when undergraduate UX design students are free to use generative AI in a team project without a prescribed citation format, and whether current citation conventions can capture their collaboration with AI. Analyzing documentation and group reflections from 35 teams and 175 students, it finds that the most common way students acknowledged AI was the general mention: naming a tool such as ChatGPT without version, date, link, or detail about how it shaped their work. Explicit citations with full metadata accounted for only 17.14% of 70 coded AI references, while indirect mentions in reflections or appendices accounted for 31.43%. The paper concludes that existing citation frameworks, built for stable, retrievable, human-authored sources, are inadequate for student–AI collaboration, and argues that AI citation should be reimagined as a reflective pedagogical practice that records how and why AI was used across the design process.

What carries the argument

The load-bearing machinery is a three-dimensional codebook built from qualitative content analysis. Content type distinguishes text generation from image generation; design phase records whether AI was used in literature review, ideation, prototyping, documentation, or other activities such as interview questions and video scripting; and citation classification sorts each reference into explicit citation, general mention, indirect mention, or no citation. The distribution tabulated from these codes carries the conclusion that existing citation norms fall short, and it also frames the paper's proposed alternative of process-aware, reflective citation models such as AI contribution statements.

What would settle it

A concrete check would be to take the same 35 project documents, have independent coders blind to the original analysis re-classify every AI reference with a reported reliability measure, and compare the resulting distribution. If general mention is no longer the majority category, or if coder agreement is low, the paper's distributional conclusion fails; likewise, if real-time observation of the teams during the project showed AI use substantially beyond what the documents report, the artifacts would not support the inadequacy claim.

Watch

Extended reading notes

Core claim

The central empirical claim is a distribution of citation behaviors among students given unrestricted AI use. In 70 coded AI-use references across 35 team projects, 51.43% were general mentions that named a tool but omitted version, date, or detailed rationale, 31.43% were indirect mentions in reflections or appendices, and 17.14% were explicit citations with full metadata; five teams made no mention of AI at all. The paper reads this distribution as evidence that current AI citation practices are inadequate for capturing the complexities of student–AI collaboration in HCI design education, because vague mentions obscure the traces of design judgment that instructors need for assessment and feedback. Its constructive claim is that citation should become a reflective documentation practice—like sketching or critique—that records when AI was used, what content it generated, how the tool was engaged, and why outputs were used, adapted, or discarded.

Load-bearing premise

The load-bearing premise is that students' written documentation and group reflections are a faithful and complete record of their AI usage and reasoning; if students omitted, downplayed, or overstated their AI use, the observed citation distribution and the conclusion that current frameworks are inadequate would be weakened.

Editorial extensions

If this is right

  • Instructors cannot reliably trace students' design judgment from current documentation, putting the validity of process-based assessment at risk.
  • Formal metadata-only citation guidelines will continue to miss how students actually use AI, so new scaffolds such as AI contribution statements are needed.
  • Reframing citation as reflective documentation would give studio critique and feedback richer material about how design decisions were reached.
  • Students who document prompts, evaluations, and adaptations of AI output are expected to develop stronger self-awareness of their design decisions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: Because the study intentionally gave no citation format, the observed predominance of general mention likely reflects default behavior; a redesigned intervention with reflective scaffolds could be tested against a control section to see whether documentation depth improves.
  • Inference: The 'general mention' habit suggests students treat AI as a utility like a search engine, which points to a need for explicit instruction on AI as a design collaborator rather than merely stricter citation rules.
  • Inference: If adopted, process-aware AI citation would have a natural extension beyond coursework: professional UX designers could use similar contribution statements to communicate AI-assisted design reasoning to stakeholders.
  • Inference: An untested corollary of the pedagogical argument is that reflection quality, not citation completeness, should be the assessed outcome; a rubric measuring how well a citation explains why AI output was used or discarded could operationalize this.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a case study of AI citation practices in an undergraduate UX design course. The authors analyzed 35 team projects and 175 student reflections, coding 70 instances of AI reference into four categories (explicit citation, general mention, indirect mention, no citation). They found that 51.43% of instances were general mentions, and argue that this shows current AI citation practices are inadequate for capturing student-AI collaboration. The paper proposes rethinking AI citation as a reflective pedagogical practice, with suggestions such as AI contribution statements and process-aware citation models.

Significance. If the empirical claims hold, the paper addresses a timely and important gap: how students attribute AI use in design education. Its strengths include a naturalistic setting with a permissive AI policy, a clearly defined coding scheme, and a substantial dataset of 70 coded instances. The proposed pedagogical reframing is valuable and builds on relevant prior work. However, the central quantitative claim is currently undermined by a unit-of-analysis mismatch, and the coding reliability is not reported; both issues require attention before the paper's main conclusion can be accepted.

major comments (3)
  1. [Section 5.1, Table 1] The claim that 'the predominant citation practice adopted by the students was the general mention approach (51.43%)' is based on percentages calculated over 70 coded instances, as reported in Table 1, not over the 35 teams or 175 students. Teams contribute multiple instances (e.g., T19a, T19b, T19c in Section 4.0.1), so a team with several general mentions inflates the instance-level percentage, and a team can appear in multiple categories. Consequently, the instance-level 51.43% does not establish that 'most of the students simply mentioned that they used AI' (Section 5.1) or that general mention was the predominant practice at the team or student level. Please report the distribution of teams by their dominant (or exclusive) citation category, or alternatively qualify the central claim to refer to instances rather than students/teams. This is a load-bearing issue because the paper's motivation for rethinking citation frameworks rests on this predominance claim.
  2. [Section 3.4.3] The coding was performed by the three course instructors and two faculty members, with no inter-rater reliability metric reported. Because the coders also graded the projects and were familiar with the teams, the coding of whether a mention is 'general' versus 'indirect' could be subject to interpretive bias. Please report a reliability check, such as independent coding of a subsample with a measure of agreement and a clear process for resolving disagreements. Also specify how 'instances of reference to AI' were delimited and segmented, since the choice of unit directly determines the percentages in Tables 1 and 2.
  3. [Section 7] The concluding statement that 'current AI citation practices are inadequate for capturing the complexities of student-AI collaboration' is an interpretive extension of the descriptive distribution; the study did not measure the consequences of the observed citation practices for assessment or learning outcomes. Consider softening the conclusion to reflect that the observed variability is in tension with the authors' design-argumentation framework, rather than asserting empirical inadequacy.
minor comments (5)
  1. [Section 2.1] The sentence beginning 'This tension about the suitability of old citation practices...' is a run-on; consider splitting it into two sentences for readability.
  2. [References] Reference [30] has formatting errors in the author list (e.g., 'Wu, Wang , Jiatong' with stray spaces and comma placements). Please correct the reference entry.
  3. [Figure 1] The text refers to Figure 1 as an example of explicit citation, but the figure's content is not described in the body text. Provide a brief explanation in the text or a more detailed caption.
  4. [Table 1] The table would benefit from a footnote stating that the five teams with 'No Citation Mentioned' were excluded from the totals, since the text mentions this but the table alone is ambiguous.
  5. [Section 4.0.2] The phrase 'where most of the students simply mentioned that they used AI' appears before the team-level analysis; consider moving the team-level interpretation to the discussion where it can be properly qualified.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the citation-practice percentages are descriptive counts of coded artifacts, and the self-citations are non-load-bearing background support.

full rationale

The paper's central empirical claim is the distribution of coded citation practices (12 explicit, 36 general mention, 22 indirect, 5 no-citation codes) reported in Section 4 and Tables 1-2. These numbers are counts of instances extracted from student documentation (Section 3.4.3), not quantities derived from a fitted model or from the paper's conclusions. No parameter is fitted and then renamed as a prediction; the claim that general mention was 'predominant' (51.43%) is a direct arithmetic summary of the coded instance counts. The interpretive conclusion that current AI citation practices are inadequate is an argument, not a quantity derived from the data by construction. The self-citations (Parsons and Gray [22]; Shukla et al. [26]) appear in the background section to support general premises about design-studio transparency and concerns about AI obscuring design thinking; the empirical findings do not depend on those premises being true, and the cited works are not used to define any coding category or percentage. The skeptical concern that 51.43% is an instance-level figure being interpreted at team/student level is a unit-of-analysis validity issue, not circularity: the percentage would be the same regardless of whether the conclusion is warranted, and it is not an input to its own calculation. Likewise, the acknowledged limitation that written documentation may not fully reflect AI use (Section 6) weakens external validity but does not make the reported distributions circular. Under the hard rule requiring a quoted reduction of a claim to its own inputs, no such reduction exists here.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim does not rest on fitted parameters or invented entities. It rests on the trustworthiness of students' self-reported documentation, the validity of the coding scheme, and the representativeness of a single-course sample. These are domain assumptions, identified in the paper's limitations.

assumptions (3)
  • domain assumption Students' written documentation and reflections provide a valid and complete record of their AI usage and reasoning.
    All empirical results are derived from these artifacts (Section 3.3). The authors note in Section 6 that documentation may not fully reflect the scope of AI integration.
  • domain assumption The four citation categories (explicit, general mention, indirect, no citation) are mutually exclusive and collectively exhaustive for coding AI references.
    The codebook (Section 3.4.2) defines these bins and the analysis depends on this categorization to produce Table 1.
  • domain assumption A single course at one institution is representative enough to support broad claims about gaps in existing citation frameworks.
    Section 3.1 describes the single-course context; Section 6 acknowledges limited generalizability, yet the provocations in the Discussion are framed broadly.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Rethinking Citation of AI Sources in Student-AI Collaboration within HCI Design Education." pith.science (2026). https://pith.science/paper/3TQUSDFY

@misc{pith2026250608467,
  author       = {Pith},
  title        = {Pith review of: Rethinking Citation of AI Sources in Student-AI Collaboration within HCI Design Education},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3TQUSDFY}},
  note         = {Machine review of arXiv:2506.08467}
}
read the original abstract

The growing integration of AI tools in student design projects presents an unresolved challenge in HCI education: how should AI-generated content be cited and documented? Traditional citation frameworks -- grounded in credibility, retrievability, and authorship -- struggle to accommodate the dynamic and ephemeral nature of AI outputs. In this paper, we examine how undergraduate students in a UX design course approached AI usage and citation when given the freedom to integrate generative tools into their design process. Through qualitative analysis of 35 team projects and reflections from 175 students, we identify varied citation practices ranging from formal attribution to indirect or absent acknowledgment. These inconsistencies reveal gaps in existing frameworks and raise questions about authorship, assessment, and pedagogical transparency. We argue for rethinking AI citation as a reflective and pedagogical practice; one that supports metacognitive engagement by prompting students to critically evaluate how and why they used AI throughout the design process. We propose alternative strategies -- such as AI contribution statements and process-aware citation models that better align with the iterative and reflective nature of design education. This work invites educators to reconsider how citation practices can support meaningful student--AI collaboration.

Figures

Figures reproduced from arXiv: 2506.08467 by the authors.

Figure 1
Figure 1. Example of explicit citation showcasing how the [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

30 extracted references · 17 canonical work pages

  1. [1]

    Nicole Alea Albada and Vanessa E. Woods. 2024. Giving Credit Where Credit is Due: An Artificial Intelligence Contribution Statement for Research Methods Writing Assignments. Teaching of Psychology (June 2024), 00986283241259750. https://doi.org/10.1177/00986283241259750

  2. [2]

    Erik Borg. 2000. Citation practices in academic writing. Patterns and perspectives: Insights into EAP writing practice (Jan. 2000), 27–45

  3. [3]

    Cecilia Ka Yuk Chan and Tom Colloton. 2024. Generative AI in Higher Educa- tion: The ChatGPT Effect (1 ed.). Routledge, London. https://doi.org/10.4324/ 9781003459026

  4. [4]

    Vigneshkumar Chellappa and Yan Luximon. 2024. Understanding the percep- tion of design students towards ChatGPT. Computers and Education: Artificial Intelligence 7 (Dec. 2024), 100281. https://doi.org/10.1016/j.caeai.2024.100281

  5. [5]

    Ricky Chen, Mychajlo Demko, Daragh Byrne, and Marti Louw. 2021. Probing Documentation Practices: Reflecting on Students’ Conceptions, Values, and Ex- periences with Documentation in Creative Inquiry. In Creativity and Cognition. ACM, Virtual Event Italy, 1–14. https://doi.org/10.1145/3450741.3465391

  6. [6]

    Debby R. E. Cotton, Peter A. Cotton, and J. Reuben Shipway. 2024. Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International 61, 2 (March 2024), 228–239. https: //doi.org/10.1080/14703297.2023.2190148

  7. [7]

    Peter Dalsgaard, Christian Dindler, and Jonas Fritsch. 2013. Design argumentation in academic design education. Nordes 1, 5 (June 2013). https://archive.nordes. org/index.php/n13/article/view/331 Number: 5

  8. [8]

    Deanna P. Dannels. 2005. Performing Tribal Rituals: A Genre Analysis of “Crits” in Design Studios. Communication Education 54, 2 (April 2005), 136–160. https://doi.org/10.1080/03634520500213165 Publisher: NCA Website _eprint: https://doi.org/10.1080/03634520500213165

Show all 30 references
  1. [9]

    Satu Elo and Helvi Kyngäs. 2008. The qualitative content analysis process. Journal of Advanced Nursing 62, 1 (2008), 107–115. https://doi.org/10.1111/j.1365- 2648.2007.04569.x _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1365- 2648.2007.04569.x

  2. [10]

    Fazilatfar, S

    Ali M. Fazilatfar, S. E. Elhambakhsh, and Hamid Allami. 2018. An Investigation of the Effects of Citation Instruction to Avoid Plagiarism in EFL Academic Writing Assignments. SAGE Open 8, 2 (April 2018), 2158244018769958. https://doi.org/ 10.1177/2158244018769958 Publisher: SA...

  3. [11]

    Jane Forman and Laura Damschroder. 2007. Qualitative Content Analysis. In Empirical Methods for Bioethics: A Primer . Vol. 11. Emerald Group Publishing Limited, 39–62. https://doi.org/10.1016/S1479-3709(07)11003-7 ISSN: 1479-3709

  4. [12]

    Colin M. Gray. 2018. Narrative Qualities of Design Argumentation. InEducational Technology and Narrative: Story and Instructional Design, Brad Hokanson, Gregory Clinton, and Karen Kaminski (Eds.). Springer International Publishing, Cham, 51–64. https://doi.org/10.1007/978-3-31...

  5. [13]

    Jessica He, Stephanie Houde, and Justin D. Weisz. 2025. Which Contributions Deserve Credit? Perceptions of Attribution in Human-AI Co-Creation. https: //doi.org/10.48550/arXiv.2502.18357 arXiv:2502.18357 [cs]

  6. [14]

    Mohammad Hosseini, David B Resnik, and Kristi Holmes. 2023. The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscripts. Research Ethics 19, 4 (Oct. 2023), 449–465. https://doi.org/10. 1177/17470161231180449

  7. [15]

    Hsiu-Fang Hsieh and Sarah E. Shannon. 2005. Three Approaches to Qualitative Content Analysis. Qualitative Health Research 15, 9 (Nov. 2005), 1277–1288. https://doi.org/10.1177/1049732305276687 Publisher: SAGE Publications Inc

  8. [16]

    Olimpius Istrate. 2023. How to cite ChatGPT and AI products.Revista de Pedagogie Digitala 2, 1 (2023), 3–8. https://doi.org/10.61071/RPD.2347

  9. [17]

    Scott Lanning. 2016. A modern, simplified citation style and student response. Reference Services Review 44, 1 (Feb. 2016), 21–37. https://doi.org/10.1108/RSR- 10-2015-0045

  10. [18]

    Seo-young Lee, Matthew Law, and Guy Hoffman. 2025. When and How to Use AI in the Design Process? Implications for Human-AI Design Collaboration. International Journal of Human–Computer Interaction (Jan. 2025). https://www. tandfonline.com/doi/full/10.1080/10447318.2024.2353451...

  11. [19]

    Jason Lively, James Hutson, and Elizabeth Melick. 2023. Integrating AI-Generative Tools in Web Design Education: Enhancing Student Aesthetic and Creative Copy Capabilities Using Image and Text-Based AI Generators. DS Journal of Artificial Intelligence and Robotics 1, 1 (Sept. ...

  12. [20]

    Timothy McAdoo. 2023. How to cite ChatGPT. https://apastyle.apa.org/blog/ how-to-cite-chatgpt

  13. [21]

    Eleftherios Papachristos, Yavuz Inal, Carlos Vicient Monllaó, Eivind Arnstein Johansen, and Mari Hermansen. 2024. Integrating AI into Design Ideation: Assessing ChatGPT’s Role in Human-Centered Design Education. https://doi. org/10.36227/techrxiv.171656320.09963657/v1

  14. [22]

    Paul C Parsons and Colin M Gray. 2022. Separating Grading and Feedback in UX Design Studios. In EduCHI’22: 4th Annual Symposium on HCI Education . New Orleans, LA, 9

  15. [23]

    Paul C Parsons, Prakash Chandra Shukla, Ali Baigelenov, and Colin M Gray. 2023. Developing Framing Judgment Ability: Student Perceptions from a Graduate UX Design Program. In Proceedings of the 5th Annual Symposium on HCI Education . ACM, Hamburg Germany, 23–32. https://doi.or...

  16. [24]

    Hauke Sandhaus, Quiquan Gu, Maria Teresa Parreira, and Wendy Ju. 2024. Student Reflections on Self-Initiated GenAI Use in HCI Education. http: //arxiv.org/abs/2410.14048 arXiv:2410.14048 [cs]

  17. [25]

    Yang Shi, Tian Gao, Xiaohan Jiao, and Nan Cao. 2023. Understanding Design Collaboration Between Designers and Artificial Intelligence: A Systematic Liter- ature Review. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2 (Sept. 2023), 1–35. https://doi.org/10.1145/3610217

  18. [26]

    Prakash Shukla, Phuong Bui, Sean S Levy, Max Kowalski, Ali Baigelenov, and Paul Parsons. 2025. De-skilling, Cognitive Offloading, and Misplaced Responsibilities: Potential Ironies of AI-Assisted Design. In Proceedings of the Extended Abstracts of the CHI Conference on Human Fa...

  19. [27]

    Prakash Shukla, Suchismita Naik, Ike Obi, Phuong Bui, and Paul Parsons. 2024. Communication Challenges Reported by UX Designers on Social Media: An Anal- ysis of Subreddit Discussions. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ...

  20. [28]

    Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. 2024. The Metacognitive Demands and Opportunities of Generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems . ACM, Honolulu ...

  21. [29]

    Yadi Wang and Susan R. Fussell. 2024. They May Have Seen My ChatGPT Tab: Exploring Social Perceptions of AI-Assisted Writing for ESL Students. In Proceedings of the Third Workshop on Intelligent and Interactive Writing Assistants . ACM, Honolulu HI USA, 13–15. https://doi.org/...

  22. [30]

    Junqi Wu, Wang , Jiatong, Lei , Shuang, Wu , Feiyan, , and Xingyu Gao. 2025. The impact of metacognitive scaffolding on deep learning in a GenAI-supported learning environment. Interactive Learning Environments 0, 0 (2025), 1–18. https: //doi.org/10.1080/10494820.2025.2479162

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.