REVIEW 4 major objections 3 minor 25 references
Despite having no team rules, hackathon participants converged on an unwritten practice of checking generative AI output before use, but time pressure and unfamiliar domains limited how well they could verify.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 12:11 UTC pith:JD4WMHEY
load-bearing objection Small, honest hackathon interview study with a solid core and one overclaim: the 'all participants' checking finding lacks P2's supporting quote. the 4 major comments →
Quick Build, Careful Check? Generative AI Use in Hackathons
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's core discovery is that informal verification norms can arise in temporary, fast-moving teams even without explicit governance. All four interviewed participants—from different teams—reported that AI output was always edited, cross-checked, or at least read before use, motivated by a sense of individual responsibility or group pressure to understand one's own code. The paper also documents the limits of that practice: participants with weak domain knowledge said they could not reliably judge AI output, and the same participants abandoned their preferred habit of reviewing small chunks of code under the 28-hour deadline. These accounts support the paper's claim that hackathons do n
What carries the argument
The central object is the unwritten checking practice: a habit of reviewing, editing, or cross-validating generative AI output before use, which the paper finds operating even where no team rules exist. The mechanism has two parts—an individual motive (being able to explain and take responsibility for submitted code) and a social one (group pressure not to let all code be AI-written). Its effectiveness is determined by two opposing forces: it persists because participants bring it from everyday work, and it degrades because time pressure and domain-knowledge limits shrink the review each output receives. This object carries the paper's argument that governance can emerge bottom-up in ad hoc
Load-bearing premise
The load-bearing premise is that a single interviewed member per team accurately described the team's shared behavior, so if those four people misremembered or were atypical, the convergence on checking may not be real.
What would settle it
A direct observation study with screen recordings and interaction logs of hackathon teams: if a notable share of AI-generated code is adopted without any edit, review, or cross-check, the universal-checking claim would be falsified.
If this is right
- Hackathon organizers can treat verification as an activity to support rather than a behavior to mandate—for example, judging 'explainability of contribution' alongside skillful AI use, as the paper suggests.
- Because time pressure curtails checking, events that build review steps into the schedule, such as mid-hack code-review checkpoints, could improve the reliability of AI-assisted output.
- Participants working outside their domain are least able to verify; pairing them with domain experts or providing domain primers could reduce uncaught AI errors.
- The presence of informal checking norms implies that interventions aimed at preventing over-reliance should reinforce existing habits rather than assume no verification happens.
- As agentic AI tools take over larger units of work, the human review window shrinks, so future support must target these tools specifically.
Where Pith is reading between the lines
- My inference: the convergence on checking may be specific to participants who self-selected into an AI-themed hackathon and already used GenAI; a broader sample with more varied AI experience might show weaker norms.
- My inference: the one-interviewee-per-team design leaves open that the 'unwritten rule' was one person's perception; interviewing all team members could reveal whether checking was collective or just individual habit.
- My inference: the finding suggests a testable extension—measuring actual verification depth (e.g., edits made to AI code, cross-model checks) across teams with different domain familiarity, to see whether the self-reported constraints correspond to behavioral differences.
- My inference: the impostor-syndrome reflection hints that hackathon AI use may affect participants' confidence beyond the event; a longitudinal follow-up could measure whether AI-heavy practice changes self-efficacy over time.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a qualitative interview study of four participants from different teams at a two-day AI-themed hackathon in central Europe. It addresses two research questions: how participants use generative AI (GenAI) across tasks and tools, and how they verify GenAI outputs under hackathon time pressure. The main findings are that participants used GenAI for learning, ideation, coding, and documentation; that they combined multiple GenAI and non-GenAI tools according to task fit; and that, despite the absence of explicit team rules, all participants supposedly converged on an unwritten practice of checking GenAI output before using it, albeit constrained by time pressure and limited domain knowledge. The paper explicitly frames the study as exploratory and proposes a follow-up mixed-methods design combining observation, surveys, and prompt-and-response logs.
Significance. If the convergence claim holds, the paper makes a useful empirical contribution: it complicates the assumption that hackathon time pressure automatically leads to uncritical acceptance of GenAI output, and it shifts attention to supporting informal verification practices rather than merely permitting or restricting GenAI use. The paper is transparent about its limitations—small sample, single event, one interviewee per team, perception-based data, single-coder thematic analysis—and provides a valuable interview guide in the appendix. Its contribution is modest but appropriate for an exploratory empirical study. However, the headline universal claim is not fully supported by the evidence actually presented, and the Discussion repeats that unsupported universal as if established. These are load-bearing issues that need to be fixed before the conclusions can be accepted as stated.
major comments (4)
- [Section 4.3 and Abstract] The headline finding—'all participants we studied converged on an unwritten practice of checking GenAI output before using it'—is supported by direct quotes only for P1, P3, and P4. P2's only quoted statements in this section are 'I often didn't really understand whether it was doing the right thing or not, because I'm not that strong mathematically' and 'couldn't really judge a lot of the time.' Wanting to check is not established by these quotes; if anything, they report an inability to verify. Since the universal quantifier is load-bearing for RQ2 and the abstract, either a supporting quote from P2 must be added or the claim should be hedged to 'three of four participants described...' or otherwise qualified.
- [Section 5.1] The Discussion repeats the unsupported universal: 'all four participants described this expectation as something they held.' This is stronger than the data in §4.3, which supplies explicit statements for P1, P3, and P4 only. The Discussion also moves from 'described wanting to check' to 'described this expectation' without additional data. Please align the Discussion with whatever evidence exists for P2, or narrow the claim to the participants who actually articulated it.
- [Section 4.3, 'Two constraints'] The text says 'Two constraints recurred across the cases.' In the reported excerpts, the domain-knowledge constraint is illustrated only by P2 and the time constraint only by P1. 'Recurred' overstates the support: the data show these constraints were salient to at least one participant each, not that they recurred across multiple cases. Either provide evidence from other participants or revise the phrasing to something like 'two constraints appeared in our data.'
- [Section 4.3 and Section 6] The claim that 'no team set explicit rules' is presented as a fact in §4.3 and the abstract, but Section 6 limits team-level claims to a single member's perception. For an exploratory study this is an acceptable limitation, but the wording should carry the same hedge that §5.1 uses ('as far as participants reported'), because the absence-of-rules premise is part of the headline contrast.
minor comments (3)
- [Section 3, first paragraph] There are spacing/formatting inconsistencies in phrases like 'relevant toRQ2' and 'relevant toRQ 1' (missing space before 'RQ' in one place).
- [References, [17]] The DOI for the Pe-Than et al. paper appears malformed or inconsistent with the publisher's usual format; please verify it.
- [Section 4, Key findings bullet] The bullet 'Despite no explicit team rules, participants checked GenAI output' would be more accurate if written as 'participants reported checking GenAI output in this small sample,' consistent with the evidence in the body.
Circularity Check
No circularity: the paper's findings are empirical summaries of interview data, not derivations from fitted inputs or from the authors' own prior results.
full rationale
This paper is a qualitative interview study with no equations, no fitted parameters, and no formal derivation chain. The central findings (GenAI used for learning, ideation, coding, documentation; tool combination by task fit; informal verification practice; time and knowledge constraints on verification) are presented as thematic summaries of participant quotes in Section 4. There is no step where an input is renamed as a prediction, no uniqueness theorem imported from prior work, and no ansatz smuggled in via self-citation. The related-work section does cite several hackathon publications co-authored by Nolte ([5], [6], [17], [18]), but these are contextual background on hackathon research and are not load-bearing for the paper's empirical claims; the Discussion's comparative points rely on external references such as [1], [8], [10], [13], [20], [22], [23], [24]. The Limitations section itself acknowledges the single-perspective-per-team and perception-based nature of the data, which is an honest validity caveat rather than evidence of circularity. The most notable weakness, that the Section 4.3 universal claim 'every participant described wanting to check or adjust GenAI output before using it' is supported by quotes from P1, P3, and P4 while P2's quoted statements emphasize inability to judge, is an internal-evidence gap about the strength of the empirical claim; it does not make the claim circular. Under the hard rules, a supported-evidence gap is a correctness/validity concern, not a circularity reduction. Therefore the appropriate circularity score is 0.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Self-reported interview accounts are treated as valid evidence of participants' behavior and reasoning during the hackathon.
- domain assumption One interviewed participant per team is sufficient to infer team-level practices such as the absence of GenAI rules.
- domain assumption Thematic analysis as implemented by a single coder with coauthor discussion yields reliable themes.
Cite this review
Pith. "Pith review of Quick Build, Careful Check? Generative AI Use in Hackathons." pith.science (2026). https://pith.science/paper/JD4WMHEY
@misc{pith2026260729178,
author = {Pith},
title = {Pith review of: Quick Build, Careful Check? Generative AI Use in Hackathons},
year = {2026},
howpublished = {\url{https://pith.science/paper/JD4WMHEY}},
note = {Machine review of arXiv:2607.29178}
}
read the original abstract
Hackathons are time-bounded events where participants form teams to rapidly build software projects. Their short-term nature makes them a natural setting for Generative AI (GenAI) use, given its promise of speed and efficiency. Yet despite GenAI being increasingly adopted in these events, we still know little about what participants use GenAI for and, crucially, what they do not use it for and why. We report an exploratory interview study with participants from different teams at a two-day AI-themed hackathon in central Europe. We found that participants used GenAI for purposes beyond coding, including learning unfamiliar topics, brainstorming, and preparing documentation. At the same time, they combined multiple GenAI and non-GenAI tools depending on task fit. We also found that, despite the absence of team rules, all participants we studied converged on an unwritten practice of checking GenAI output before using it, yet their ability to actually verify was constrained by time pressure and limited domain knowledge. This work aims to identify avenues for further investigation by outlining a follow-up study combining observation, surveys, and prompt-and-response logs.
Reference graph
Works this paper leans on
-
[1]
Sadia Afroz, Zixuan Feng, Tyler Menezes, Katie Kimura, Bianca Trinkenreich, Igor Steinmacher, and Anita Sarma. 2026. The Fast and Spurious: Developer Productivity with GenAI. arXiv:2510.24265 [cs.SE] https://arxiv.org/abs/2510.24265
Pith/arXiv arXiv 2026
-
[2]
Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psychology.Qualitative Research in Psychology3, 2 (2006), 77–101. doi:10.1191/1478088706qp063oa
-
[3]
Connie W. Chau and Elizabeth M. Gerber. 2023. On Hackathons: A Multidisciplinary Literature Review. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems(Hamburg, Germany)(CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 637, 21 pages. doi:10.1145/3544548.3581234
arXiv 2023
-
[4]
Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, and Tobias Salz. 2025. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. (August 2025). doi:10.2139/ssrn.4945566
-
[5]
Jeanette Falk, Alexander Nolte, Daniela Huppenkothen, Marion Weinzierl, Kiev Gama, Daniel Spikol, Erik Tollerud, Neil Chue Hong, Ines Knäpper, and Linda Bailey Hayden. 2024. The future of hackathon research and practice.IEEE Access(2024)
2024
-
[6]
Kiev Gama, Filipe Calegario, Victoria Jackson, Alexander Nolte, Luiz Augusto Morais, and Vinicius Garcia. 2025. " Can you feel the vibes?": An exploration of novice programmer engagement with vibe coding.arXiv preprint arXiv:2512.02750(2025)
arXiv 2025
-
[7]
Kiev Gama, George Valença, Pedro Alessio, Rafael Formiga, André Neves, and Nycolas Lacerda. 2023. The developers’ design thinking toolbox in hackathons: a study on the recurring design methods in software development marathons.International Journal of Human-Computer Interaction39, 12 (2023), 2269–2291. doi:10.1080/10447318.2022.2075601
arXiv 2023
-
[8]
Paloma Guenes, Rafael Tomaz, Marcos Kalinowski, Maria Teresa Baldassarre, and Margaret-Anne Storey. 2024. Impostor Phenomenon in Software Engineers. InProceedings of the 46th International Conference on Software Engineering: Software Engineering in Society (ICSE-SEIS ’24). doi:10.1145/ 3639475.3640114
arXiv 2024
-
[9]
Siw Elisabeth Hove and Bente Anda. 2005. Experiences from conducting semi-structured interviews in empirical software engineering research. In 11th IEEE International Software Metrics Symposium (METRICS’05). IEEE, 10–pp
2005
-
[10]
Ranim Khojah, Mazen Mohamad, Philipp Leitner, and Francisco Gomes de Oliveira Neto. 2024. Beyond Code Generation: An Observational Study of ChatGPT Usage in Software Engineering Practice.Proceedings of the ACM on Software Engineering1, FSE (2024), 1819–1840. doi:10.1145/3660788
doi:10.1145/3660788 2024
-
[11]
Marko Komssi, Danielle Pichlis, Mikko Raatikainen, Klas Kindström, and Janne Järvinen. 2015. What are Hackathons for?IEEE Software32, 5 (2015), 60–67. doi:10.1109/MS.2014.78
-
[12]
Miikka Kuutila, Mika Mäntylä, Umar Farooq, and Maëlick Claes. 2020. Time Pressure in Software Engineering: A Systematic Review.Information and Software Technology121 (2020), 106257. doi:10.1016/j.infsof.2020.106257
arXiv 2020
-
[13]
Lu Li, Savindu Herath, and Cyrille Grumbach. 2026. Hack-Agents: A Multi-Agent System for Innovation - Proof of Concept and Implications for AI-Augmented Hackathons(CHI EA ’26). Association for Computing Machinery, New York, NY, USA, Article 359, 6 pages. doi:10.1145/3772363.3798678
arXiv 2026
-
[14]
Lincoln and Egon G
Yvonna S. Lincoln and Egon G. Guba. 1985.Naturalistic Inquiry. Sage Publications, Beverly Hills, CA. Manuscript submitted to ACM 8 Wangyiyao Zhou, Alexander Serebrenik, and Alexander Nolte
1985
-
[15]
Amr Mohamed, Maram Assi, and Mariam Guizani. 2026. The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and Mapping Study.ACM Transactions on Software Engineering and Methodology(2026). doi:10.1145/3809494
doi:10.1145/3809494 2026
-
[16]
Anh Nguyen-Duc, Beatriz Cabrero-Daniel, Adam Przybylek, Chetan Arora, Dron Khanna, Tomas Herda, Usman Rafiq, Jorge Melegati, Eduardo Guerra, Kai-Kristian Kemell, Mika Saari, Zheying Zhang, Huy Le, Tho Quan, and Pekka Abrahamsson. 2025. Generative Artificial Intelligence for Software Engineering—A Research Agenda.Software: Practice and Experience55, 11 (20...
- [17]
-
[18]
Ei Pa Pa Pe-Than, Alexander Nolte, Anna Filippova, Christian Bird, Steve Scallen, and James D. Herbsleb. 2022. Corporate Hackathons, How and Why? A Multiple Case Study of Motivation, Projects Proposal and Selection, Goal Setting, Coordination, and Outcomes.Human-Computer Interaction37, 4 (2022), 281–313. doi:10.1080/07370024.2020.1760869
arXiv 2022
-
[19]
Jari Porras, Jayden Khakurel, Jouni Ikonen, Ari Happonen, Antti Knutas, Antti Herala, and Olaf Drögehorn. 2018. Hackathons in Software Engineering Education: Lessons Learned from a Decade of Events. InProceedings of the 2nd International Workshop on Software Engineering Education for Millennials (SEEM ’18). ACM, New York, NY, USA, 40–47. doi:10.1145/31947...
arXiv 2018
-
[20]
Ramteja Sajja, Carlos Erazo Ramirez, Zhouyayan Li, Bekir Z Demiray, Yusuf Sermet, and Ibrahim Demir. 2024. Integrating generative AI in hackathons: Opportunities, challenges, and educational implications.Big Data and Cognitive Computing8, 12 (2024), 188
2024
-
[21]
Italo Santos, Cleyton Magalhaes, and Ronnie De Souza Santos. 2025. Model-Assisted and Human-Guided: Perceptions and Practices of Software Professionals Using LLMs for Coding. In2025 2nd IEEE/ACM International Conference on AI-powered Software (AIware). 105–112. doi:10.1109/ AIware69974.2025.00019
arXiv 2025
-
[22]
Igor Steinmacher, Marco Aurélio Graciotto Silva, and Marco Aurélio Gerosa. 2014. Barriers faced by newcomers to open source projects: a systematic review. InIFIP International Conference on Open Source Systems. Springer, 153–163
2014
-
[23]
Kelly B. Wagman, Matthew T. Dearing, and Marshini Chetty. 2025. Generative AI Uses and Risks for Knowledge Workers in a Science Organization. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM. doi:10.1145/3706598.3713827
arXiv 2025
-
[24]
Wenhao Yang, Runzhi He, and Minghui Zhou. 2026. Beyond Banning AI: A First Look at GenAI Governance in Open Source Software Communities. arXiv:2603.26487 [cs.SE] https://arxiv.org/abs/2603.26487
Pith/arXiv arXiv 2026
-
[25]
Runlong Ye, Oliver Huang, Jessica He, and Michael Liut. 2026. Exploring Emerging Norms of AI Attribution and Disclosure in Programming Education. arXiv:2602.04023 [cs.HC] https://arxiv.org/abs/2602.04023 A Interview Guide IntroductionThis interview is part of a research project about how people use Generative AI tools like ChatGPT, Gemini, or image genera...
arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.