Pith. sign in

REVIEW 3 major objections 6 minor 39 references

Teaching AI coding in a visualization course made final projects more polished but more visually homogeneous, while students mostly refined prompts and almost never asked AI to explain code.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 14:27 UTC pith:2ZXLKLFM

load-bearing objection Useful classroom evidence on vibe-coding prompts and optional-AI uptake, with an honest but soft polish/homogenization claim against one prior offering. the 3 major comments →

arxiv 2607.09938 v1 pith:2ZXLKLFM submitted 2026-07-10 cs.HC

"Code Is Cheap. Show Me the Talk.": Lessons from Teaching and Managing AI Coding Tool Usage in a Visualization Course

classification cs.HC
keywords Generative AIAI codingvisualization educationvibe codingprompt injectionhomogenizationD3.jsprompting
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Generative AI coding tools help students implement visualizations, but they also let students skip the work a course intends them to do. This paper reports one upper-level visualization course that managed AI with hidden prompt injections, oral checkout questions, and two vibe-coding labs after foundational D3 practice. Before those labs, at least half the students already used AI on assignments. In the labs, about half of analyzable prompts were refinements and explanation was nearly absent; when AI coding was optional, 56 percent preferred scaffolded handouts over writing their own prompts. Final projects looked more complete and polished than in the prior offering, yet shared recurring visual patterns such as card grids, large metric summaries, and gradient-filled charts. The authors argue that instructors need clearer AI boundaries, real prompting instruction, and deliberate practice in questioning generic AI designs and adapting them to a specific dataset and story.

Core claim

In a project-based upper-level visualization course that taught AI coding after D3 foundations, student final projects became more polished than in the previous offering but also more visually homogeneous, with shared AI-typical patterns; student prompt logs were dominated by refinement rather than explanation, and a majority preferred scaffolded instructions when AI coding was optional.

What carries the argument

Mixed, goal-specific AI policies enforced by hidden prompt injections in lab handouts and oral conceptual checkouts, plus two vibe-coding labs whose exported prompt histories were coded as GENERATE, DEBUG, REFINE, or EXPLAIN.

Load-bearing premise

The claim that AI teaching raised polish and visual sameness rests on an informal, non-blind comparison to one prior course offering without matched cohorts or controlled grading.

What would settle it

Blind external raters score polish and cross-project visual similarity on final projects from the AI-taught semester versus a matched prior semester; if polish and similarity do not rise together under the AI-taught design, the central observational claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Visualization instructors should teach prompting and critique of generic AI designs, not only tool interfaces.
  • Clearer activity-specific AI boundaries reduce student confusion about what is allowed.
  • Visual homogenization (card grids, metric summaries, shared palettes) is a predictable side effect of unrestricted AI coding.
  • Oral checkouts can surface understanding in fully project-based courses without exams.
  • When a lab is well-scaffolded, many students will choose handouts over free-form vibe coding.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Introducing vibe coding earlier may trade low-level coding skill for faster polish and push assessment toward design rationale and story.
  • Prompt-injection guardrails will weaken as models get better at ignoring hidden instructions.
  • Shared visual signatures (gradient fills, card grids, accent borders) could serve as a practical signal of heavy AI reliance in project courses.
  • Optional AI tracks may self-select students who already find debugging generated code costly.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This experience report describes how an upper-level CS visualization course managed and taught GenAI coding tools across a 15-week, project-based offering (n≈87). The authors applied mixed AI policies by activity, used hidden prompt injections and oral lab checkouts as guardrails, and ran two vibe-coding labs (one required, one optional) with exported prompt histories. They systematically coded 698 (Lab 11) and 212 (Lab 12) analyzable student prompts into GENERATE / DEBUG / REFINE / EXPLAIN, finding REFINE ≈ half of turns and EXPLAIN nearly absent; in the optional lab, 44/78 (56.4%) submissions preferred scaffolded manual instructions. Assignment AI disclosure rates and informal comparison of 23 final projects to a prior offering are used to argue that projects became more polished yet more visually homogeneous, motivating clearer AI boundaries, prompting instruction, and teaching students to question generic AI designs.

Significance. If the observations hold, the paper offers timely, concrete guidance for visualization and CS educators navigating GenAI: a transparent dual-pass prompt taxonomy with reported category shares, a rare optional-track choice result (majority preferring scaffolding), and operational guardrails (prompt injections, oral checkouts) that others can reuse. The systematic prompt-log analysis and the honest student-motivation notes are strengths that go beyond pure anecdote. The homogenization reflection connects classroom practice to a broader HCI concern about AI-driven design convergence and yields a clear pedagogical stance (“Code is cheap. Show me the talk.”). As a reflective experience paper rather than a controlled trial, its value is in transferable practice and falsifiable classroom patterns, not causal identification.

major comments (3)
  1. Sec. 4.3 and the Abstract state that final projects were “more polished” and “more visually homogeneous” than the previous offering, with recurring patterns (card grids and large numeric summaries each in 8/23 projects, gradient-filled lines, map+side panels, accent UI). This comparison is informal, non-blind, single-author pattern spotting with no metrics, inter-rater reliability, or controls for cohort ability, dataset choice, or grading drift (also noted in Sec. 5.1). Because this claim underpins the core recommendation to teach students to question generic AI designs (Sec. 5.2, 5.5), the manuscript should either (a) reframe it explicitly as unblinded instructor impression with no causal attribution, or (b) add a minimal structured coding of both offerings (even post hoc) and state limitations more sharply in Abstract and Sec. 4.3.
  2. Sec. 4.3 attributes increased polish partly to “teaching AI coding tools explicitly” (echoed in Sec. 5.1), yet AI was unrestricted on assignments/final projects and at least half of students already used AI before the vibe labs (Sec. 4.1). The design cannot separate teaching vibe coding from simply permitting AI, cohort effects, or extra-credit incentives for AI disclosure. The causal language should be softened to “coincided with” / “consistent with,” and Sec. 5 should list alternative explanations as first-class limitations rather than only “what we still do not know.”
  3. Table 1 / Sec. 4.2: EXPLAIN is reported as 3% (Lab 11) and 0% (Lab 12), contrasted with ~28% in a prior web-programming study [16]. The labs explicitly instructed students to generate visualizations with AI and, in Lab 11, not to hand-write most code; that task framing may suppress explanation prompts by design. The manuscript should discuss this demand-characteristic confound before treating low EXPLAIN rates as a general student behavior finding that motivates more prompting instruction.
minor comments (6)
  1. Fig. 3 and Fig. 4 captions note redrawn student work; state more clearly in the figure notes that counts (e.g., N=55 bubble charts) come from original submissions, not the redraws, so readers do not confuse illustration with data.
  2. Sec. 3.2: “Labs 0–10” in Fig. 1 vs. “eight D3.js labs” plus two vibe labs in text is slightly inconsistent; align the lab numbering and grade weights so the timeline is unambiguous.
  3. Sec. 4.2 coding procedure: Claude assigned initial tags then one author corrected them. Briefly report agreement rate or number of corrections so dual-pass reliability is inspectable.
  4. Prompt Injection Example 1/2 and Fig. 2 are useful; a short note on whether students discovered or circumvented injections later in the semester would strengthen the “what worked” reflection in Sec. 5.1.
  5. Typos / polish: “Y ang” spacing in author line; “claude-sonnet-5” model string may be nonstandard; footnote 2’s Torvalds inversion is effective but the arXiv date strings (e.g., “July 2026”) look like placeholders—verify before camera-ready.
  6. Related Work (Sec. 2) could cite one additional visualization-education or design-homogenization classroom study if available; current coverage is adequate but slightly thin on assessment redesign under GenAI.

Circularity Check

0 steps flagged

No circularity: retrospective experience report with independent observational claims, not a derivation or prediction chain.

full rationale

This paper reports course design choices, prompt-log coding, track preferences, and informal comparisons of final projects to a prior offering. It contains no equations, fitted parameters, uniqueness theorems, or first-principles derivations whose outputs reduce to their inputs by construction. Prompt categories (GENERATE/DEBUG/REFINE/EXPLAIN) are applied post-hoc to student logs and are not used to define or force the reported percentages. The polish/homogenization claim is an informal retrospective observation against one prior semester, not a fitted prediction or self-definitional result. Self-citations (e.g., related author work on assessment tools) are ordinary background and are not load-bearing for the central empirical claims or pedagogical reflections. The paper is self-contained as an experience report; no circular step of the enumerated kinds is present.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 1 invented entities

As an experience report the paper rests on ordinary educational and HCI assumptions rather than free parameters or new physical entities. The main load-bearing premises are that the prior course offering is a usable baseline, that observed visual patterns indicate AI influence, and that student self-reports and prompt logs are sufficiently honest.

axioms (3)
  • domain assumption The previous course offering (same instructor, similar structure) constitutes a fair informal baseline for judging changes in project polish and visual diversity.
    Invoked throughout Sec. 4.3 and 5.1; no quantitative matching of cohorts or grading rubrics is supplied.
  • domain assumption Recurring visual patterns (card grids, large metric cards, gradient fills, accent borders) that were absent in the prior offering are indicators of AI coding-tool influence rather than independent student fashion or template reuse.
    Central to the homogenization claim in Sec. 4.3; patterns are identified by one author without inter-rater reliability.
  • domain assumption Student-exported prompt histories and self-reported AI use are sufficiently complete and truthful for category analysis.
    Used for Table 1 and disclosure rates in Sec. 4.1–4.2; no independent verification of omitted chats.
invented entities (1)
  • Prompt-injection guardrails embedded in lab handouts no independent evidence
    purpose: Detect or deter direct pasting of lab materials into AI tools by forcing a refusal or a visible marker (e.g., rounded bars).
    Described in Sec. 3.2 and Fig. 2; useful classroom tactic but not independently validated outside this course.

pith-pipeline@v1.1.0-grok45 · 18600 in / 2463 out tokens · 31032 ms · 2026-07-14T14:27:12.955680+00:00 · methodology

0 comments
read the original abstract

Generative Artificial Intelligence (GenAI) coding tools are transforming visualization education. They can assist with implementation and design, but they can also let students bypass intended learning trajectories. In this paper, we share our retrospective experience managing and teaching AI use in an upper-level visualization course. We implemented prompt injections, asked oral checkout questions, and taught two AI coding labs. Prior to our coding labs, at least half of the students had already used AI tools in their assignments. In both AI coding labs, refinement accounted for about half of students' prompting logs, and explanation was almost absent. In the lab where AI coding was optional, 44 of 78 (56.4%) submissions preferred the scaffolded instructions over designing their own prompts. Students' final projects were more polished than in our previous offering, but also more visually homogeneous. Our reflections point to the need for clearer AI use boundaries and instruction on prompting, and for teaching students to question generic AI designs and adapt them to their data and story.

Figures

Figures reproduced from arXiv: 2607.09938 by Fumeng Yang, Taehyun Yang, Zhongzheng Xu.

Figure 1
Figure 1. Figure 1: Timeline of our 15-week course schedule: 13 labs in total, [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 3
Figure 3. Figure 3: Examples from Lab 11. (a) The example created by the instructors using the provided prompts. (b) Student work on the same dataset using vibe coding: some results are highly similar and oth￾ers differ. Similarity is unsurprising here given shared prompts. (c) Student work on the COVID cases dataset. A substantial number of submissions resulted in a bubble chart (e.g., new cases vs. deaths) because they clos… view at source ↗
Figure 4
Figure 4. Figure 4: Examples from Lab 12. We received 37 vibe-coding sub￾missions. Overall, they resembled the instructor’s example, with AI use indicators such as smoothed curves, rounded boxes, and gradi￾ent backgrounds. learn what each part of the code did, and make sure they could complete the task themselves. Several described the lab as rel￾atively small and well-scaffolded, making the detailed instruc￾tions more useful… view at source ↗
Figure 5
Figure 5. Figure 5: Summarized patterns from final projects. We summa￾rize visual patterns that recurred in the 23 projects but not in the previous offering. Because all projects are publicly available online, we redrew these examples to avoid identifying students. student-designed visuals and customized algorithmic heuristics; we perceived a high degree of student agency in this project. Another used affective visualization … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 9 canonical work pages · 5 internal anchors

  1. [1]

    Adiguzel, M

    T. Adiguzel, M. H. Kaya, and F. K. Cansu. Revolutionizing education with ai: Exploring the transformative potential of chatgpt.Contempo- rary educational technology, 15(3), 2023. 1

  2. [2]

    Agarwal, M

    D. Agarwal, M. Naaman, and A. Vashistha. AI Suggestions Ho- mogenize Writing Toward Western Styles and Diminish Cultural Nu- ances. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pp. 1–21, Apr. 2025. doi: 10.1145/3706598. 3713564 2

  3. [3]

    Ahn and N

    Y . Ahn and N. W. Kim. Understanding Why ChatGPT Outperforms Humans in Visualization Design Advice, Aug. 2025. doi: 10.48550/ arXiv.2508.01547 2

  4. [4]

    B. R. Anderson, J. H. Shah, and M. Kreminski. Homogenization Ef- fects of Large Language Models on Human Creative Ideation. InCre- ativity and Cognition, pp. 413–425, June 2024. doi: 10.1145/3635636 .3656204 2

  5. [5]

    Ashkinaze, J

    J. Ashkinaze, J. Mendelsohn, L. Qiwei, C. Budak, and E. Gilbert. How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment. InProceedings of the ACM Collective Intelligence Conference, pp. 198–213, Aug

  6. [6]

    doi: 10.1145/3715928.3737481 2

  7. [7]

    Challenges and Opportunities in Data Visualization Education: A Call to Action

    B. Bach, M. Keck, F. Rajabiyazdi, T. Losev, I. Meirelles, J. Dykes, R. S. Laramee, M. AlKadi, C. Stoiber, S. Huron, C. Perin, L. Morais, W. Aigner, D. Kosminsky, M. Boucher, S. Knudsen, A. Manataki, J. Aerts, U. Hinrichs, J. C. Roberts, and S. Carpendale. Challenges and Opportunities in Data Visualization Education: A Call to Action, Aug. 2023. doi: 10.48...

  8. [8]

    Bower, J

    M. Bower, J. Torrington, J. W. M. Lai, P. Petocz, and M. Alfano. How should we change teaching and assessment in response to increasingly powerful generative Artificial Intelligence? Outcomes of the ChatGPT teacher survey.Education and Information Technologies, Jan. 2024. doi: 10.1007/s10639-023-12405-0 1

  9. [9]

    Dear Diary: A randomized controlled trial of Generative AI coding tools in the workplace

    J. Butler, J. Suh, S. Haniyur, and C. Hadley. Dear Diary: A random- ized controlled trial of Generative AI coding tools in the workplace, Oct. 2024. doi: 10.48550/arXiv.2410.18334 1

  10. [10]

    L. Chen, Y . Song, C. Zheng, Q. Jing, P. Hansen, and L. Sun. Un- derstanding Design Fixation in Generative AI, Feb. 2025. doi: 10. 48550/arXiv.2502.05870 2

  11. [11]

    Z. Chen, C. Zhang, Q. Wang, J. Troidl, S. Warchol, J. Beyer, N. Gehlenborg, and H. Pfister. Beyond Generating Code: Evaluating GPT on a Data Visualization Course. In2023 IEEE VIS Workshop on Visualization Education, Literacy, and Activities (EduVis), pp. 16–21. IEEE, Melbourne, Australia, Oct. 2023. doi: 10.1109/EduVis60792. 2023.00009 1, 2

  12. [12]

    Cheng, J

    Z. Cheng, J. Xu, and H. Jin. TreeQuestion: Assessing conceptual learning outcomes with llm-generated multiple-choice questions.Pro- ceedings of the ACM on Human-Computer Interaction, 8, Nov. 2024. doi: 10.1145/3686970 1

  13. [13]

    Y . Cui, A. M. Goldman, J. Zhou, X. Liu, C. M. Shieh, J. Yao, M. Cheng, M. Kay, and F. Yang. Codesigning ripplet: An LLM- assisted assessment authoring system grounded in a conceptual model of teachers’ workflows. InProceedings of the CHI Conference on Human Factors in Computing Systems. ACM, 2026. doi: 10.1145/ 3772318.3790418 1

  14. [14]

    R. Denkin. On Perception of Prevalence of Cheating and Usage of Generative AI, May 2024. doi: 10.48550/arXiv.2405.18889 1

  15. [15]

    F. Geng, A. Shah, H. Li, N. Mulla, S. Swanson, G. S. Raj, D. Zingaro, and L. Porter. Exploring Student-AI Interactions in Vibe Coding, Nov

  16. [16]

    doi: 10.48550/arXiv.2507.22614 1

  17. [17]

    Inoshita, M

    K. Inoshita, M. Omura, T. Yamanaka, G. Maeda, and K. Tsuji. Does AI Homogenize Student Thinking? A Multi-Dimensional Analysis of Structural Convergence in AI-Augmented Essays, Mar. 2026. doi: 10 .48550/arXiv.2603.21228 2

  18. [18]

    H.-Y . Isa, M. Weston, M. R. Wellyanto, I. Karna, J. O. Talton III, and R. Kumar. From code generation to conceptual learning: Student use of llms in a web programming course. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems, pp. 1–13,

  19. [19]

    Kazemitabaar, R

    M. Kazemitabaar, R. Ye, X. Wang, A. Z. Henley, P. Denny, M. Craig, and T. Grossman. CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs. InProceedings of the CHI Conference on Human Factors in Computing Systems, pp. 1–20, May 2024. doi: 10.1145/ 3613904.3642773 1

  20. [20]

    N. W. Kim, Y . Ahn, G. Myers, and B. Bach. How Good Is CHAT- GPT in Giving Advice on Your Visualization Design?ACM Trans- actions on Computer-Human Interaction, 32(5):1–33, Oct. 2025. doi: 10.1145/3745768 1, 2

  21. [21]

    Will I be replaced?

    M. A. Kuhail, S. S. Mathew, A. Khalil, J. Berengueres, and S. J. H. Shah. “Will I be replaced?” Assessing ChatGPT’s effect on soft- ware development and programmer perceptions of AI tools.Science of Computer Programming, 235:103111, July 2024. doi: 10.1016/j. scico.2024.103111 1

  22. [22]

    Q. Lang, M. Wang, M. Yin, S. Liang, and W. Song. Transforming education with generative ai (gai): Key insights and future prospects. IEEE Transactions on Learning Technologies, 18:230–242, 2025. doi: 10.1109/TLT.2025.3537618 1

  23. [23]

    R. Liu, C. Zenke, C. Liu, A. Holmes, P. Thornton, and D. J. Malan. Teaching CS50 with AI: Leveraging Generative Artificial Intelligence in Computer Science Education. InProceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1, pp. 750–

  24. [24]

    ACM, Portland OR USA, Mar. 2024. doi: 10.1145/3626252. 3630938 1

  25. [25]

    E. N. S. Lockhart. AI is not creative, and the debate is a distraction.AI & Society, 41:4207–4208, 2026. doi: 10.1007/s00146-026-02942-w 6

  26. [26]

    X. Lu, S. Fan, J. Houghton, L. Wang, and X. Wang. ReadingQuiz- Maker: A human-NLP collaborative system that supports instructors 6 to design high-quality reading quiz questions. InProceedings of the CHI Conference on Human Factors in Computing Systems, 2023. doi: 10.1145/3544548.3580957 1

  27. [27]

    W. Lyu, S. Zhang, Tingting, Chung, Y . Sun, and Y . Zhang. Under- standing the Practices, Perceptions, and (Dis)Trust of Generative AI among Instructors: A Mixed-methods Study in the U.S. Higher Edu- cation, Feb. 2025. doi: 10.48550/arXiv.2502.05770 1

  28. [28]

    Q. Ma, H. Shen, K. Koedinger, and T. Wu. How to Teach Program- ming in the AI Era? Using LLMs as a Teachable Agent for Debugging. vol. 14829, pp. 265–279. 2024. doi: 10.1007/978-3-031-64302-6 19 2

  29. [29]

    Mozannar, G

    H. Mozannar, G. Bansal, A. Fourney, and E. Horvitz. Reading Be- tween the Lines: Modeling User Behavior and Costs in AI-Assisted Programming. InProceedings of the CHI Conference on Human Fac- tors in Computing Systems, pp. 1–16. ACM, Honolulu HI USA, May

  30. [30]

    doi: 10.1145/3613904.3641936 1

  31. [31]

    A. Osmani. Loop engineering.https://addyosmani.com/blog/ loop-engineering/, June 2026. Blog post; accessed July 2026. 5

  32. [32]

    Paradis, K

    E. Paradis, K. Grey, Q. Madison, D. Nam, A. Macvean, V . Meimand, N. Zhang, B. Ferrari-Church, and S. Chandra. How Much Does AI Impact Development Speed? an Enterprise-Based Randomized Controlled Trial. In2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE- SEIP), pp. 618–629. IEEE, Ottawa, ON, Canad...

  33. [33]

    Prather, B

    J. Prather, B. Reeves, J. Leinonen, S. MacNeil, A. S. Randrianasolo, B. Becker, B. Kimmel, J. Wright, and B. Briggs. The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers, May 2024. doi: 10.48550/arXiv.2405.17739 1

  34. [34]

    J. C. Roberts, P. Butcher, and P. D. Ritsos. From Data to Insight: Using Contextual Scenarios to Teach Critical Thinking in Data Visualisation, Aug. 2025. doi: 10.48550/arXiv.2508.08737 2

  35. [35]

    Sapkota, K

    R. Sapkota, K. I. Roumeliotis, and M. Karkee. Vibe coding vs. agentic coding: Fundamentals and practical implications of agentic ai, 2025. 1

  36. [36]

    Sarkar and I

    A. Sarkar and I. Drosos. Vibe coding: Programming through conver- sation with artificial intelligence, 2025. 1

  37. [37]

    Stamper, R

    J. Stamper, R. Xiao, and X. Hou. Enhancing LLM-Based Feedback: Insights from Intelligent Tutoring Systems and the Learning Sciences, May 2024. doi: 10.48550/arXiv.2405.04645 1

  38. [38]

    Talk is cheap. Show me the code

    L. Torvalds. Message to the linux-kernel mailing list.https: //lkml.org/lkml/2000/8/25/132, Aug. 2000. “Talk is cheap. Show me the code.”. 6

  39. [39]

    Wright, S

    D. Wright, S. Masud, J. Moore, S. Yadav, M. Antoniak, P. E. Chris- tensen, C. Y . Park, and I. Augenstein. Epistemic Diversity and Knowl- edge Collapse in Large Language Models, Jan. 2026. doi: 10.48550/ arXiv.2510.04226 2 7