REVIEW 4 major objections 3 minor 6 references
A vibe coding learning design to enhance EFL students' talking to, through, and about AI
T0 review · 4 major / 3 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Differences in how students talk to, through, and about AI explain whether vibe coding succeeds.
desk verdict Genuinely useful pilot for teaching vibe coding in EFL, but the causal claim is confounded by platform/access differences; treat the explanation as a hypothesis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The human-AI metalanguaging framework. It names three dimensions—talking to AI (prompt engineering), talking through AI (negotiating authorship and rhetorical load), and talking about AI (mental models of AI)—and treats them as interrelated. The framework is the lens through which the paper reads screen recordings, think-aloud protocols, worksheets, and AI-generated images; the contrast between the two students is meant to show the framework predicts who ships a coherent app.
What would settle it
Run a matched workshop with a dozen or more students, randomly assigning half to structured-prompt training and half to free-form prompting, on the same platform with equal message quotas; the claim would lose support if app functionality and design coherence do not track the training, or if a conversational prompter succeeds while a structured prompter fails.
Extended reading notes
Core claim
The paper's central claim is that AI acts as a beneficial languaging machine in vibe coding: it provokes learners' language, reflection, and self-regulation. More specifically, the authors argue that three interrelated metalanguaging dimensions—prompt engineering (talking to AI), authorship negotiation (talking through AI), and mental models of AI (talking about AI)—shape whether a learner's natural-language app-building succeeds. In the reported contrast, the student who wrote long, structured, one-shot prompts built a functional app matching her intended design, while the student who used short conversational prompts and encountered platform limits ended with a static page that did not mat
Load-bearing premise
The argument depends on being able to separate the three kinds of talking in the data and on those differences causing the different outcomes, rather than prior skill, platform choice, or technical luck doing the work.
Editorial extensions
If this is right
- Vibe coding instruction should teach structured prompt engineering: headings, one-shot examples, and planning prompts before entering them.
- Teachers should ask students to reflect on what their prompt style implies about human versus AI authorship.
- Students need explicit vocabulary to describe their mental models of AI; simple pre/post image-generation tasks can reveal gaps.
- Ambitious design ideas should be paired with pragmatic technical planning so platform limits do not derail the whole project.
- EFL learners can turn authentic writing problems into functional AI-assisted applications within a few hours when the languaging dimensions are supported.
Reading between the lines
- A testable extension is to run the same backward-designed workshop with a larger class, measure the three dimensions separately, and see whether they predict app functionality beyond prior skill and platform luck.
- The three-dimension lens should transfer outside EFL: prompt style and perceived authorship may serve as a quick diagnostic for why any vibe-coding project stalls.
- The students' unchanged minimal image prompts suggest that vibe coding alone may not build mental-model vocabulary; naming AI concepts before the task could produce larger shifts in both prompts and design.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This Innovative Practice article introduces a three-dimension 'human-AI metalanguaging' framework—talking to AI (prompt engineering), talking through AI (negotiating authorship), and talking about AI (mental models)—and applies it to a four-hour EFL vibe-coding workshop. Two Form Three students designed AI-assisted writing apps; one (Student A) produced a functional app on POE's App Creator, while the other (Student G) encountered technical barriers on lovable.dev and submitted a partially functional static page. The authors analyze the contrasting cases through worksheets, think-aloud protocols, screen recordings, and AI-generated images, and argue that differences in how students talked to, through, and about AI explain the outcome variations. They conclude that explicit metalanguaging scaffolding should be part of vibe-coding instruction.
Significance. If the proposed framework is validated, it provides a useful vocabulary for AI-mediated language learning and actionable instructional design for EFL vibe-coding. The paper's strengths include a clearly described workshop design, rich multimodal data collection, transparent think-aloud questions, and concrete examples of prompts and app artifacts. The contrasting-case method is appropriate for a hypothesis-generating pilot. However, the central causal claim—that metalanguaging differences explain success and failure—is not supported by the current evidence: the two cases differ in platform, technical barriers, and access, and the analytic categories derive from the same framework used to design the intervention. The paper is a promising proof-of-concept, but its interpretive claims need to be reframed as hypotheses rather than demonstrated causes.
major comments (4)
- [Abstract and §4] The claim that differences in talking to, through, and about AI 'explain vibe coding outcome variations' is too strong. Section 4 goes further: Student A's success 'can be attributed directly to how she talked to AI.' With only two self-selected students, no comparison group, and no experimental manipulation, a causal attribution is not warranted. The case is confounded by platform differences: Student G used lovable.dev and was blocked by a free daily message limit, upload errors, and inability to access Gemini (§3.2.3), while Student A used POE's App Creator with no such barriers. The alternative explanation—that Student G would have succeeded with a different platform or unlimited quota—is not ruled out. Please reframe the conclusion in terms of association and hypothesis generation.
- [§3] The content analysis lacks a documented coding protocol and inter-rater reliability. The paper states that prompts were analyzed 'in terms of their number, sequence, frequency, length, and content' and that images were 'analyzed and compared,' but no codebook, unitization rules, or reliability indices are provided. Since the central distinctions—structured vs. conversational prompting, high vs. low metalinguistic awareness, mental models as 'precise tool' vs. 'collaborative partner'—are inferential, the absence of reliability evidence weakens the empirical basis for the conclusions. A systematic coding scheme, or at least an explicit acknowledgment that the analysis is illustrative and interpretive, is needed.
- [§2.2 and §3–4] There is a circularity concern: the same three-dimension framework was used to design the learning activities and to define the analytic categories for interpreting the outcomes. For example, the workshop scaffolds 'structured prompt engineering,' and the analysis then finds that Student A used structured prompts and succeeded. This does not invalidate the observation, but it means the framework is partly built into the data. The paper should explicitly acknowledge this design/analysis coupling and consider triangulating with alternative analytic lenses or independent outcome measures to strengthen the interpretation.
- [§4] The cognitive-load explanation for authorship perceptions is post hoc. Student A planned 85% authorship but reported 30–40%, and Student G reported 80% authorship despite a collaborative mental model. The paper attributes this to 'high cognitive load,' but no cognitive-load measure was taken; the inference rests on self-reported affect and task behavior. This is a plausible hypothesis, but it should be presented as such, not as a demonstrated mechanism. A direct measure (e.g., a validated cognitive-load scale) or additional evidence from think-aloud transcripts would be needed to support the claim.
minor comments (3)
- [§3.2.3] The platform name is spelled inconsistently: 'loveable.dev' appears once, while 'lovable.dev' appears later. Please standardize.
- [§3.2.4 and §3.1.4] The pre/post draw-a-picture prompts for Student A are identical, and for Student G only one adds 'Upload the picture.' The paper interprets this as 'misunderstood the task or possessed a limited lexicon,' but it may simply reflect a literal response to an under-specified elicitation. Consider softening this interpretation or probing the students' intentions directly.
- [General] The paper lacks a dedicated limitations section. Given the small sample, self-selection of high-performing students, and the teacher–student relationship (the first author taught both students), a brief limitations paragraph would help calibrate reader expectations and strengthen the innovative-practice framing.
Circularity Check
No circularity: the three-dimension framework is theory-laden but the outcome contrast is independently assessed; causal overreach is a validity threat, not a circular one.
full rationale
The paper's central derivation—that differences in talking to, through, and about AI explain vibe coding outcomes—is not equivalent to its inputs by construction. The three dimensions are defined independently of the outcome variable: talking to AI is operationalized as prompt length, structure, and sequencing; talking through AI as students' self-reported authorship percentages; talking about AI as image prompts and think-aloud mental-model inferences. The outcome (functional app cohering to intended design) is assessed from the submitted app artifacts and design sketches, not from the framework categories. Thus the success/failure contrast provides independent content, and the framework's use in both workshop design and data interpretation is theory-laden observation, not circular reasoning. No parameter is fitted to the outcome and then relabeled as a prediction; no conclusion is forced by a definitional identity. The self-citations (Woo et al., 2024; Guo & Li, 2024) appear in background and a post-hoc cognitive-load explanation, but they are not load-bearing: removing them would not alter the central claim. The paper's real weakness is causal overreach—Student G's failure is confounded with platform choice and lovable.dev's paywall, so 'attributed directly to how she talked to AI' (Section 4) is not warranted. But that is a confound and validity limitation, not a circularity.
Assumptions & free parameters
assumptions (5)
- domain assumption EFL students can learn to build software by writing natural-language prompts to LLMs during a four-hour workshop.
- ad hoc to paper The three metalanguaging dimensions (talking to, through, and about AI) are distinct, interrelated, and observable in worksheets, videos, and prompts.
- domain assumption Students' prompt engineering strategies reflect their mental models of AI.
- domain assumption Reported percentages of authorship are meaningful indicators of 'talking through AI' rather than measurement artifacts.
- ad hoc to paper Cognitive load can explain the paradox that Student A planned 85% authorship but reported 30-40%, and Student G reported 80% authorship despite a collaborative mental model.
invented entities (1)
-
human-AI metalanguaging framework (talking to, through, and about AI)
Cite this review
Pith. "Pith review of A vibe coding learning design to enhance EFL students' talking to, through, and about AI." pith.science (2026). https://pith.science/paper/RMCDEVLK
@misc{pith2026250908854,
author = {Pith},
title = {Pith review of: A vibe coding learning design to enhance EFL students' talking to, through, and about AI},
year = {2026},
howpublished = {\url{https://pith.science/paper/RMCDEVLK}},
note = {Machine review of arXiv:2509.08854}
}
read the original abstract
This innovative practice article reports on the piloting of vibe coding (using natural language to create software applications with AI) for English as a Foreign Language (EFL) education. We developed a human-AI meta-languaging framework with three dimensions: talking to AI (prompt engineering), talking through AI (negotiating authorship), and talking about AI (mental models of AI). Using backward design principles, we created a four-hour workshop where two students designed applications addressing authentic EFL writing challenges. We adopted a case study methodology, collecting data from worksheets and video recordings, think-aloud protocols, screen recordings, and AI-generated images. Contrasting cases showed one student successfully vibe coding a functional application cohering to her intended design, while another encountered technical difficulties with major gaps between intended design and actual functionality. Analysis reveals differences in students' prompt engineering approaches, suggesting different AI mental models and tensions in attributing authorship. We argue that AI functions as a beneficial languaging machine, and that differences in how students talk to, through, and about AI explain vibe coding outcome variations. Findings indicate that effective vibe coding instruction requires explicit meta-languaging scaffolding, teaching structured prompt engineering, facilitating critical authorship discussions, and developing vocabulary for articulating AI mental models.
Figures
Reference graph
Works this paper leans on
-
[1]
Introduction Artificial intelligence (AI) applications for English as a foreign language (EFL) have expanded as teachers and students use large language models (LLMs) with chatbot interfaces like ChatGPT to generate language outputs that support learning (Barrot, 2023). Significantly, LLMs have mastered human and computer languages and can be equipped wit...
work page 2023
-
[2]
Materials and Methods 2.1. Context and Participants 4 We developed the innovative practice as a pilot for an annual human–AI creative writing contest delivering AI and English literacy workshops in Hong Kong secondary schools. The first author piloted the 2025-26 contest learning design in a 4-hour English-medium workshop on July 17, 2025, at an all-girls...
work page 2025
-
[4]
Discussion This study demonstrates that in the context of vibe coding, AI can be a beneficial languaging machine (Jones, 2025), provoking EFL students’ language, reflection, and self-regulation. Simultaneously, the study reveals insights into the complex interplay between how students talk to, through, and about AI. The differences in how Students A and G...
work page 2025
-
[5]
Conclusion This innovative practice has demonstrated vibe coding’s potential in EFL contexts, showing how students use natural language to design software for authentic learning scenarios. The cases illustrate that successful vibe coding is intertwined with metalanguaging practice: how students talk to, through, and about AI. Consequently, effective instr...
-
[6]
https://so04.tci-thaijo.org/index.php/LEARN/article/view/256743 Guo, K., & Li, D. (2024). Understanding EFL students’ use of self-made AI chatbots as personalized writing assistance tools: A mixed methods study. System, 124, 103362. https://doi.org/10.1016/j.system.2024.103362 Harvard Graduate School of Education. (2022). PZ’s Thinking Routines Toolbox. P...
-
[2022]
DT serves as a guiding framework for students to approach vibe coding
and writing (Wible, 2020). DT serves as a guiding framework for students to approach vibe coding. We specified essential understandings through intended learning outcomes (ILOs). For DT ILOs, we adopted those from the Middle Years Program (MYP) design cycle because it emphasizes creative, critical thinking for real-world problems, and a coherent four-stag...
work page 2020
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.